TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
Jieneng Chen, Yongyi Lu, Qihang Yu, Xiangde Luo, Ehsan Adeli, Yan Wang, Le Lu, Alan L. Yuille, Yuyin Zhou
Introduction
Convolutional neural networks (CNNs), especially fully convolutional networks (FCNs) , have become dominant in medical image segmentation. Among different variants, U-Net , which consists of a symmetric encoder-decoder network with skip-connections to enhance detail retention, has become the de-facto choice. Based on this line of approach, tremendous success has been achieved in a wide range of medical applications such as cardiac segmentation from magnetic resonance (MR) , organ segmentation from computed tomography (CT) and polyp segmentation from colonoscopy videos.
In spite of their exceptional representational power, CNN-based approaches generally exhibit limitations for modeling explicit long-range relation, due to the intrinsic locality of convolution operations. Therefore, these architectures generally yield weak performances especially for target structures that show large inter-patient variation in terms of texture, shape and size. To overcome this limitation, existing studies propose to establish self-attention mechanisms based on CNN features . On the other hand, Transformers, designed for sequence-to-sequence prediction, have emerged as alternative architectures which employ dispense convolution operators entirely and solely rely on attention mechanisms instead . Unlike prior CNN-based methods, Transformers are not only powerful at modeling global contexts but also demonstrate superior transferability for downstream tasks under large-scale pre-training. The success has been widely witnessed in the field of machine translation and natural language processing (NLP) . More recently, attempts have also matched or even exceeded state-of-the-art performances for various image recognition tasks .
In this paper, we present the first study which explores the potential of transformers in the context of medical image segmentation. However, interestingly, we found that a naive usage (i.e., use a transformer for encoding the tokenized image patches, and then directly upsamples the hidden feature representations into a dense output of full resolution) cannot produce a satisfactory result.
This is due to that Transformers treat the input as 1D sequences and exclusively focus on modeling the global context at all stages, therefore result in low-resolution features which lack detailed localization information. And this information cannot be effectively recovered by direct upsampling to the full resolution, therefore leads to a coarse segmentation outcome. On the other hand, CNN architectures (e.g., U-Net ) provide an avenue for extracting low-level visual cues which can well remedy such fine spatial details.
To this end, we propose TransUNet, the first medical image segmentation framework, which establishes self-attention mechanisms from the perspective of sequence-to-sequence prediction. To compensate for the loss of feature resolution brought by Transformers, TransUNet employs a hybrid CNN-Transformer architecture to leverage both detailed high-resolution spatial information from CNN features and the global context encoded by Transformers. Inspired by the u-shaped architectural design, the self-attentive feature encoded by Transformers is then upsampled to be combined with different high-resolution CNN features skipped from the encoding path, for enabling precise localization. We show that such a design allows our framework to preserve the advantages of Transformers and also benefit medical image segmentation. Empirical results suggest that our Transformer-based architecture presents a better way to leverage self-attention compared with previous CNN-based self-attention methods. Additionally, we observe that more intensive incorporation of low-level features generally leads to a better segmentation accuracy. Extensive experiments demonstrate the superiority of our method against other competing methods on various medical image segmentation tasks.
Related Works
Combining CNNs with self-attention mechanisms. Various studies have attempted to integrate self-attention mechanisms into CNNs by modeling global interactions of all pixels based on the feature maps. For instance, Wang et al. designed a non-local operator, which can be plugged into multiple intermediate convolution layers . Built upon the encoder-decoder u-shaped architecture, Schlemper et al. proposed additive attention gate modules which are integrated into the skip-connections. Different from these approaches, we employ Transformers for embedding global self-attention in our method.
Transformers. Transformers were first proposed by for machine translation and established state-of-the-arts in many NLP tasks. To make Transformers also applicable for computer vision tasks, several modifications have been made. For instance, Parmar et al. applied the self-attention only in local neighborhoods for each query pixel instead of globally. Child et al. proposed Sparse Transformers, which employ scalable approximations to global self-attention. Recently, Vision Transformer (ViT) achieved state-of-the-art on ImageNet classification by directly applying Transformers with global self-attention to full-sized images. To the best of our knowledge, the proposed TransUNet is the first Transformer-based medical image segmentation framework, which builds upon the highly successful ViT.
Method
2 TransUNet
Although combining a Transformer with naive upsampling already yields a reasonable performance, as mentioned above, this strategy is not the optimal usage of Transformers in segmentation since is usually much smaller than the original image resolution , therefore inevitably results in a loss of low-level details (e.g., shape and boundary of the organ). Therefore, to compensate for such information loss, TransUNet employs a hybrid CNN-Transformer architecture as the encoder as well as a cascaded upsampler to enable precise localization. The overview of the proposed TransUNet is depicted in Figure 1.
CNN-Transformer Hybrid as Encoder. Rather than using the pure Transformer as the encoder (Section 3.1), TransUNet employs a CNN-Transformer hybrid model where CNN is first used as a feature extractor to generate a feature map for the input. Patch embedding is applied to patches extracted from the CNN feature map instead of from raw images.
We choose this design since 1) it allows us to leverage the intermediate high-resolution CNN feature maps in the decoding path; and 2) we find that the hybrid CNN-Transformer encoder performs better than simply using a pure Transformer as the encoder.
We can see that CUP together with the hybrid encoder form a u-shaped architecture which enables feature aggregation at different resolution levels via skip-connections. The detailed architecture of CUP as well as the intermediate skip-connections can be found in Figure 1(b).
Experiments and Discussion
Synapse multi-organ segmentation datasethttps://www.synapse.org/#!Synapse:syn3193805/wiki/217789. We use the 30 abdominal CT scans in the MICCAI 2015 Multi-Atlas Abdomen Labeling Challenge, with axial contrast-enhanced abdominal clinical CT images in total.
Each CT volume consists of slices of pixels, with a voxel spatial resolution of . Following , we report the average DSC and average Hausdorff Distance (HD) on 8 abdominal organs (aorta, gallbladder, spleen, left kidney, right kidney, liver, pancreas, spleen, stomach with a random split of 18 training cases (2212 axial slices) and 12 cases for validation.
Automated cardiac diagnosis challengehttps://www.creatis.insa-lyon.fr/Challenge/acdc/. The ACDC challenge collects exams from different patients acquired from MRI scanners. Cine MR images were acquired in breath hold, and a series of short-axis slices cover the heart from the base to the apex of the left ventricle, with a slice thickness of 5 to 8 mm. The short-axis in-plane spatial resolution goes from 0.83 to 1.75 mm2/pixel.
Each patient scan is manually annotated with ground truth for left ventricle (LV), right ventricle (RV) and myocardium (MYO). We report the average DSC with a random split of 70 training cases (1930 axial slices), 10 cases for validation and 20 for testing.
2 Implementation Details
For all experiments, we apply simple data augmentations, e.g., random rotation and flipping.
For pure Transformer-based encoder, we simply adopt ViT with 12 Transformer layers. For the hybrid encoder design, we combine ResNet-50 and ViT, denoted as “R50-ViT”, throught this paper. All Transformer backbones (i.e., ViT) and ResNet-50 (denoted as “R-50”) were pretrained on ImageNet . The input resolution and patch size are set as 224224 and 16, unless otherwise specified. Therefore, we need to cascade four upsampling blocks consecutively in CUP to reach the full resolution. And for Models are trained with SGD optimizer with learning rate 0.01, momentum 0.9 and weight decay 1e-4. The default batch size is 24 and the default number of training iterations are 20k for ACDC dataset and 14k for Synapse dataset respectively. All experiments are conducted using a single Nvidia RTX2080Ti GPU.
Following , all 3D volumes are inferenced in a slice-by-slice fashion and the predicted 2D slices are stacked together to reconstruct the 3D prediction for evaluation.
3 Comparison with State-of-the-arts
We conduct main experiments on Synapse multi-organ segmentation dataset by comparing our TransUNet with four previous state-of-the-arts: 1) V-Net ; 2) DARR ; 3) U-Net and 4) AttnUNet .
To demonstrate the effectiveness of our CUP decoder, we use ViT as the encoder, and compare results using naive upsampling (“None”) and CUP as the decoder, respectively; To demonstrate the effectiveness of our hybrid encoder design, we use CUP as the decoder, and compare results using ViT and R50-ViT as the encoder, respectively. In order to make the comparison with the ViT-hybrid baseline (R50-ViT-CUP) and our TransUNet to be fair, we also replace the original encoder of U-Net and AttnUNet with ImageNet pretrained ResNet-50. The results in terms of DSC and mean hausdorff distance (in mm) are reported in Table 1.
Firstly, we can see that compared with ViT-None, ViT-CUP observes an improvement of and mm in terms of average DSC and Hausdorff distance respectively. This improvement suggests that our CUP design presents a better decoding strategy than direct upsampling. Similarly, compared with ViT-CUP, R50-ViT-CUP also suggests an additional improvement of in DSC and mm in Hausdorff distance, which demonstrates the effectiveness of our hybrid encoder. Built upon R50-ViT-CUP, our TransUNet which is also equipped with skip-connections, achieves the best result among different variants of Transformer-based models.
Secondly, Table 1 also shows that the proposed TransUNet has significant improvements over prior arts, e.g., performance gains range from 1.91% to 8.67% considering average DSC. In particular, directly applying Transformers for multi-organ segmentation yields reasonable results (67.86% DSC for ViT-CUP), but cannot match the performance of U-Net or attnUNet. This is due to that Transformers can well capture high-level semantics which are favorable for classification task but lack of low-level cues for segmenting the fine shape of medical images. On the other hand, combining Transformers with CNN, i.e., R50-ViT-CUP, outperforms V-Net and DARR but still yield inferior results than pure CNN-based R50-U-Net and R50-AttnUNet. Finally, when combined with the U-Net structure via skip-connections, the proposed TransUNet sets a new state-of-the-art, outperforming R50-ViT-CUP and previous best R50-AttnUNet by 6.19% and 1.91% respectively, showing the strong ability of TransUNet to learn both high-level semantic features as well as low-level details, which is crucial in medical image segmentation. A similar trend can be also witnessed for the average Hausdorff distance, which further demonstrates the advantages of our TransUNet over these CNN-based approaches.
4 Analytical Study
To thoroughly evaluate the proposed TransUNet framework and validate the performance under different settings, a variety of ablation studies were performed, including: 1) the number of skip-connections; 2) input resolution; 3) sequence length and patch size and 4) model scaling.
The Number of Skip-connections. As discussed above, integrating U-Net-like skip-connections help enhance finer segmentation details by recovering low-level spatial information. The goal of this ablation is to test the impact of adding different numbers of skip-connections in TransUNet. By varying the number of skip-connections to be 0 (R50-ViT-CUP)/1/3, the segmentation performance in average DSC on all 8 testing organs are summarized in Figure 2. Note that in the “1-skip” setting, we add the skip-connection only at the 1/4 resolution scale. We can see that adding more skip-connections generally leads to a better segmentation performance. The best average DSC and HD are achieved by inserting skip-connections to all three intermediate upsampling steps of CUP except the output layer, i.e., at 1/2, 1/4, and 1/8 resolution scales (illustrated in Figure 1). Thus, we adopt this configuration for our TransUNet. It is also worth mentioning that the performance gain of smaller organs (i.e., aorta, gallbladder, kidneys, pancreas) is more evident than that of larger organs (i.e., liver, spleen, stomach). These results reinforce our initial intuition of integrating U-Net-like skip-connections into the Transformer design to enable learning precise low-level details.
As an interesting study, we apply additive Transformers in the skip-connections, similar to , and find this new type of skip-connection can even further the segmentation performance. Due to the GPU memory constraint, we employ a light Transformer in the 1/8 resolution scale skip-connection while keeping the other two skip-connections unchanged. As a result, this simple alteration leads to a performance boost of 1.4 % DSC.
On the Influence of Input Resolution. The default input resolution for TransUNet is 224224. Here, we also provide results of training TransUNet on a high-resolution 512512, as shown in Table 2. When using 512512 as input, we keep the same patch size (i.e., 16), which results in an approximate 5 larger sequence length for the Transformer. As indicated, increasing the effective sequence length shows robust improvements. For TransUNet, changing the resolution scale from 224224 to 512512 results in 6.88% improvement in average DSC, at the expense of a much larger computational cost. Therefore, considering the computation cost, all experimental comparisons in this paper are conducted with a default resolution of to demonstrate the effectiveness of TransUNet.
On the Influence of Patch Size/Sequence Length.
We also investigate the influence of patch size on TransUNet. The results are summarized in Table 3. It is observed that a higher segmentation performance is usually obtained with smaller patch size. Note that the Transformer’s sequence length is inversely proportional to the square of the patch size (e.g., patch size 16 corresponds to a sequence length of 196 while patch size 32 has a shorter sequence length of 49), therefore decreasing the patch size (or increasing the effective sequence length) shows robust improvements, as the Transformer encodes more complex dependencies between each element for longer input sequences. Following the setting in ViT , we use 1616 as the default patch size throughout this paper.
Model Scaling. Last but not least, we provide ablation study on different model sizes of TransUNet. In particular, we investigate two different TransUNet configurations, the “Base” and “Large” models. For the “base” model, the hidden size , number of layers, MLP size, and number of heads are set to be 12, 768, 3072, and 12, respectively while those hyperparamters for “large” model are 24, 1024, 4096, and 16. From Table 4 we conclude that larger model results in a better performance. Considering the computation cost, we adopt “Base” model for all the experiments.
5 Visualizations
We provide qualitative comparison results on the Synapse dataset, as shown in Figure 3. It can be seen that: 1) pure CNN-based methods U-Net and AttnUNet are more likely to over-segment or under-segment the organs (e.g., in the second row, the spleen is over-segmented by AttnUNet while under-segmented by UNet), which shows that Transformer-based models, e.g., our TransUNet or R50-ViT-CUP have stronger power to encode global contexts and distinguish the semantics. 2) Results in the first row show that our TransUNet predicts fewer false positives compared to others, which suggests that TransUNet would be more advantageous than other methods in suppressing those noisy predictions. 3) For comparison within Transformer-based models, we can observe that the predictions by R50-ViT-CUP tend to be coarser than those by TransUNet regarding the boundary and shape (e.g., predictions of the pancreas in the second row). Moreover, in the third row, TransUNet correctly predicts both left and right kidneys while R50-ViT-CUP erroneously fills the inner hole of left kidney. These observations suggest that TransUNet is capable of finer segmentation and preserving detailed shape information. The reason is that TransUNet enjoys the benefits of both high-level global contextual information and low-level details, while R50-ViT-CUP solely relies on high-level semantic features. This again validates our initial intuition of integrating U-Net-like skip-connections into the Transformer design to enable precise localization.
6 Generalization to Other Datasets
To show the generalization ability of our TransUNet, we further evaluate on other imaging modalities, i.e., an MR dataset ACDC aiming at automated cardiac segmentation. We observe consistent improvements of TransUNet over pure CNN-based methods (R50-UNet and R50-AttnUnet) and other Transformer-based baselines (ViT-CUP and R50-ViT-CUP), which are similar to previous results on the Synapse CT dataset.
Conclusion
Transformers are known as architectures with strong innate self-attention mechanisms. In this paper, we present the first study to investigate the usage of Transformers for general medical image segmentation. To fully leverage the power of Transformers, TransUNet was proposed, which not only encodes strong global context by treating the image features as sequences but also well utilizes the low-level CNN features via a u-shaped hybrid architectural design. As an alternative framework to the dominant FCN-based approaches for medical image segmentation, TransUNet achieves superior performances than various competing methods, including CNN-based self-attention methods.
Acknowledgements. This work was supported by the Lustgarten Foundation for Pancreatic Cancer Research.