Self-Supervised Learning with Swin Transformers

Zhenda Xie, Yutong Lin, Zhuliang Yao, Zheng Zhang, Qi Dai, Yue Cao, Han Hu

Introduction

The vision field is undergoing two revolutionary trends since about two years ago. The first trend is self-supervised visual representation learning pioneered by MoCo , which for the first time demonstrated superior transferring performance on seven downstream tasks over the previous standard supervised methods by ImageNet-1K classification. The second is the Transformer-based backbone architecture , which has strong potential to replace the previous standard convolutional neural networks such as ResNet . The pioneer work is ViT , which demonstrated strong performance on image classification by directly applying the standard Transformer encoder in NLP on non-overlapping image patches. The follow-up work, DeiT , tuned several training strategies to make ViT work well on ImageNet-1K image classification. While ViT/DeiT are designed for the image classification task and has not been well tamed for downstream tasks requiring dense prediction, Swin Transformer is proposed to serve as a general-purpose vision backbone by introducing useful inductive biases of locality, hierarchy and translation invariance.

While the two revolutionary waves appeared independently, the community is curious about what kind of adaptation is needed and what it will behave when they meet each other. Nevertheless, until very recently, a few works started to explore this space: MoCo v3 presents a training recipe to let ViT perform reasonably well on ImageNet-1K linear evaluation; DINO presents a new self-supervised learning method which shows good synergy with the Transformer architecture.

Although these works produce encouraging results on ImageNet-1K linear evaluation, there are no assessment of the transferring performance on downstream tasks such as object detection and semantic segmentation, probably due to that ViT/DeiT are not well tamed for these downstream tasks. To enable more comprehensive evaluations of the self-supervised learnt representations on also these downstream tasks, we propose to adopt Swin Transformer as the backbone architecture instead of the previous used ViT architecture, thanks to that Swin Transformer is designed as general-purpose and performs strong on downstream tasks.

In addition to this backbone architecture change, we also present a self-supervised learning approach by combining MoCo v2 and BYOL , named MoBY (by picking the first two letters of each). We tune a training recipe to make the approach performing reasonably high on ImageNet-1K linear evaluation: 72.8% top-1 accuracy using DeiT-S with 300-epoch training which is slightly better than that in MoCo v3 and DINO but with lighter tricks. Using Swin-T architecture instead of DeiT-S, it achieves 75.0% top-1 accuracy with 300-epoch training, which is 2.2% higher than that using DeiT-S. Initial study shows that some tricks in MoCo v3 and DINO are also useful for MoBY, e.g. replacing the LayerNorm layers before the MLP blocks by BatchNorm like that in MoCo v3 bring additional +1.1% gains using 100 epoch training, indicating the strong potential of MoBY.

When transferred to downstream tasks of COCO object detection and ADE20K semantic segmentation, the representations learnt by this self-supervised learning approach achieves on par performance compared to the supervised method. Noting self-supervised learning with ResNet architectures has shown significantly stronger transferring performance on downstream tasks than supervised methods , the results indicate large space to improve for self-supervised learning with Transformers.

The proposed approach basically has no new inventions. What we provide is an approach which combines the previous good practice but with lighter tricks, associated with tuned hyper-parameters to achieve reasonably high accuracy on ImageNet-1K linear evaluation. We also provide baselines to aid the evaluation of transferring performance on downstream tasks for the future study of self-supervised learning on Transformer architectures.

A Baseline SSL Method with Swin Transformers

MoBY is a combination of two popular self-supervised learning approaches: MoCo v2 and BYOL . It inherits the momentum design, the key queue, and the contrastive loss used in MoCo v2, and inherits the asymmetric encoders, asymmetric data augmentations and the momentum scheduler in BYOL. We name it MoBY by picking the first two letters of each method.

The MoBY approach is illustrated in Figure 1. There are two encoders: an online encoder and a target encoder. Both two encoders consist of a backbone and a projector head (2-layer MLP), and the online encoder introduces an additional prediction head (2-layer MLP), which makes the two encoders asymmetric. The online encoder is updated by gradients, and the target encoder is a moving average of the online encoder by momentum updating in each training iteration. A gradually increasing momentum updating strategy is applied for on the target encoder: the value of momentum term is gradually increased to 1 during the course of training. The default starting value is 0.99.

A contrastive loss is applied to learn the representations. Specifically, for an online view qq, its contrastive loss is computed as

where k+k_{+} is the target feature for the other view of the same image; kik_{i} is a target feature in the key queue; τ\tau is a temperature term; KK is the size of the key queue (4096 by default).

In training, like most Transformer-based methods, we also adopt the AdamW optimizer, in contrast to previous self-supervised learning approaches built on ResNet backbone where usually SGD or LARS is used. We also introduce a regularization method of asymmetric drop path which proves crucial for the final performance.

In the experiments, we adopt a fixed learning rate of 0.001 and a fixed weight decay of 0.05, which performs stably well. We tune hyper-parameters of the key queue size KK, the starting momentum value of the target branch, the temperature τ\tau, and the drop path rates.

A pseudo code of MoBY in a PyTorch-like style is shown in Algorithm 1.

Swin Transformer as the backbone

Swin Transformer is a general-purpose backbone for computer vision and achieved state-of-the-art performance on various vision tasks such as COCO object detection (58.7 box AP and 51.1 mask AP on test-dev set) and ADE20K semantic segmentation (53.5 mIoU on validation set). It is basically a hierarchical Transformer whose representation is computed with shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection.

In this work, we adopt the tiny version of Swin Transformer (Swin-T) as our default backbone, such that the transferring performance on downstream tasks of object detection and semantic segmentation can be also evaluated. The Swin-T has similar complexity with ResNet-50 and DeiT-S. The details of specific architecture design and hyper-parameters can be found in .

Experiments

Linear evaluation on ImageNet-1K dataset is a common evaluation protocol to assess the quality of learnt representations . In this protocol, a linear classifier is applied on the backbone, with the backbone weights frozen and only the linear classifier trained. After training this linear classifier, the top-1 accuracy using center crop is reported on the validation set.

During training, we follow to use random resize cropping with scale from [0.08,1][0.08,1] and horizontal flipping as the data augmentation. 100-epoch training with a 5-epoch linear warm-up stage is conducted. The weight decay is set as 0. The learning rate is set as the optimal one of {0.5,0.75,1.0,1.25}\{0.5,0.75,1.0,1.25\} through grid search for each pre-trained model.

Table 1 listed the major results of pre-trained models using different self-supervised learning methods and backbone architectures.

Regarding previous methods such as MoCo v3 and DINO adopt ViT/DeiT as their backbone architecture, we first report results of MoBY using DeiT-S for fair comparison with them. Under 300-epoch training, MoBY achieves 72.8% top-1 accuracy, which is slightly better than MoCo v3 and DINO (without the multi-crop trick), as shown in Table 1.

We note that MoCo v3 and DINO adopt heavy tricks to achieve the same accuracy as ours:

Tricks in MoCo v3 . MoCo v3 adopts a fixed patch embedding, batch normalization layers to replace the layer normalization ones before the MLP blocks, and a 3-layer MLP head. It also uses large batch size (i.e. 4096) which is unaffordable for many research labs.

Tricks in DINO . DINO adopts asymmetric temperatures between student and teacher, a linearly warmed-up teacher temperature, varying weight decay during pre-training, the last layer fixed at the first epoch, tuning whether to put weight normalization in the head, a concatenation of the last few blocks or CLS tokens as the input to the linear classifier, and etc.

In contrast, we mainly adopt standard settings from MoCo v2 and BYOL , and use a small batch size of 512 such that the experimental settings will be affordable for most labs. We have also started to try applying some tricks of MoCo v3 /DINO to MoBY, though they are not included in the standard settings. Our initial exploration reveals that the fixed patch embedding has no use to MoBY, and replacing the layer normalization layers before the MLP blocks by batch normalization can bring +1.1% top-1 accuracy using 100-epoch training, as shown in Table 2. This indicates that some of these tricks may be useful for the MoBY approach, and the MoBY approach has potential to achieve much higher accuracy on ImageNet-1K linear evaluation. This will be left as our future study.

Swin-T v.s. DeiT-S

We also compare the use of different Transformer architectures in self-supervised learning. As shown in Table 1, Swin-T achieves 75.0% top-1 accuracy, surpassing DeiT-S by +2.2%. Also note the performance gap is larger than that of using supervised learning (+1.5%).

2 Transferring Performance on Downstream Tasks

We evaluate the transferring performance of the learnt representation on downstream tasks of COCO object detection/instance segmentation and ADE20K semantic segmentation.

Two detectors are adopted in the evaluation: Mask R-CNN and Cascade Mask R-CNN , following the implementation of https://github.com/SwinTransformer/Swin-Transformer-Object-Detection. Table 3 shows the comparison of the learnt representation by MoBY and the pretrained supervised method in , in both 1x and 3x settings. For each experiment, we follow all the settings used for supervised pre-trained models , except that we tune the drop path rate in {0,0.1,0.2}\{0,0.1,0.2\} and report the best results (for also supervised models).

It can be seen that the representations learnt by the self-supervised method (MoBY) and the supervised method are similarly well on transferring performance. While we note that previous SSL works using ResNet as the backbone architecture usually report stronger performance over the supervised methods , no gains over supervised methods are observed using Transformer architectures. We hypothesis it is partly because the supervised pre-training on Transformers has involved strong data augmentations , while supervised training of ResNet usually employs much weaker data augmentation. These results also imply space to improve for self-supervised learning using Transformer architectures.

ADE20K Semantic Segmentation

The UPerNet approach and the ADE20K dataset are adopted in the evaluation, following https://github.com/SwinTransformer/Swin-Transformer-Semantic-Segmentation. The fine-tuning and testing settings also follow except that the learning rate of each experiment is tuned using {3×10−5,6×10−5,1×10−4}\{3\times 10^{-5},6\times 10^{-5},1\times 10^{-4}\}. Table 4 shows the comparisons of supervised and self-supervised pre-trained models on this evaluation. It indicates that MoBY performs slightly worse than the supervised method, implying a space to improve for self-supervised learning using Transformer architectures.

3 Ablation Study

We perform ablation study using the ImageNet-1K linear evaluation protocol. Swin-T is used as the backbone architecture. In each ablation, we vary one hyper-parameter and other hyper-parameters are set as the default ones.

Drop path has proved a useful regularization for supervised representation learning using the image classification task and Transformer architectures . We also ablate the effect of this regularization in Table 5. Increasing the drop path regularization from 0.05 to 0.1 to the online encoder is beneficial for representation learning, especially in longer training, probably due to the relief of over-fitting. Additionally adding drop path regularization to the target encoder results in 1.9% top-1 accuracy drop (70.9% to 69.0%), indicating a harm. We thus adopt an asymmetric drop path rates in pre-training.

Other hyper-parameters

Table 6(a) ablates the effect of key queue size KK from 1024 to 16384. The approach stably performs across various KK (from 1024 to 16384), and we adopt 4096 as default. Table 6(b) ablates the effect of temperature τ\tau and 0.2 performs best which is set as the default value. Table 6(c) ablates the effect of the starting momentum value of the target encoder. 0.99 performs best and is set as the default value.

Conclusion

In this paper, we present a self-supervised learning approach called MoBY, with Vision Transformers as its backbone architecture. With a proper training recipe and much lighter tricks than MoCo v3/DINO, MoBY can achieve reasonably high performance on ImageNet-1K linear evaluation: 72.8% and 75.0% top-1 accuracy using DeiT-S and Swin-T, respectively, by 300-epoch training. More importantly, in contrast to ViT/DeiT, the general-purpose Swin Transformer backbone enables us to also evaluate the learnt representations on downstream tasks such as object detection and semantic segmentation. MoBY can perform comparably or slightly worse than the supervised methods, indicating a space to improve for self-supervised learning with Transformer architectures. We hope our results can facilitate more comprehensive evaluation of self-supervised learning methods designed for Transformer architectures. Our code and models are available and will be continually enriched at https://github.com/SwinTransformer/Transformer-SSL.

References