PVT v2: Improved Baselines with Pyramid Vision Transformer

Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, Ling Shao

Introduction

Recent studies on vision Transformer are converging on the backbone network designed for downstream vision tasks, such as image classification, object detection, instance and semantic segmentation. To date, there have been some promising results. For example, Vision Transformer (ViT) first proves that a pure Transformer can archive state-of-the-art performance in image classification. Pyramid Vision Transformer (PVT v1) shows that a pure Transformer backbone can also surpass CNN counterparts in dense prediction tasks such as detection and segmentation tasks . After that, Swin Transformer , CoaT , LeViT , and Twins further improve the classification, detection, and segmentation performance with Transformer backbones.

This work aims to establish stronger and more feasible baselines built on the PVT v1 framework. We report that three design improvements, namely (1) linear complexity attention layer, (2) overlapping patch embedding, and (3) convolutional feed-forward network are orthogonal to the PVT v1 framework, and when used with PVT v1, they can bring better image classification, object detection, instance and semantic segmentation performance. The improved framework is termed as PVT v2. Specifically, PVT v2-B5 PVT v2 has 6 different size variants, from B0 to B5 according to the parameter number. yields 83.8% top-1 error on ImageNet, which is better than Swin-B and Twins-SVT-L , while our model has fewer parameters and GFLOPs. Moreover, GFL with PVT-B2 archives 50.2 AP on COCO val2017, 2.6 AP higher than the one with Swin-T , 5.7 AP higher than the one with ResNet50 . We hope these improved baselines will provide a reference for future research in vision Transformer.

Related Work

We mainly discuss transformer backbones related to this work. ViT treats each image as a sequence of tokens (patches) with a fixed length, and then feeds them to multiple Transformer layers to perform classification. It is the first work to prove that a pure Transformer can also archive state-of-the-art performance in image classification when training data is sufficient (e.g., ImageNet-22k , JFT-300M). DeiT further explores a data-efficient training strategy and a distillation approach for ViT.

To improve image classification performance, recent methods make tailored changes to ViT. T2T ViT concatenates tokens within an overlapping sliding window into one token progressively. TNT utilizes inner and outer Transformer blocks to generate pixel and patch embeddings respectively. CPVT replaces the fixed size position embedding in ViT with conditional position encodings, making it easier to process images of arbitrary resolution. CrossViT processes image patches of different sizes via a dual-branch Transformer. LocalViT incorporates depth-wise convolution into vision Transformers to improve the local continuity of features.

To adapt to dense prediction tasks such as object detection, instance and semantic segmentation, there are also some methods to introduce the pyramid structure in CNNs to the design of Transformer backbones. PVT v1 is the first pyramid structure Transformer, which presents a hierarchical Transformer with four stages, showing that a pure Transformer backbone can be as versatile as CNN counterparts and performs better in detection and segmentation tasks. After that, some improvements are made to enhance the local continuity of features and to remove fixed size position embedding. For example, Swin Transformer replaces fixed size position embedding with relative position biases, and restricts self-attention within shifted windows. CvT , CoaT , and LeViT introduce convolution-like operations into vision Transformers. Twins combines local attention and global attention mechanisms to obtain stronger feature representation.

Methodology

There are three main limitations in PVT v1 as follows: (1) Similar to ViT , when processing high-resolution input (e.g., shorter side being 800 pixels), the computational complexity of PVT v1 is relatively large. (2) PVT v1 treats an image as a sequence of non-overlapping patches, which loses the local continuity of the image to a certain extent; (3) The position encoding in PVT v1 is fixed-size, which is inflexible for process images of arbitrary size. These problems limit the performance of PVT v1 on vision tasks.

To address these issues, we propose PVT v2, which improves PVT v1 through three designs, which are listed in Sec 3.2, 3.3, and 3.4.

2 Linear Spatial Reduction Attention

First, to reduce the high computational cost caused by attention operations, we propose linear spatial reduction attention (SRA) layer as illustrated in Fig. 1. Different from SRA which uses convolutions for spatial reduction, linear SRA uses average pooling to reduce the spatial dimension (i.e., h×wh\times w) to a fixed size (i.e., P×PP\times P) before the attention operation. So linear SRA enjoys linear computational and memory costs like a convolutional layer. Specifically, given an input of size h×w×ch\times w\times c, the complexity of SRA and linear SRA are:

where RR is the spatial reduction ratio of SRA . PP is the pooling size of linear SRA, which is set to 7.

3 Overlapping Patch Embedding

Second, to model the local continuity information, we utilize overlapping patch embedding to tokenize images. As shown in Fig. 2(a), we enlarge the patch window, making adjacent windows overlap by half of the area, and pad the feature map with zeros to keep the resolution. In this work, we use convolution with zero paddings to implement overlapping patch embedding. Specifically, given an input of size h×w×ch\times w\times c, we feed it to a convolution with the stride of SS, the kernel size of 2S−12S-1, the padding size of S−1S-1, and the kernel number of c′c^{\prime}. The output size is hS×wS×C′\frac{h}{S}\times\frac{w}{S}\times C^{\prime}.

4 Convolutional Feed-Forward

Third, inspired by , we remove the fixed-size position encoding , and introduce zero padding position encoding into PVT. As shown in Fig. 2(b), we add a 3×33\times 3 depth-wise convolution with the padding size of 1 between the first fully-connected (FC) layer and GELU in feed-forward networks.

5 Details of PVT v2 Series

We scale up PVT v2 from B0 to B5 By changing the hyper-parameters. which are listed as follows:

SiS_{i}: the stride of the overlapping patch embedding in Stage ii;

CiC_{i}: the channel number of the output of Stage ii;

LiL_{i}: the number of encoder layers in Stage ii;

RiR_{i}: the reduction ratio of the SRA in Stage ii;

PiP_{i}: the adaptive average pooling size of the linear SRA in Stage ii;

NiN_{i}: the head number of the Efficient Self-Attention in Stage ii;

EiE_{i}: the expansion ratio of the feed-forward layer in Stage ii;

Tab. 1 shows the detailed information of PVT v2 series. Our design follows the principles of ResNet . (1) the channel dimension increase while the spatial resolution shrink with the layer goes deeper. (2) Stage 3 is assigned to most of the computation cost.

6 Advantages of PVT v2

Combining these improvements, PVT v2 can (1) obtain more local continuity of images and feature maps; (2) process variable-resolution input more flexibly; (3) enjoy the same linear complexity as CNN.

Experiment

Settings. Image classification experiments are performed on the ImageNet-1K dataset , which comprises 1.28 million training images and 50K validation images from 1,000 categories. All models are trained on the training set for fair comparison and report the top-1 error on the validation set. We follow DeiT and apply random cropping, random horizontal flipping , label-smoothing regularization , mixup , and random erasing as data augmentations. During training, we employ AdamW with a momentum of 0.9, a mini-batch size of 128, and a weight decay of 5×10−25\times 10^{-2} to optimize models. The initial learning rate is set to 1×10−31\times 10^{-3} and decreases following the cosine schedule . All models are trained for 300 epochs from scratch on 8 V100 GPUs. We apply a center crop on the validation set to benchmark, where a 224×\times 224 patch is cropped to evaluate the classification accuracy.

Results. In Tab. 2, we see that PVT v2 is the state-of-the-art method on ImageNet-1K classification. Compared to PVT, PVT v2 has similar flops and parameters, but the image classification accuracy is greatly improved. For example, PVT v2-B1 is 3.6% higher than PVT v1-Tiny, and PVT v2-B4 is 1.9% higher than PVT-Large.

Compared to other recent counterparts, PVT v2 series also has large advantages in terms of accuracy and model size. For example, PVT v2-B5 achieves 83.8% ImageNet top-1 accuracy, which is 0.5% higher than Swin Transformer and Twins , while our parameters and FLOPS are fewer.

2 Object Detection

Settings. Object detection experiments are conducted on the challenging COCO benchmark . All models are trained on COCO train2017 (118k images) and evaluated on val2017 (5k images). We verify the effectiveness of PVT v2 backbones on top of mainstream detectors, including RetinaNet , Mask R-CNN , Cascade Mask R-CNN , ATSS , GFL , and Sparse R-CNN . Before training, we use the weights pre-trained on ImageNet to initialize the backbone and Xavier to initialize the newly added layers. We train all the models with batch size 16 on 8 V100 GPUs, and adopt AdamW with an initial learning rate of 1×10−41\times 10^{-4} as optimizer. Following common practices , we adopt 1×\times or 3×\times training schedule (i.e., 12 or 36 epochs) to train all detection models. The training image is resized to have a shorter side of 800 pixels, while the longer side does not exceed 1,333 pixels. When using the 3×\times training schedule, we randomly resize the shorter side of the input image within the range of $$. In the testing phase, the shorter side of the input image is fixed to 800 pixels.

Results. As reported in Tab. 3, PVT v2 significantly outperforms PVT v1 on both one-stage and two-stage object detectors with similar model size. For example, PVT v2-B4 archive 46.1 AP on top of RetinaNet , and 47.5 APb on top of Mask R-CNN , surpassing the models with PVT v1 by 3.5 AP and 4.6 APb, respectively. We present some qualitative object detection and instance segmentation results on COCO val2017 in Fig. 3, which also shows the good performance of our models.

For a fair comparison between PVT v2 and Swin Transformer , we keep all settings the same, including ImageNet-1K pre-training and COCO fine-tuning strategies. We evaluate Swin Transformer and PVT v2 on four state-of-the-arts detectors, including Cascade R-CNN , ATSS , GFL , and Sparse R-CNN . We see PVT v2 obtain much better AP than Swin Transformer among all the detectors, showing its better feature representation ability. For example, on ATSS, PVT v2 has similar parameters and flops compared to Swin-T, but PVT v2 achieves 49.9 AP, which is 2.7 higher than Swin-T. Our PVT v2-Li can largely reduce the computation from 258 to 194 GFLOPs, while only sacrificing a little performance.

3 Semantic Segmentation

Settings. Following PVT v1 , we choose ADE20K to benchmark the performance of semantic segmentation. For a fair comparison, we test the performance of PVT v2 backbones by applying it to Semantic FPN . In the training phase, the backbone is initialized with the weights pre-trained on ImageNet , and the newly added layers are initialized with Xavier . We optimize our models using AdamW with an initial learning rate of 1e-4. Following common practices , we train our models for 40k iterations with a batch size of 16 on 4 V100 GPUs. The learning rate is decayed following the polynomial decay schedule with a power of 0.9. We randomly resize and crop the image to 512×512512\times 512 for training, and rescale to have a shorter side of 512 pixels during testing.

Results. As shown in Tab. 5, when using Semantic FPN for semantic segmentation, PVT v2 consistently outperforms PVT v1 and other counterparts. For example, with almost the same number of parameters and GFLOPs, PVT v2-B1/B2/B3/B4 are at least 5.3% higher than PVT v1-Tiny/Small/Medium/Large. Moreover, although the GFLOPs of PVT-Large are 12% lower than those of ResNeXt101-64x4d, the mIoU is still 8.5 points higher (48.7 vs 40.2). In Fig. 3, we also visualize some qualitative semantic segmentation results on ADE20K . These results demonstrate that PVT v2 backbones can extract powerful features for semantic segmentation, benefiting from the improved designs.

4 Ablation Study

Ablation experiments of PVT v2 is reported in Tab. 6. We see that all three designs can improve the model in terms of performance, parameter number, or computation overhead.

Overlapping patch embedding (OPE) is important. Comparing #1 and #2 in Tab. 6, the model with OPE obtains better top-1 accuracy (81.1% vs. 79.8%) on ImageNet and better AP (42.2% vs. 40.4%) on COCO than the one with original patch embedding (PE) . OPE is effective because it can model the local continuity of images and feature map via the overlapping sliding window.

Convolutional feed-forward network (CFFN) matters. Compared to original feed-forward network (FFN) , our CFFN contains a zero-padding convolutional layer. which can capture the local continuity of the input tensor. In addition, due to the positional information introduced by zero-padding in OPE and CFFN, we can remove the fixed-size positional embeddings used in PVT v1, making the model flexible to handle variable resolution inputs. As reported in #2 and #3 in Tab. 6, CFFN brings 0.9 points improvement on ImageNet (82.0% vs. 81.1%) and 2.4 points improvement on COCO, which demonstrates its effectiveness.

Linear SRA (LSRA) contributes to a better model. As reported in #3 and #4 in Tab. 6, compared to SRA , our LSRA significantly reduces the computation overhead (GFLOPs) of the model by 22%, while keeping a comparable top-1 accuracy on ImageNet (82.1% vs. 82.0%), and only 1 point lower AP on COCO (43.6 vs. 44.6). These results show the low computational cost and good effect of LSRA.

4.2 Computation Overhead Analysis

As shown in Figure 4, with increasing input scale, the GFLOPs growth rate of the proposed PVT v2-B2-Li is much lower than that of PVT v1-Small , and is similar to that of ResNet-50 . This result proves that our PVT v2-Li successfully addresses the high computational overhead problem caused by the attention layer.

Conclusion

We study the limitations of Pyramid Vision Transformer (PVT v1) and improve it with three designs, which are overlapping patch embedding, convolutional feed-forward network, and linear spatial reduction attention layer. Extensive experiments on different tasks, such as image classification, object detection, and semantic segmentation demonstrate that the proposed PVT v2 is stronger than its predecessor PVT v1 and other state-of-the-art transformer-based backbones, under comparable numbers of parameters. We hope these improved baselines will provide a reference for future research in vision Transformer.

References