ScalableViT: Rethinking the Context-oriented Generalization of Vision Transformer

Rui Yang, Hailong Ma, Jie Wu, Yansong Tang, Xuefeng Xiao, Min Zheng, Xiu Li

Introduction

Convolutional Neural Networks (CNNs) dominated the computer vision field last few years, which attributes to their capacity in modeling realistic images from a local to global perception. Although they have been widely applied in various vision tasks, there are still deficiencies in global visual perception. This global view is essential for downstream tasks, such as object detection and semantic segmentation. Recently, ViT and its follow-ups employed transformer encoders to address the image task and achieved comparable performance against their CNN counterparts because of the global receptive field. However, the global perception of the Transformer entails an unaffordable computation since self-attention (the primary operation of the Transformer) is quadratically computed on the whole sequence. To alleviate this overhead, typical Swin transformer employed Window-based Self-Attention (WSA), which partitioned a feature map into many non-overlapped sub-regions and enabled it to process large-scale images with linear complexity. They also proposed a novel Shifted Window-based Self-Attention (SWSA) to compensate for losses of potential long-range dependency. Twins combined the WSA with Global Sub-sampled Attention (GSA) for better performance.

To gain an insight into the WSA , we visualize feature maps after the second block. As shown in Fig. 1, features captured by the WSA are dispersed, and their responses incline to partial rather than object-oriented. It may attribute to an invariably fixed dimension that results in limited learning ability, thereby the final performance of the model being highly determined by the difficulty of input data. To alleviate this problem, we develop a novel self-attention mechanism, termed Scalable Self-Attention (SSA), which simultaneously introduces two scaling factors (rnr_{n} and rcr_{c}) to spatial and channel dimensions. Namely, SSA selectively applies these factors to queryquery, keykey, and valuevalue matrices (QQ, KK, and VV), ensuring the dimension is more elastic and no longer deeply bound by the input. On the one hand, SSA aggregates redundant tokens with similar semantic information to a more compact one via spatial scalability. Consequently, unnecessary intermediate multiplication operations are eliminated, and the computational complexity is reduced significantly. In the third row of Fig. 1, we can easily observe that spatial scalability can bring nearly contiguous visual modeling for objects, but some contextual cues are still lost. Hence, on the other hand, we expand the channel dimension to learn a more graphic representation. As depicted in the last row of Fig. 1, SSA successfully obtains complete object activation while maintaining context-oriented generalization via channel scalability. For instance, the contextual cues of the cat in the last column are represented in detail. Such scaling factors also restore the output dimension to align with the input, which makes the residual connection feasible.

Moreover, we propose an Interactive Window-based Self-Attention (IWSA) that consists of a regular WSA and a local interactive module (LIM). The IWSA establishes information connections by re-merging independent valuevalue tokens and aggregating spatial information from adjacent windows. Therefore, it no longer limits the self-attention to local windows, particularly non-overlapping windows. Such characteristic enhances the desired global receptive field and takes good advantage of the most significant superiority of the Transformer in a single layer. The effectiveness of LIM for WSA is validated in Tab. 5(b). To achieve a more efficient backbone for general vision tasks, we adopt a hierarchical design and propose a new Vision Transformer architecture, termed ScalableViT, which alternately arranges IWSA and SSA blocks in each stage.

Main contributions of our ScalableViT lie in two aspects:

For the global self-attention, we propose SSA to supply context-oriented generalization in the vanilla self-attention block, which significantly reduces computational overhead without sacrificing contextual expressiveness.

For the local self-attention, we design LIM to enhance the global perception ability of WSA.

Both SSA and IWSA can model long-range dependency in a single layer instead of stacking more self-attention layers; hence, the ScalableViT is more suitable for visual tasks. We employ ScalableViT on several vision tasks, including image-level classification on ImageNet , pixel-level object detection and instance segmentation on COCO , and semantic segmentation on ADE20K . Extensive experiments demonstrate that the ScalableViT outperforms other state-of-the-art Vision Transformers with similar or less computational cost. For example, ScalableViT-S achieves +1.4% gains against Twins-SVT-S and +1.8% gains against Swin-T on ImageNet-1K classification.

Related Work

The Transformer architecture has become a common template for natural language processing (NLP) tasks due to its solid global modeling capabilities and convenient parallelization ability. Inspired by this, many researchers tried to equip CNNs with the self-attention to modulate and augment outputs of convolutions . DETR employed the self-attention mechanism to model relations between objects for end-to-end detection. Others combined self-attention with convolutions for full-image contextual information. Recently, the emergence of ViT , DeiT , and a series of follow-ups proved the bright prospect of the Vision Transformer.

ViT applied standard Transformer encoders to build a convolution-free image classifier by decomposing the image into a sequence of non-overlapping patches directly. Although it harvested promising results, a gap still existed between data-hungry Transformers and top-performing CNNs when only training on the midsize ImageNet-1K from scratch. In order to bridge this gap, DeiT proposed a token-based distillation procedure and a data-efficient training strategy to optimize the Transformer effectively. Later, the follow-ups improved different aspects of the ViT, making them more suitable for vision tasks. T2T-ViT optimized the tokenization by concatenating the neighboring tokens into one token. DynamicViT pruned the tokens of less importance in a dynamic way for a better lightweight module. Cvt , CeiT incorporated the convolution designs into the self-attention or the FFN to enhance the locality. CPVT utilized the implicit position representation ability from convolutions (with zero padding) to encode the conditional position information for inputs with the arbitrary size. Then, hierarchical pyramid structures were performed by progressively shrinking the number of tokens and replacing the class token with the average pooling. Thus, the Transformer, supported by multi-level features , can handle object detection and image segmentation tasks conveniently. In this paper, we develop a Vision Transformer, ScalableViT, which achieves a better accuracy and cost trade-off on visual tasks.

2 Local Self-Attention

The computational complexity of the self-attention mechanism is a barrier that confines it in only downsampled feature maps or small images. Thus, several previous studies proposed decomposing the global self-attention into much paralleled local self-attention to handle expensive computation burdens. However, this local self-attention limits the receptive field that is critical to dense predict tasks. proposed generating the sparse attention map on a criss-cross path to realize global interaction. captured the information from all the other positions via interlacing elements between different local windows. HaloNet used the overlapped local windows to add the interactions between independent windows. After ViT showed competitiveness, several follow-ups applied the self-attention within non-overlapped local windows for linear computational complexity. To compensate for lost information, Swin Transformer introduced a novel shifted window strategy, and Twins Transformer baked sparse global attention after WSA. We design the IWSA, which can aggregate information from a collection of discrete valuevalue tokens and enable local self-attention to model long-range dependency in a single block.

Method

In this section, we elaborately introduce the architecture of ScalableViT and mainly focus on SSA and IWSA mechanisms. SSA simultaneously introduces different scale factors into spatial and channel dimensions to maintain context-oriented generalization while reducing computational overhead. IWSA enhances the receptive field of local self-attention by aggregating information from a set of discrete valuevalue tokens. Both have linear computational complexity and can learn long-range dependency in a single layer.

The architecture of ScalableViT is illustrated in Fig. 2. For an input image with size H×W×3H\times W\times 3, a convolutional patch embedding layer (7×77\times 7, stride 44) is used to obtain a collection of tokens (H4×W4\frac{H}{4}\times\frac{W}{4}) and project the channel dimension to CC. Then, these initial tokens will pass through four stages which contain a series of Transformer blocks. Between two adjacent stages, another convolutional patch embedding layer (3×33\times 3, stride 22) is utilized to merge tokens and double the channel dimension. For the ithi^{th} stage, there are H2i+1×W2i+1\frac{H}{2^{i+1}}\times\frac{W}{2^{i+1}} input tokens with 2i−1C2^{i-1}C channels and LiL_{i} Transformer blocks. As a result, the quantity of tokens will eventually be reduced to H32×W32\frac{H}{32}\times\frac{W}{32}. This architecture enables us to obtain a hierarchical representation similar to the typical backbones based on CNNs . This merit allows ScalableViT to naturally migrate to various vision tasks, such as object detection and segmentation. In each stage, we devise an alternate arrangement of IW-MSA and S-MSA blocks to organize the topological structure. In the front of each stage, a position encoding generator (PEG) is inserted between two Transformer blocks to generate position embedding dynamically.

2 Scalable Self-Attention

Self-attention is a critical mechanism in the Transformer, and the vanilla self-attention can be calculated as:

Generally, there is much homologous information in natural images, but vanilla self-attention still calculates their similarity. Notably, not all information is necessary to calculate self-attention in the Vision Transformer. For example, similar background tokens should be aggregated as one representative token to attend to other foreground tokens. Namely, the dimension of Q(X)Q(X), K(X)K(X), and V(X)V(X) should not be bounded with the input XX. More importantly, the fixed dimension results in limited learning ability. Thus, we develop the Scalable Self-Attention (SSA), where two scaling factors (rnr_{n} and rcr_{c}) are introduced to spatial and channel dimensions, respectively, resulting a more efficient intermediate calculation than the vanilla one. As illustrated in Fig. 4, the spatial dimension NN and channel dimension CC are selectively scaled to N×rnN\times r_{n} and C×rcC\times r_{c}, respectively, by three transformation functions fq(⋅)f_{q}(\cdot), fk(⋅)f_{k}(\cdot), and fv(⋅)f_{v}(\cdot). These scaling factors can also restore the output dimension to align with the input, making the subsequent FFN layers and residual connections feasible. As a result, the intermediate dimension is more elastic and no longer deeply bound with the input XX. The model can reap context-oriented generalization while dwindling computational overhead significantly. SSA can be naturally written as:

More importantly, the introduced spatial and channel scalability can bring context-oriented generalization (see Fig. 1). If only spatial scalability is introduced (rc≡1r_{c}\equiv 1), there would realize nearly contiguous visual modeling for objects but a lack of critical graphic representation. When further introducing channel scalability, SSA can successfully maintain contextual cues and obtain complete object activation, which is essential in visual tasks. The values of these two scaling factors vary with model configurations and different network stages. As the network gradually deepens, the quantity of tokens shrinks, and the degree of redundancy is also dropped. Thus, rnr_{n} is largen with the stage depth. Similarly, the channel dimension does not always mismatch with spatial dimension in the self-attention operation. Thus, we set rc≥1r_{c}\geq 1 in ScalableViT-S and ScalableViT-B. Because of a too-large channel dimension, we set rc≤1r_{c}\leq 1 in ScalableViT-L. Details about two scale factors are displayed in Table 1.

3 Interactive Window-based Self-Attention

Besides the efficient self-attention , earlier researches have developed the local self-attention to avoid the quadratic computational complexity with the number of tokens. For example, WSA divides an image (H×W×CH\times W\times C) into multiple partial windows which contains M×MM\times M tokens. Then, the self-attention would be calculated in every isolated window and produce a set of discrete outputs {Zn}n=1HM×WM\{Z_{n}\}_{n=1}^{\frac{H}{M}\times\frac{W}{M}}, where ZnZ_{n} can be calculated as:

CoaT also introduced a depth-wise convolution into self-attention. However, they only considered the convolution as a positional encoding method and inserted it deeply into the calculation. If this convolution is expanded into the WSA, it would be limited in the discrete Vn(Xn)V_{n}(X_{n}), which is denoted as local enhanced module (LEM). Differently, we regard our LIM as a matchmaker, which is applied on the spliced valuevalue map VV and parallels with self-attention. By making the sufficient ablation study in Section 4.4, we demonstrate that LIM is capable of delivering stable improvements, especially for downstream tasks.

4 Position Encoding

Besides the position information introduced by LIM, we utilize the positional encoding generator (PEG) , composed of a convolution layer with fixed weights, to acquire implicit positional information. As illustrated in Fig. 2, it is plugged between two consecutive Transformer blocks, with only one in the front of each stage. After the PEG, input tokens are sent to subsequent blocks where position bias could enable the Transformer to realize the input permutation.

5 Architecture Variants

In order to fairly compare with other models under similar computation complexity, we set three models: ScalableViT-S, ScalableViT-B, and ScalableViT-L. The detailed configurations are provided in Table 1, where rcr_{c} and rnr_{n} denote expansion or reduction factors for channel and spatial dimensions, respectively, as described in Section 3.2. Due to the varying representational capability, we set different rcr_{c} for three models. Additionally, the number of blocks, channels, and heads varies with the computational cost.

Experiments

In the following, we compare the proposed model with other state-of-the-art works on ImageNet-1K , COCO , and ADE20K . Then, we conduct ablation studies on the upgraded parts to verify their effectiveness.

Settings. Image classification experiments are conducted on the ImageNet-1K dataset. All settings mainly follow DeiT . During training, we apply data augmentation and regularization strategies in . We employ the AdamW optimizer to train models for 300 epochs from scratch. The learning rate is set to 0.001 initially and varies with the cosine scheduler. The global batchsize is set to 1024 on 8 V100 GPUs. During testing on the validation set, the shorter side of an input image is first resized to 256, and a center crop of 224 × 224 is used to evaluate the classification accuracy.

Result. Classification results on ImageNet-1K are reported in Table 2, where all models are divided into small (around 4G), base (around 9G), and large (around 15G) levels according to computation complexity (FLOPs). ScalableViT-S with a two-layer head outperforms comparable models (1.4%1.4\% better than Twins-SVT-S, and 1.8%1.8\% better than Swin-T). Moreover, it can even approach or exceed other base models. For the base level, ScalableViT-B surpasses Twins-SVT-B by 0.9%0.9\% and SWin-S by 1.1%1.1\% with similar FLOPs. ScalableViT-L also achieves a prominent accuracy-cost trade-off. Additionally, our ScalableViT outperforms the EfficientNet by 0.2%0.2\%, 0.5%0.5\%, and 0.4%0.4\% under three magnitude receptively.

2 Object Detection on COCO

Settings. Object detection experiments are conducted on COCO 2017 dataset. We verify the model effectiveness on RetinaNet and Mask R-CNN detection frameworks using the MMDetection . Before training, we initialize the backbone with the weight pre-trained on ImageNet-1K, FPN with Xavier scheme, and other new layers with Normal scheme (std=0.01std=0.01). All models utilize the same settings as : AdamW optimizer, 1×1\times (12 epochs), and 3×3\times (36 epochs) schedules with a global batchsize of 16 on 8 GPUs. For the 1×1\times schedule, the short side of images is resized to 800 pixels, and the long side is never more than 1333 pixels. The learning rate is declined at the 8th and 11th epoch with a decay rate of 0.10.1. For the 3×3\times schedule, we adopt the multi-scale training, which randomly resizes the short side of images within the range of while keeping the longer side at most 1333. The learning rate is declined at the 27th and 33rd with a decay rate of 0.10.1.

Result. We present results of RetinaNet and Mask R-CNN frameworks in Table 3, where APb\text{AP}^{b} and APm\text{AP}^{m} refer to box mAP and mask mAP, respectively. For object detection with RetinaNet, ScalableViT performs a notable advantage against its CNN and Transformer counterparts. With the 1×1\times schedule, our ScalableViT brings 7.3-8.9 APb\text{AP}^{b} against ResNet at comparable settings. Compared with the popular Swin and Twins Transformers, our ScalableViT performs 3.5-3.7 APb\text{AP}^{b} and 0.5-2.2 APb\text{AP}^{b} improvements, respectively. With the 3×3\times schedule, our ScalableViT still achieves competitive performance. For Mask R-CNN, our ScalableViT-S outperforms ResNet-50 by 7.8 APb\text{AP}^{b} and 7.3 APm\text{AP}^{m} with the 1×1\times schedule. ScalableViT-S achieves 3.6 APb\text{AP}^{b} and 2.6 APm\text{AP}^{m} gains than Swin-T. With the 3×3\times schedule, ScalableViT-S brings 7.7 APb\text{AP}^{b} and 6.5 APm\text{AP}^{m} against ResNet-50. Similarly, it also surpasses Swin-T and Twins-SVT-S Transformers. Under base level, there is also a similar improvement, demonstrating its stronger context-oriented generalization. Additionally, Fig. 5 depicts some qualitative object detection and instance segmentation results from ScalableViT-S-based RetinaNet and Mask R-CNN, which show that contextual representation from the backbone enables the model to detect objects better.

3 Semantic Segmentation on ADE20K

Settings. Semantic segmentation experiments are conducted on the challenging ADE20K dataset. We use the typical Semantic FPN and the UperNet as segmentation frameworks to evaluate our models. We use the MMSegmentation to implement all related experiments, and the settings follow . For the Semantic FPN, we train 80K iterations with a batch size 16 on 4 GPUs. For the UperNet, we train 160K iterations with a batch size 16 on 8 GPUs. During training, we first resize the short side of input images to 512 pixels, and the long side is never more than 2048 pixels, then they are randomly cropped to 512×512512\times 512. During testing, we resize input images as the training phase but without cropping. We also use the test time augmentation for UperNet, including multi-scale test ([0.5,0.75,1.0,1.25,1.5,1.75]×[0.5,0.75,1.0,1.25,1.5,1.75]\times resolution) and flip.

Result. Table 4 reports the segmentation results. For the Semantic FPN, our ScalableViT outperforms Swin Transformer by +3.4 mIoU, +3.2 mIoU, and +3.4 mIoU, respectively, under three FLOPs levels. Compared with CrossFormer-S , ScalableViT-S performs a modest mIoU but has a fewer computation. When equipped into the UperNet, the ScalableViT achieves +4 mIoU, +1.9 mIoU, and +1.6 mIoU gains than Swin Transformer under different model sizes. The same competitive results are achieved when test time augmentation is adopted. In addition, ScalableViT-S outperforms CrossFormer-S by +0.9 mIoU and achieves comparable performance on the base and large size. Fig. 5(c) shows some qualitative results from ScalableViT-S-based Semantic FPN on validation split. These results indicate that the ScalableViT can obtain high-quality semantic segmentation results under contextual-oriented generalization.

4 Ablation Study

Analysis for Self-Attention mechanisms. Our ScalableViT contains two important designs: SSA and IWSA. We ablate their benefits in Table 5(a). Firstly, all attention modules in ScalableViT-S are replaced with the regular window-based self-attention (WSA). Although WSA achieves 82.4%82.4\% top-1 accuracy, the dispersed feature (see Fig. 1) hinders it from better performance on the downstream visual task. Then, we substituted all attention modules with our IWSA and SSA, respectively. Both of them outperform WSA 0.4%0.4\% top-1 accuracy. More importantly, they bring +4.6+4.6 mIoU and +5.5+5.5 mIoU improvements on ADE20K because of the ability modeling long-range dependency. With spatial scalability (rc≡1r_{c}\equiv 1), SSA only achieve 82.6%82.6\% top-1 accuracy and 43.743.7 mIoU. Thus, the context-oriented generalization from the cooperation between spatial and channel scalability plays a critical role in visual tasks. Additionally, we examine the topology by rearranging IWSA and SSA. Results demonstrate that prioritizing IWSA followed by SSA performs best. We also compare IWSA with SWSA in ScalableViT, where our IWSA is more appropriate than SWSA.

Speed analysis. Following , we measure throughput of the ScalableViT-S on single 3090 GPU with a batch size of 64 in Table 6. ScalableViT-S achieves 859.0859.0 img/simg/s, which perform better speed-accuracy trade-offs than Swin-S.

Effectiveness of Local Interactive Module. We examine the effectiveness of LIM in Table 5(b). The ScalableViT-S without position encoding generator (PEG), locally enhanced module (LEM), or LIM is regarded as a baseline model which achieves 82.7%82.7\% top-1 accuracy on ImageNet. Then, three modules are inserted and yield +0.2%+0.2\%, +0.1%+0.1\%, and +0.3%+0.3\% gains than baseline, respectively. It demonstrates that the reasonable convolution can help the model perform better. Due to the window connection, LIM outperforms LEM by +0.2%+0.2\% top-1 accuracy, proving the significance of the information interaction. Additionally, we combine PEG with LEM or LIM, whose results are better than only using a single module. Note that the combination of PEG and LIM outperforms the PEG and LEM under the same overhead. LIM aims to bring global perception into the single Transformer block. Its effectiveness is greatly demonstrated on downstream tasks. Using Semantic FPN with ScalableViT-S on ADE20K, LIM obtains +2.0+2.0 mIoU, and associating PEG with LIM brings +3.2+3.2 mIoU gains.

Conclusion

In this paper, we have presented a Vision Transformer backbone named ScalableViT, composed of two highly effective self-attention mechanisms (SSA and IWSA). SSA employs two cooperated scaling factors in spatial and channel dimensions for context-oriented generalization, which maintains more contextual cues and learns graphic representations. IWSA develops a local interactive module to establish information connections between independent windows. Both of them owns the capability to model long-range dependency in a single layer. The proposed ScalableViT alternately stakes these two self-attention modules. It pushes the whole framework into a more effective trade-off state and achieves state-of-the-art performance on various vision tasks.

Acknowledgements

This work was supported by the National Key R&D Program of China 505 (Grant No.2020AAA0108303), the National Natural Science Foundation of China (Grant No.41876098) and the Shenzhen Science and Technology Project (Grant No.JCYJ20200109143041798).

Appendix 0.A Additional Analyses for IWSA

As shown in Fig. 6, IWSA is composed of a window-based self-attention (WSA) and a local interactive module (LIM). WSA splits the global self-attention into many limited windows and yields a collection of discrete valuevalue matrices. LIM build connections between these valuevalue matrices through a fusion function F\mathcal{F}. In practice, this function is replaced with a 3×33\times 3 depth-wise convolution. Additionally, WSA can be viewed as a 7×77\times 7 depth-wise convolution with an adaptive weight. Thus, F\mathcal{F} brings information exchange through a kind of interleaving effect (illustrated by yellow squares in Fig. 6). This parallel stagger makes IWSA realize a global receptive field in a single layer.

In Table 7, we compare the LIM and the LEM on the ADE20K using Semantic FPN framework. All settings are recorded in the Section 0.C. ScalableViT-S with the LIM achieves +3.8+3.8 mIoU than the LEM under the same overhead because IWSA can model the long-range dependency in single layer. This result also proves that the global receptive field plays a more critical role on the downstream vision task. Moreover, the LIM can be expanded to other window-based self-attention with different window division styles.

Appendix 0.B Comparing visualizations from other blocks

We visualize the feature maps after the 2nd, 4th, and 24th blocks in Figure 7. In the 2nd and 4th blocks, the WSA focuses on local regions, especially the ears and nose. In the latter 24th block, the WSA attends to contextual information but losses some semantic cues. Since feature aggregation from a large downsampling ratio (1616) causes the foreground and background to be poorly separated. By contrast, the SSA can retain a trail of details although the feature map of the later block are not as continuous as the earlier ones.

Appendix 0.C More Implementary Details

Classification. The classification settings mainly follow DeiT . All variants are trained under a resolution of 224×224224\times 224. During training from scratch, we employ the AdamW optimizer with a weight decay of 0.05 and a momentum of 0.9 to train models for 300 epochs. The learning rate is set to 0.001 initially and varies with the cosine scheduler, where a 5-epochs linear warm-up is used to stabilize training. The global batchsize is set to 1024 on 8 V100 GPUs. Moreover, we apply data augmentations and regularizations, including random cropping, random horizontal flipping , mixup , CutMix , random erasing , label-smoothing , stochastic depth , and repeated augmentation . For stochastic depth augmentation, we set the drop rate to 0.20.2, 0.50.5, and 0.50.5 for ScalableViT-S, ScalableViT-B, and ScalableViT-L, respectively. During testing on the validation set, the shorter side of an input image is first resized to 256, and a center crop of 224 × 224 is used to evaluate the classification accuracy.

Object Detection. We adopt RetinaNet and Mask R-CNN detection frameworks on COCO that contains 118K training images and 5K validation images. Before training, we initialize the backbone with the weight pre-trained on ImageNet-1K, FPN with Xavier scheme, and other new layers with Normal scheme (std=0.01std=0.01). All models utilize AdamW optimizer, 500-iteration warm-up, 1×1\times (12 epochs), and 3×3\times (36 epochs) schedule with a global batch size of 16 on 8 GPUs. Settings of initial learning rate and weight decay are shown in Table 8. For 1×1\times schedule, the short side of training images is resized to 800 pixels, and the long side is never more than 1333 pixels. The learning rate is declined at the 8th8th and 11th11th epoch with a decay rate of 0.1. For the 3×3\times schedule, we adopt the multi-scale training, which randomly resizes the short side of the input images within the range of while keeping the longer side at most 1333. The learning rate is declined at the 27th27th and 33rd33rd with a decay rate of 0.1. When testing, the image size is set as the same as the 1×1\times schedule.

Semantic Segmentation. Semantic segmentation experiments are conducted on the challenging ADE20K , with 20K images for training and 2K images for validation. We use the typical Semantic FPN and UperNet as segmentation frameworks to evaluate our models. Following the common practice, we use the MMSegmentation to implement all related experiments. We employ the AdamW to optimize two models. The initial learning rate and weight decay are shown in Table 8. For the Semantic FPN, we train 80K iterations with a batch size 16 on 4 GPUs. The polynomial policy schedules the learning rate with a power of 0.90.9. For the UperNet, we train 160K iterations with a batch size 16 on 8 GPUs. The polynomial policy schedules the learning rate with a power of 1.01.0. During training, we first resize the short side of input images to 512 pixels, and the long side is never more than 2048 pixels, then randomly crop to 512×512512\times 512. During testing, we resize input images the same as the training phase but without cropping. We also use the test time augmentation for UperNet, including multi-scale test ([0.5,0.75,1.0,1.25,1.5,1.75]×[0.5,0.75,1.0,1.25,1.5,1.75]\times resolution) and flip, for better results.

References