CrossFormer++: A Versatile Vision Transformer Hinging on Cross-scale Attention
Wenxiao Wang, Wei Chen, Qibo Qiu, Long Chen, Boxi Wu, Binbin Lin, Xiaofei He, Wei Liu
Introduction
Transformer based vision backbones have achieved great success in many computer vision tasks such as image classification, object detection, semantic segmentation, etc. Compared with convolutional neural networks (CNNs), vision transformers enable long range dependencies by introducing a self-attention module, endowing the models with higher capacities.
A transformer requires a sequence of tokens (e.g., word embeddings) as input. To adapt this requirement to typical vision tasks, most existing vision transformers produce tokens by splitting an input image into equally-sized patches. For example, a image can be split into patches of size , and these patches are embedded through a linear layer to yield a token sequence. Inside a certain transformer, self-attention is engaged to build the interactions between any two tokens. Thus, the computational or memory cost of the self-attention module is , where is the length of a token sequence. Such a cost is too big for a visual input because its token sequence is much longer than that of NLP tasks. Therefore, the recently proposed vision transformers develop multiple substitutes (e.g., Swin restricts the self-attention in a small local region instead of a global region) to approximate the vanilla self-attention module with a lower cost.
Though the aforementioned vision transformers have made progress, they still face challenges in utilizing visual features of different scales, whereas multi-scale features are very vital for a lot of vision tasks. Particularly, an image often contains many objects of different sizes, and to fully understand the image, the model is required to extract features at different scales (i.e., different ranges and sizes). Existing vision transformers fail to deal with the above case due to two reasons: (1) Those models’ input tokens are generated from equally-sized patches. Though these patches theoretically have a chance to extract any scale features if only the patch size is large enough, it is difficult to promise that they can learn appropriate multi-scale features automatically in practice. (2) Some vision transformers such as Swin restrict the attention in a small local region, giving up the long range attention.
In this paper, we propose a cross-scale embedding layer (CEL) and a long-short distance attention (LSDA) to fill the cross-scale gap from two perspectives:
Cross-scale Embedding Layer (CEL). Following PVT , we also employ a pyramid structure for our designed transformer, which naturally splits the vision transformer model into multiple stages. CEL appears at the start of each stage, which receives last stage’s output (or an input image) as input and samples patches with multiple kernels of different scales (e.g., 4×4 or 8 × 8). Then, each token is constructed by embedding and concatenating these patches. Through this way, we enforce some dimensions (e.g., from 4 × 4 patches) to focus on small-scale features only, while others (e.g., from 8 × 8 patches) have a chance to learn large-scale features, leading to a token with explicit cross-scale features.
Long-Short Distance Attention (LSDA). Instead of using a shifted window based self-attention like Swin, we argue that both short distance and long distance attentions are imperative for a visual input and thereby propose our LSDA. In particular, we split the self-attention module into a Short Distance Attention (SDA) and a Long Distance Attention (LDA). SDA builds the dependencies among neighboring embeddings, while LDA takes charge of the dependencies among embeddings far away from each other. SDA and LDA appear alternately in consecutive layers of a CrossFormer, which also reduce the cost of the self-attention module while keep both the short distance and long distance attentions.
Dynamic Position Bias (DPB). Besides, following prior work , we employ a relative position bias for tokens’ position representations. The Relative Position Bias (RPB) only supports fixed image/group size. However, image size for many vision tasks such as object detection is variable, so does group size for many architectures, including ours. To make the RPB more flexible, we further introduce a trainable module called Dynamic Position Bias (DPB), which receives two tokens’ relative distance as input and outputs their position bias. The DPB module is optimized end-to-end in the training phase, inducing an ignorable cost but making RPB apply to variable group size.
Armed with the above three modules, we name our architecture as CrossFormer and design four variants of different depths and channels. While these variants achieve great performance, we also find that a further expansion of CrossFormer does not bring a persistent accuracy gain, but even a degradation. To this end, we analyze the output of each layer (i.e., the self-attention maps and the MLPs) and observe another two issues which affect the models’ performance, i.e., the enlarging self-attention maps and amplitude explosion. Thus, we further propose a progressive group size (PGS) and an amplitude cooling layer (ACL) to alleviate the problems.
Progressive Group Size (PGS). In terms of the self-attention maps, we observe that tokens at shallow layers always attend to tokens around themselves. Whereas tokens at deep layers pay nearly equal attention to all other tokens. This phenomenon shows that vision transformers perform similarly as CNN, i.e., extracting local features at shallow layers and global features at deep layers, respectively. It will hinder the models’ performance if adopting a fixed group size for all stages like most existing vision transformers. Hence, we propose to enlarge the group size progressively (PGS) from shallow to deep layers and implement a manually designed group size policy under this guide. Though it is a simple policy, PGS is actually a general paradigm. We hope to present its importance and appeal other researchers to explore more automated and adaptive group size policies.
Amplitude Cooling Layer (ACL). In the vision transformers, the activation’s amplitude grows dramatically as the layer goes deeper. For example, in a CrossFormer-B, the maximal amplitude in the 22nd layer is 300 times larger than that in the 1st layer. The tremendous amplitude discrepancy among layers makes the training process unstable. However, we also find that our proposed CEL can effectively suppress the amplitude. Considering that it is a bit cumbersome to put more CELs into a CrossFormer, we design a CEL-like but more lightweight layer dubbed Amplitude Cooling Layer (ACL). ACL is inserted after some CrossFormer blocks to cool down the amplitude.
We improve CrossFormer by introducing PGS and ACL, yielding CrossFormer++, and propose four new variants. Extensive experiments on four downstream tasks (i.e., image classification, object detection, semantic segmentation, and instance segmentation) show that CrossFormer++ outperforms CrossFormer and other existing vision transformers on all these tasks.
The following sections are organized as follows: the background and related works will be introduced in Sec. 2. Then, we will retrospect CrossFormerAn earlier version of this paper has appeared in ICLR 2022: https://openreview.net/pdf?id=_PHymLIxuI and its CEL, LSDA, and DPB in Sec. 3. PGS, ACL, and CrossFormer++ are introduced in detail in Sec. 4. Thereafter, the experiments will be shown in Sec. 5. Finally, we will present our conclusions and future work in Sec. 6.
Background and Related Works
Vision Transformers. Motivated by the great success achieved by transformers in NLP, researchers have tried to design specific visual transformers for vision tasks to take advantage of the powerful attention mechanism. Early vision transformers like ViT and DeiT transfer the original transformer to vision tasks, achieving impressive results and demonstrating great potentials. T2T-ViT and VOLO inherit the ViT’s architecture while improving its input tokens. They introduce locality into tokens, which is more suitable for visual input. Furthermore, VOLO also incorporates Token Labeling and achieves state-of-the-art results on several downstream vision tasks. Later works like PVT , Swin , and MViTv2 combine the pyramid structure with the transformer and remove the class token used in the original architecture, enabling the use in further vision tasks like object detection and image segmentation. As is mentioned in MViTv2 , such a pyramid structure is partly inspired by CNN, which hierarchically expands the channel capacity while reducing the spatial resolution. Our work proposes that other insights from CNN may also be valuable for transformers. To be specific, the lower layers in the network tend to refine local features while the higher layers focus more on global information communications.
Self-supervised ViTs. In addition to architectural design, another field of ViT is exploring the self-supervised pre-training scheme to enhance its performance. Most self-supervised ViT fall into two categories: masked image modeling and contrastive learning. Masked image modeling such as MAE , data2vec , and CAE gives the model an masked image, and the ViT learns to predict the masked parts (pixels or hidden representations). Some papers also adapts masked image modeling to pyramid ViTs. The belief behind masked image modeling is that the model can predict the masked parts only if it understands the image and extracts good features. While contrastive learning generates different views for each image. The model is training to predict whether these views are from the same or different image. Self-supervised pretraining is orthogonal to architectural design and they can be combined to further improve the ViT’s performance.
Efficient Self-Attention and Universal Vision Backbone. Vanilla self-attention module in the transformer suffers from a quadratic complexity with respect to the image size, which is unacceptable for dense vision tasks like object detection and image segmentation. Therefore, in order to build transformer-based universal vision backbones, researchers have proposed ways to do efficient self-attention. The first choice is to do sparse self-attention. Instead of doing self-attention among all embeddings, some transformers divide the embeddings into groups and do self-attention within each group. The ways of choosing the group while also enabling global information interactions become the core design for works like Swin , CAT , CvT , and CSwin . Another choice is to reduce the cost of global self-attention through reducing the size of input queries and keys. PVT uses average pooling to reduce keys size, while Scalable ViT uses convolution with large intervals. MViTv2 inserts a pooling layer after the linear transformation layer of the self-attention module. It should also be noted that the results reported in the above works are sometimes obtained under different settings such as different depths and embedding methods, which disables the direct comparison between different ways to divide the group. Our work proposes a novel way of group division and uses the same experiment setting and architecture to compare the effectiveness of different division ways.
Position Representations. Transformers are permutation invariant models, which are unsatisfying for both NLP and vision tasks, where the permutation of the embeddings affects the semantic information. To make the model aware of position information, many different position representations are proposed. For example, Dosovitskiy et al. (2021) directly added the embeddings with the vectors that contain absolute position information, while Relative Position Bias (RPB) (Shaw et al., 2018) shows that relative position information is more important for vision tasks and resorts to position information indicating the relative distance of two embeddings. In contrast, MaxViT proposes that depth-wise convolutions can also be seen as a kind of conditional position representations, thus no explicit position representation is needed. Besides, Xiangxiang Chu et al. pointed out that a successful positional encoding for vision tasks should meet the requirements that: (1) being permutation-variant but translation-invariant; (2) being able to handle different lengths of inputs; (3) containing absolute position information. Based on such requirements, they further proposed conditional positional encoding (CPE) that is dynamically generated and conditioned on the local neighborhood of the input tokens. CPE is used in latest transformers like MaxViT. Furthermore, since it could be implemented by a depthwise convolution layer, CPE is able to be combined with other designs like Squeeze-and-Excitation layer to remove explicit positional encodings layers.
Neural Network Design. There have been a number of works in CNN aimed at designing architectures that achieve a good trade-off between efficiency and accuracy. Works such as X3D have proposed greedy methods to search for good hyperparameters like spatial resolution and depth for each stage of a CNN, which attach importance to the design choice other than a novel convolution design. However, to the best of our knowledge, such an exploration in vision transformers is clearly insufficient. AutoFormer and S3 (also known as AutoFormerV2) uses Neural Architecture Search (NAS) methods to find good embedding dimension for each self-attention layer. GLiT searches a good permutation of global and local self-attention modules in the transformer. Anyway, most works that focus on transformer architecture design fix their group size as 7 in order to keep pace with Swin . Such choices are made out of convenience and clearly not the best. Our work makes a small step forward through pointing out that choosing good group size for embeddings allows the transformers to achieve a better efficiency-accuracy tradeoff.
CrossFormer
The overall architecture of CrossFormer is plotted in Fig. 1. Following , CrossFormer also employs a pyramid structure, which naturally splits the transformer model into four stages. Each stage consists of a cross-scale embedding layer (CEL, Sec. 3.1) and several CrossFormer blocks (Sec. 3.2). A CEL receives last stage’s output (or an input image) as input and generates cross-scale tokens through an embedding layer. In this process, CEL (except that in Stage-1) reduces the number of embeddings to a quarter while doubles their dimensions for a pyramid structure. Then, several CrossFormer blocks, each of which involves long short distance attention (LSDA) and dynamic position bias (DPB), are set up after CEL. A specialized head (e.g., the classification head in Fig. 1) follows after the final stage accounting for a specific task.
Cross-scale embedding layer (CEL) is leveraged to generate input tokens for each stage. Fig. 3 takes the first CEL, which is ahead of Stage-1, as an example. It receives an image as input, then sampling patches using four kernels of different sizes. The stride of four kernels is kept the same so that they generate the same number of tokensThe image will be padded if necessary.. As we can observe in Fig. 3, every four corresponding patches have the same center but different scales, and all these four patches will be embedded and concatenated as one embedding. In practice, the process of sampling and embedding can be fulfilled through four convolutional layers.
For a cross-scale token, one problem is how to set the embedded dimension of each scale. An intuitive way is allocating the dimension equally. Take a 96-dimensional token with four kernels as an example, each kernel outputs a 24-dimensional vector and concatenating these vectors yields a 96-dimensional token. However, the computational budget of a convolutional layer is proportional to , where and represent kernel size and input/output dimension, respectively (assuming that the input dimension equals to the output dimension). Therefore, given the same dimension, a large kernel consumes a greater budget than a smaller one. To control the total budget of the CEL, we use a lower dimension for large kernels while a higher dimension for small kernels. Fig. 3 provides the specific allocation rule in its subtable, and a -dimensional example is given. Compared with allocating the dimension equally, our scheme saves much computational cost but does not explicitly affect the model’s performance. The cross-scale embedding layers in other stages work in a similar way. As shown in Fig. 1, CELs for Stage-2/3/4 use two different kernels ( and ). Further, to form a pyramid structure, the strides of CELs for Stage-2/3/4 are set as to reduce the number of embeddings to a quarter.
2 CrossFormer Block
Each CrossFormer block consists of a long short distance attention module (i.e., LSDA, which involves a short distance attention (SDA) module or a long distance attention (LDA) module) and a multi-layer perceptron (MLP). As shown in Fig. 1b, SDA and LDA appear alternately in different blocks, and the dynamic position bias (DPB) module works in both SDA and LDA for obtaining embeddings’ position representations. Following the prior vision transformers, residual connections are used in each block.
We split the self-attention module into two parts: short distance attention (SDA) and long distance attention (LDA). For SDA, all adjacent embeddings are grouped together. Fig. 2a gives an example where . For LDA with input of size , the embeddings are sampled with a fixed interval . For example in Fig. 2b (), all embeddings with a red border belong to a group, and those with a yellow border comprise another group. The group’s height or width for LDA is computed as (i.e., in this example). After grouping embeddings, both SDA and LDA employ the vanilla self-attention within each group. As a result, the memory/computational cost of the self-attention module is reduced from to .
It is worth noting that the effectiveness of LDA also benefits from cross-scale embeddings. Specifically, we draw all the patches comprising two embeddings in Fig. 2b. As we can see, the small-scale patches of two embeddings are non-adjacent, so it is difficult to judge their relationship without the help of the context. In other words, it will be hard to build the dependency between these two embeddings if they are only constructed by small-scale patches (i.e., single-scale feature). On the contrary, adjacent large-scale patches provide sufficient context to link these two embeddings, which makes long-distance cross-scale attention easier to compute and more meaningful.
2.2 Dynamic Position Bias (DPB)
Relative position bias (RPB) indicates embeddings’ relative positions by adding a bias to their attentions. Formally, the LSDA’s attention map with RPB becomes:
The structure of DPB is displayed in Fig. 2c. Its non-linear transformation consists of three fully-connected layers with layer normalization and ReLU . The input dimension of DPB is , i.e., , and intermediate layers’ dimension is set as , where is the dimension of embeddings. The output is a scalar, encoding the relative position feature between the and embeddings. DPB is a trainable module optimized along with the whole transformer model. It can deal with any image/group size without worrying about the bound of .
3 Variants of CrossFormer
TABLE I lists the detailed configurations of CrossFormer’s four variants for image classification. To re-use the pre-trained weights, the models for other tasks (e.g., object detection) employ the same backbones as classification except that they may use different and . Specifically, besides the configurations same to classification, we also test with , and for the detection (and segmentation) models’ first two stages to adapt to larger images. Notably, the group size or the interval (i.e., or ) does not affect the shape of weight tensors, so the backbones pre-trained on ImageNet can be readily fine-tuned on other tasks even though they use different or .
CrossFormer++
In this section, we propose a progressive group size (PGS) paradigm to adapt to vision transformers’ gradually expanding group size. Moreover, an amplitude cooling layer (ACL) is proposed to alleviate the amplitude explosion issue. The PGS and ACL are plugged into CrossFormer, yielding an improved version, dubbed CrossFormer++.
Existing work has explored the vanilla ViTs’ mechanisms and found that ViTs prefer global attentions even from early layers, which work in a different way from CNNs. However, vision transformers with a pyramid structure take smaller patches as input, resorting to group self-attention, and may perform differently from the vanilla ViTs. We take CrossFormer as an example and compute its average attention maps. Specifically, the attention maps of a certain group can be represented as:
where represent batch size, number of heads, and group size, respectively. It means that there are tokens in all, and that the attention map for each token is of size . The attention map of a token at image , head , and position is represented as:
For the token at position , the attention map averaged over batches and multi-heads is:
We train a CrossFormer-B with a large group size of and compute the average attention of each token. The results of a random token’s are shown in Fig. 4. As we can see, the attention regions gradually expand from shallow to deep layers. For example, tokens in the first two stages mainly attend to regions of size around themselves, while the attention maps of deep layers from stage-3 are evenly distributed. The results indicate that tokens from shallow layers prefer local dependencies, while those from deep layers prefer global attentions.
To this end, we propose a PGS paradigm, i.e., adopting a smaller group size in shallow layers to lower the computational budget and a larger group size in deep layers for global attentions. Under this guide, we first empirically set the group size to The last stage’s group size decreases because its feature maps size is , and a group size of already means a global attention. for four stages, respectively. Then, a linear scaling group size policy is proposed, i.e., expanding the group size linearly from shallow to deep layers.
Previous work S3 has proposed a design guideline similar to PGS. However, the intuitions behind and the details are different. The guideline from S3 is inspired by the phenomenon observed during the process of NAS. The search space of group size they use is limited to only two integers, i.e. , and the effect of different group sizes is modeled with linear approximation, which is relatively coarser compared with PGS. Contrarily, PGS gets intuition from attention matrix visualization and adopts a more aggressive and flexible strategy that enables different transformer blocks to choose different group size from a set of consecutive integers within interval $$, e.g. linear scaling group size policy. We compare experiment results and provide detailed ablation experiments to prove the effectiveness of PGS in section 5. The combination of PGS and NAS is left for future research.
2 Amplitude Cooling Layer (ACL)
In addition to attention maps, we also explore the output’s amplitude of each block. As shown in Fig. 5, for a CrossFormer-B model, the amplitude increases greatly as the block goes deeper. In particular, the maximal output for the 22nd block becomes over 3000, about 300 times larger than the value for the 1st block. Besides, the average amplitude of the 22nd also becomes about 15 times larger than that of the 1st block. The extreme value makes the training process unstable and hinders the model from converging.
Fortunately, we also observe that the amplitude shrinks back to a small value at the start of each stage (e.g., the and blocks). We think that all block’s outputs are gradually accumulated through the residual connections in the model. The CEL at the beginning of each stage does not have a residual connection and cuts off the accumulation process, so it can effectively cool down the amplitude. While CEL still contains cumbersome normal convolution layers, we propose a more lightweight counterpart, dubbed amplitude cooling layer (ACL). As shown in Fig. 6, similar to CEL, an ACL does not use any residual connection, either. It only consists of a depth-wise convolution layer and a normalization layer. The comparison in Fig. 5 shows that ACL can also cool down the amplitude, but it introduces fewer parameters and a less computational budget than CEL because a depth-wise convolution with a small kernel (instead of a normal convolutional layer) is used.
However, ACL without a residual connection will prolong the back-propagation path and aggravate the vanishing gradient issue. To prevent this, we put an ACL layer after each blocks with , as shown in Fig. 6. Empirically, achieves a satisfying trade-off between amplitude cooling and back-propagation.
3 Variants of CrossFormer++
Armed with PGS and ACL, we further improve CrossFormer and propose CrossFormer++. The architectures are listed as TABLE I. Wherein, CrossFormer++-B and CrossFormer++-L inherit each stage’s depth and channel from CrossFormer-B and CrossFormer-L, respectively. Besides, CrossFormer++ focuses more on larger models, so the “Tiny (-T)” version is ignored, and a new “Huge (-H)” version is constructed. Particularly, CrossFormer++-H puts more layers in the first two stages for refined low-level features. Moreover, we find that for the “Small (-S)” version, a deep slim model performs similar to a shallow wide one but introduces fewer parameters, so CrossFormer++-S resorts to a deeper model than CrossFormer-S while using fewer channels for each block.
Experiments
The experiments are carried out on four challenging tasks: image classification, object detection, instance segmentation, and semantic segmentation. To entail a fair comparison, we keep the same data augmentation and training settings as the other vision transformers as far as possible. The competitors are all competitive vision transformers, including DeiT , PVT , T2T-ViT , TNT , CViT , Twins , Swin , S3 , NesT , CvT , ViL , CAT , ResT , TransCNN , Shuffle , BoTNet , RegionViT , ViTAEv2 , MPViT , ScalableViT , DaViT , and CoAtNet .
Experimental Settings. The experiments on image classification are conducted on the ImageNet dataset. It contains 1.28M natural images for training and 50,000 images for evaluation. The images are resized to for both training and evaluation by default. The same training settings as the other vision transformers are adopted. In particular, we use an AdamW optimizer training for 300 epochs with a cosine decay learning rate scheduler, and 20 epochs of linear warm-up are used. The batch size is 1,024 split on 8 V100 GPUs. An initial learning rate of 0.001 and a weight decay of 0.05 are used. Besides, we use drop path rate of for CrossFormer-T, CrossFormer-S, CrossFormer-B, and CrossFormer-L, respectively. As well, the drop path rates of are used for CrossFormer++-S, CrossFormer++-B, CrossFormer++-L, and CrossFormer++-H, respectively. Further, similar to Swin , RandAugment , Mixup , Cutmix , random erasing , and stochastic depth are used for data augmentation.
Results. The results are shown in TABLE II. CrossFormer outperforms all other contemporaneous vision transformers. In specific, compared against strong baselines DeiT, PVT, and Swin, our CrossFormer outperforms them at least absolute in accuracy on small models. Compared with CrossFormer, CrossFormer++ brings about 0.8 accuracy improvement in average. For example, CrossFormer-B achieves 83.4 in accuracy, while CrossFormer++-B reaches 84.2 with a negligible extra computational budget. Moreover, CrossFormer++ outperforms all existing vision transformers with similar parameters and a comparable computational budget and throughput.
2 Object Detection and Instance Segmentation
Experimental Settings. The experiments on object detection and instance segmentation are both done on the COCO 2017 dataset , which contains K training and K val images. We use MMDetection-based RetinaNet and Mask R-CNN as the object detection and instance segmentation head, respectively. For both tasks, the backbones are initialized with the weights pre-trained on ImageNet. Then the whole models are trained with batch size on V100 GPUs, and an AdamW optimizer with an initial learning rate of is used. Following previous works, we adopt training schedule (i.e., the models are trained for epochs) when taking RetinaNets as detectors, and images are resized to 800 pixels for the short side. While for Mask R-CNN, both and training schedules are used. It is noted that multi-scale training is also employed when taking training schedules.
Results. The results on RetinaNet and Mask R-CNN are shown in TABLE III and TABLE IV, respectively. As we can see, CrossFormer outperforms most existing vision transformers and shows higher superiority than on image classification task. We think it is because CrossFormer utilizes cross-scale features explicitly, and cross-scale features are particularly vital for dense prediction tasks like objection detection and instance segmentation. However, there are still some other architectures that perform better than CrossFormer, e.g., PVTv2-B3 achieves AP higher than CrossFormer-S. Instead, CrossFormer++ surpasses CrossFormer by at least AP and outperforms all existing methods. Further, its performance gain over the other architectures gets sharper when enlarging the model, indicating that CrossFormer++ enjoys greater potentials. For example, CrossFormer++-S outperforms ScalableViT-S by 0.8 in AP (48.7 vs. 49.5) when using Mask R-CNN as the detection head, while the CrossFormer++-B outperforms ScalableViT-B by 1.2 in AP.
3 Semantic Segmentation
Experimental Settings. ADE20K is used as the benchmark for semantic segmentation. It covers a broad range of semantic categories, including K images for training and K for validation. Similar to models for detection, we initialize the backbones with weights pre-trained on ImageNet, and MMSegmentation-based semantic FPN and UPerNet are taken as the segmentation head. For FPN , we use an AdamW optimizer with learning rate and weight deacy of . Models are trained for K iterations with batch size . For UPerNet, an AdamW optimizer with an initial learning rate of and a weight decay of is used, and models are trained for K iterations.
Results. All results are shown in TABLE V. Compared with architectures released earlier, CrossFormer exhibits a greater performance gain over the others when enlarging the model, which is similar to object detection. For example, CrossFormer-T achieves absolutely higher in IOU than Twins-SVT-B, but CrossFormer-B achieves absolutely higher in IOU than Twins-SVT-L. Totally, CrossFormer shows a more significant advantage over the others on dense prediction tasks (e.g., detection and segmentation) than on classification, implying that cross-scale interactions in the attention module are more important for dense prediction tasks than for classification.
Moreover, CrossFormer++ shows a more significant advantage than CrossFormer when using a more powerful segmentation head. As we can see in TABLE V, UPerNet is a more powerful model than semantic FPN. CrossFormer++S outperforms CrossFormer-S by 1.0 AP when using semantic FPN (47.4 vs. 46.4), while outperforms 2.0 AP when using UPerNet as the segmentation head.
4 Ablation Studies
We conduct the experiments by replacing cross-scale embedding layers with single-scale ones. As we can see in TABLE VI, when using single-scale embeddings, the kernel in Stage-1 brings ( vs. ) absolute improvement compared with the kernel. It tells us that overlapping receptive fields help improve the model’s performance. Besides, all models with cross-scale embeddings perform better than those with single-scale embeddings. In particular, our CrossFormer achieves ( vs. ) absolute performance gain compared with using single-scale embeddings for all stages. For cross-scale embeddings, we also try several different combinations of kernel sizes, and they all show similar performance (). In summary, cross-scale embeddings can bring a large performance gain, yet the model is relatively robust with respect to different choices of kernel size.
In addition, we also test different dimension allocation schemes for CEL. As described in Fig. 3, we allocate different number of dimensions for different sampling kernels, and we also compare with “allocating equally”. The experiments are conducted with CrossFormer-S and the results are in Table VII. The results indicate that the two allocation schemes achieve similar accuracy while our scheme owns less parameters.
4.2 LSDA vs. Other Self-attentions
Self-attention mechanisms used in other vision transformers are also compared. Specifically, we replace LSDA in CrossFormer-S with self-attention modules proposed in previous work, and CEL and DPB are retained (except for PVT because DPB cannot be applied to PVT). As shown in TABLE VIII, the self-attention mechanisms have great influences on the models’ accuracies. In particular, LSDA outperforms CAT-like self-attention by 1.4% (82.5% vs. 81.1%). Besides, some self-attention mechanisms can achieve similar accuracy to LSDA (CSwin and ScalableViT), which indicates that the most proper self-attention mechanism is not unique.
4.3 Ablation Studies about PGS and ACL
Manually Designed Group Size. We adopt manually designed group size for the CrossFormer++, i.e., for four stages, respectively. We also test other group size, and the results are shown in TABLE IX. Compare CrossFormer++-S with CrossFormer++-S3 and CrossFormer++-S4, and the results show that a group size for the first two stages is inessential, which leads to a greater computational budget but no accuracy improvement. The comparison between CrossFormer++-S and CrossFormer++-S1 (83.2% vs 82.6%) shows that a large group size for Stage-3 is pivotal. These conclusions also coincide with our observation that the self-attention at shallow layers concentrates on a small region around each token, while the attention gradually disperses at deep layers.
Linearly Scaling Group Size. Additionally, we also try a simple non-manually designed group size, i.e., we expand group size from to linearly from Stage-1 to Stage-3. However, the linear group size does not work as well as a manually designed group size.
Attention Mechanisms vs. Group Size. It is also worth emphasizing that the conclusions about group size also apply to other vision transformers. As shown in TABLE IX, an appropriate group size for Swin-T also brings a significant accuracy improvement, from 81.3% to 82.8%. In comparison, replacing shifted window in Swin with LSDA only brings 0.6% improvement (as shown in TABLE VIII), which indicates that an appropriate group size may be more important than the choice of the self-attention mechanism.
Ablation Studies about ACL. Regarding ACL, the ACL layer brings about 0.3%0.4% accuracy improvement. Concretely, CrossFormer++-B improves from 83.9% to 84.2% after plugging an ACL, and CrossFormer++-L improves from 84.3% to 84.7%. Moreover, ACL is a universal layer that also applies to other vison transformers. For example, plugging an ACL into Swin-T can bring 0.4% accuracy improvement (82.8% vs. 83.2%).
4.4 DPB vs. Other Position Representations
We compare the parameters, FLOPs, throughputs, and accuracies of the models among absolute position embedding (APE), relative position bias (RPB), and DPB. The results are shown in TABLE X. DPB-residual means DPB with residual connections. Both DPB and RPB outperform APE with absolute accuracy, which indicates that relative position representations are more beneficial than absolute ones.
Further, DPB achieves the same accuracy () as RPB with an ignorable extra cost; however, as we described in Sec. 3.2.2, it is more flexible than RPB and applies to variable image size or group size. Besides, the results also show that residual connection in DPB does not help improve or even degrades the model’s performance (82.5% vs. 82.4%).
Moreover, we design two interpolation strategies to adapt RPB to variable group size, called offline interpolated RPB and online interpolated RPB:
Offline interpolated RPB: After training the model on the ImageNet with a small group size (e.g., ), we first interpolate RPB to a sufficient large size (e.g., applying to at most group size for the COCO dataset). During the fine-tuning process, the enlarged RPB is fine-tuned along the whole model. For the inference stage on downstream tasks, the enlarged RPB is fixed and no interpolation is required.
Online interpolated RPB: The RPB is kept as the original (e.g., fit to group size) after training. During the fine-tuning process, we dynamically resize RPB with a differentiable bilinear interpolation for each image. Though kept as a fixed small size, the RPB is fine-tuned to apply to variable group size. For the inference stage on downstream tasks, the RPB needs to do an online interpolation for each image.
As object detection and instance segmentation are the most common tasks with variable input image sizes, experiments are done on the COCO dataset, and results are shown in Table XI. As we can see, DPB outperforms both offline and online interpolated DPB. Offline interpolated RPB is 1.5 lower than DPB in absolute APb. We assume it is because there are too many different group sizes (), resulting in each of them has limited training samples, so the interpolated RPB may be underfitting. In contrast, online interpolation RPB encodes performs a little better than the offline version, but it is still not as profitable as DPB. Moreover, since online interpolation RPB executes an online interpolation for each image, its throughput is on par with DPB, which is a bit smaller than that of offline interpolated RPB (16 imgs/sec vs. 14 imgs/sec). The throughput indicates that the extra computational budgets for online interpolated RPB and DPB are both acceptable.
Conclusions and Future Works
In this paper, we first proposed a universal visual backbone, dubbed CrossFormer. It utilizes cross-scale features through our proposed CEL and LSDA. The experimental results show that utilizing cross-scale features explicitly can significantly improve the vision transformers’ performance on image classification and other downstream tasks. Besides, a more flexible relation position representation, DPB, is also proposed to make the CrossFormer apply to dynamic group size. Based on CrossFormer, we further analyzed the self-attention module’s and MLPs’ outputs in the CrossFormer and proposed progressive group size (PGS) and amplitude cooling layer (ACL). Based on PGS and ACL, CrossFormer++ achieves better performance than CrossFormer and outperforms all the existing state-of-the-art vision transformers on representative visual tasks. Besides, extensive experiments show that CEL, PGS, and ACL are all universal and can consistently bring performance gains when plugged into other vision transformers.
Despite these contributions, the proposed modules still have some limitations. Concretely, though we demonstrated the importance of a progressive appropriate group size, an empirically designed group size is used finally. An adaptive and automated group size policy is still expected. Then, to prevent the gradient vanishing issue, the ACL layer cannot be used in large quantities in vision transformers. So, we will also explore how to cool down the vision transformers’ amplitude without cutting the residual connection.
Acknowledgments
This work was supported in part by The National Nature Science Foundation of China (Grant Nos: 62036009, 62273302, 62273303, 62303406, 61936006, U1909203), in part by Ningbo Key R&D Program (No.2023Z231, 2023Z229), in part by the Key R&D Program of Zhejiang Province, China (2023C01135), in part by Yongjiang Talent Introduction Programme (Grant No: 2022A-240-G, 2023A-194-G), in part by the Key Research and Development Program of Zhejiang Province (No. 2021C01012).