Scale-Aware Modulation Meet Transformer

Weifeng Lin, Ziheng Wu, Jiayu Chen, Jun Huang, Lianwen Jin

Introduction

Since the groundbreaking work on Vision Transformers (ViT) , Transformers have gained significant attention from both industry and academia, achieving remarkable success in various computer vision tasks, such as image classification , object detection , and semantic segmentation . Unlike convolutional networks, which only allow for interactions within a local region using a shared kernel, ViT divides the input image into a sequence of patches and updates token features via self-attention (SA), enabling global interactions. However, self-attention still faces challenges in downstream tasks due to the quadratic complexity in the number of visual tokens, particularly for high-resolution inputs.

To address these challenges, several efficient spatial attention techniques have been proposed. For example, Swin Transformer employs window attention to limit the number of tokens and establish cross-window connections via shifting. PVT and Focal reduce the cost of self-attention by combining token merging with spatial reduction. Shunted effectively models objects at multiple scales simultaneously while performing spatial reduction. Other techniques such as dynamic token selection have also proven to be effective improvements.

Rather than directly improving self-attention, several works have investigated hybrid CNN-Transformer architectures that combine efficient convolutional blocks with powerful Transformer blocks. We observed that most hybrid networks replace shallow Transformer blocks with convolution blocks to reduce the high computational cost of self-attention in the early stages. However, these simplistic stacking strategies hinder them from achieving a better balance between accuracy and latency. Therefore, one of the objectives of this paper is to present a new perspective on the integration of Transformer and convolution blocks.

Based on the research conducted in , which performed a quantitative analysis of different depths of self-attention blocks and discovered that shallow blocks tend to capture short-range dependencies while deeper ones capture long-range dependencies, we propose that substituting convolution blocks for Transformer blocks in shallow networks offers a promising strategy for two primary reasons: (1)(1) self-attention induces significant computational costs in shallow networks due to high-resolution input, and (2)(2) convolution blocks, which inherently possess a capacity for local modeling, are more proficient at capturing short-range dependencies than SA blocks in shallow networks. However, we observed that simply applying the convolution directly to the feature map does not lead to the desired performance. Taking inspiration from recent convolutional modulation networks , we discovered that convolutional modulation can aggregate surrounding contexts and adaptively self-modulate, giving it a stronger modeling capability than using convolution blocks alone. Therefore, we proposed a novel convolutional modulation, termed Scale-Aware Modulation (SAM), which incorporates two new modules: Multi-Head Mixed Convolution (MHMC) and Scale-Aware Aggregation (SAA). The MHMC module is designed to enhance the receptive field and capture multi-scale features simultaneously. The SAA module is designed to effectively aggregate features across different heads while maintaining a lightweight architecture. Despite these improvements, we find that SAM falls short of the self-attention mechanism in capturing long-range dependencies. To address this, we propose a new hybrid Modulation-Transformer architecture called the Evolutionary Hybrid Network (EHN). Specifically, we incorporate SAM blocks in the top two stages and Transformer blocks in the last two stages, while introducing a new stacking strategy in the penultimate stage. This architecture not only simulates changes in long-range dependencies from shallow to deep layers but also enables each block in each stage to better match its computational characteristics, leading to improved performance on various downstream tasks. Collectively, we refer to our proposed architecture as Scale-Aware Modulation Transformer (SMT).

As shown in Fig. 1, our SMT significantly outperforms other SOTA vision Transformers and convolutional networks on ImageNet-1K . It is worth noting that our SMT achieves top-1 accuracy of 82.2% and 84.3% with the tiny and base model sizes, respectively. Moreover, our SMT consistently outperforms other SOTA models on COCO and ADE20K for object detection, instance segmentation, and semantic segmentation tasks.

Overall, the contributions of this paper are as follows.

We introduce the Scale-Aware Modulation (SAM) which incorporates a potent Multi-Head Mixed Convolution (MHMC) and an innovative, lightweight Scale-Aware Aggregation (SAA). The SAM facilitates the integration of multi-scale contexts and enables adaptive modulation of tokens to achieve more precise predictions.

We propose a new evolutionary hybrid network that effectively models the transition from capturing local to global dependencies as the network increases in depth, leading to improved performance and high efficiency.

We evaluated our proposed Scale-Aware Modulation Transformer (SMT) on several widely used benchmarks, including classification, object detection, and segmentation. The experimental results indicated that SMT consistently outperformed the SOTA Vision Transformers while requiring fewer parameters and incurring lower computational costs.

Related Work

The Transformer was initially developed for natural language processing tasks and has since been adapted for computer vision tasks through the introduction of the Vision Transformer (ViT) . Further improvements to ViT have been achieved through knowledge distillation or more intricate data augmentation, as demonstrated by DeiT . However, Transformers do not consider the quadratic complexity of high-resolution images or the 2D structure of images, which are challenges in vision tasks. To address these issues and improve the performance of vision Transformers, various methods have been proposed, including multi-scale architectures , lightweight convolution layers , and local self-attention mechanisms .

2 Convolutional Neural Networks

Convolutional neural networks (CNNs) have been the main force behind the revival of deep neural networks in computer vision. Since the introduction of AlexNet , VGGNet , and ResNet , CNNs have rapidly become the standard framework for computer vision tasks. The design principles of CNNs have been advanced by subsequent models such as Inception , ResNeXt , Res2Net and MixNet , which promote the use of building blocks with multiple parallel convolutional paths. Other works such as MobileNet and ShuffleNet have focused on the efficiency of CNNs. To further improve the performance of CNNs, attention-based models such as SE-Net , Non-local Networks , and CBAM have been proposed to enhance the modeling of channel or spatial attention. EfficientNets and MobileNetV3 have employed neural architecture search (NAS) to develop efficient network architectures. ConvNeXt adopts the hierarchical design of Vision Transformers to enhance CNN performance while retaining the simplicity and effectiveness of CNNs. Recently, several studies have utilized convolutional modulation as a replacement for self-attention, resulting in improved performance. Specifically, FocalNet utilizes a stack of depth-wise convolutional layers to encode features across short to long ranges and then injects the modulator into the tokens using an element-wise affine transformation. Conv2Former achieves good recognition performance using a simple 11×1111\times 11 depth-wise convolution. In contrast, our scale-aware modulation also employs depth-wise convolution as a basic operation but introduces multi-head mixed convolution and scale-aware aggregation.

3 Hybrid CNN-Transformer Networks

A popular topic in visual recognition is the development of hybrid CNN-Transformer architectures. Recently, several studies have demonstrated the effectiveness of combining Transformers and convolutions to leverage the strengths of both architectures. CvT first introduced depth-wise and point-wise convolutions before self-attention. CMT proposed a hybrid network that utilizes Transformers to capture long-range dependencies and CNNs to model local features. MobileViT , EdgeNeXt , MobileFormer , and EfficientFormer reintroduced convolutions to Transformers for efficient network design and demonstrated exceptional performance in image classification and downstream applications. However, the current hybrid networks lack the ability to model range dependency transitions, making it challenging to improve their performance. In this paper, we propose an evolutionary hybrid network that addresses this limitation and showcases its importance.

Method

The overall architecture of our proposed Scale-Aware Modulation Transformer (SMT) is illustrated in Fig. 2. The network comprises four stages, each with downsampling rates of {4,8,16,32}\{4,8,16,32\}. Instead of constructing an attention-free network, we first adopt our proposed Scale-Aware Modulation (SAM) in the top two stages, followed by a penultimate stage where we sequentially stack one SAM block and one Multi-Head Self-Attention (MSA) block to model the transition from capturing local to global dependencies. For the last stage, we solely use MSA blocks to capture long-range dependencies effectively. For the Feed-Forward Network (FFN) in each block, we adopt the detail-specific feedforward layers as used in Shunted .

2 Scale-Aware Modulation

We propose the Multi-Head Mixed Convolution (MHMC), which introduces multiple convolutions with different kernel sizes, enabling it to capture various spatial features across multiple scales. Furthermore, MHMC can expand the receptive field using a large convolutional kernel, enhancing its ability to model long-range dependencies. As depicted in Fig. 3(b), MHMC partitions input channels into N heads and applies distinct depth-wise separable convolutions to each head, which reduces the parameter size and computational cost. To simplify our design process, we initialize the kernel size with 3×\times3 and gradually increase it by 2 per head. This approach enables us to regulate the range of receptive fields and multi-granularity information by merely adjusting the number of heads. Our proposed MHMC can be formulated as follows:

where x=[x1,x2,...,xn]x=[x_{1},x_{2},...,x_{n}] means to split up the input feature xx into multiple heads in the channel dimension and ki∈{3,5,…,K}k_{i}\in\{3,5,\dots,K\} denotes the kernel size increases monotonically by 2 per head.

As shown in Fig. 4(a), each distinct convolution feature map learns to focus on different granularity features in an adaptive manner, as expected. Notably, when we compare the single-head and multi-head by visualizing modulation maps in Fig. 4(b), we find that the visualization under multi-head depicts the foreground and target objects accurately in stage 1, while filtering out background information effectively. Moreover, it can still present the overall shape of the target object as the network becomes deeper, while the information related to the details is lost under the single-head convolution. This indicates that MHMC has the ability to capture local details better than a single head at the shallow stage, while maintaining detailed and semantic information about the target object as the network becomes deeper.

Scale-Aware Aggregation

Fig. 5 shows that our SAA module explicitly strengthens the semantically relevant low-frequency signals and precisely focuses on the most important parts of the target object. For instance, in stage 2, the eyes, head and body are clearly highlighted as essential features of the target object, resulting in significant improvements in classification performance. Compared to the convolution maps before aggregation, our SAA module demonstrates a better ability to capture and represent essential features for visual recognition tasks. (More visualizations can be found in Appendix E).

Scale-Aware Modulation

where ⊙\odot is the element-wise multiplication, WvW_{v} and WsW_{s} are weight martices of linear layers. Since the modulator is calculated via Eq. 3, it changes dynamically with different inputs, thereby achieving adaptively self-modulation. Moreover, unlike self-attention, which computes an N×NN\times N attention map, the modulator retains the channel dimension. This feature allows for spatial- and channel-specific modulation of the value after element-wise multiplication, while also being memory-efficient, particularly when processing high-resolution images.

3 Scale-Aware Modulation Transformer

In this section, we propose to reallocate the appropriate computational modules according to the variation pattern in the network’s capture range dependencies to achieve better computational performance. We propose using MSA blocks only from the penultimate stage to reduce the computational burden. Furthermore, to effectively simulate the transition pattern, we put forth two hybrid stacking strategies for the penultimate stage: (i)(i) sequentially stacking one SAM block and one MSA block, which can be formulated as (SAM×1+MSA×1)×N2(SAM\times 1+MSA\times 1)\times\frac{N}{2}, depicted in Fig. 6(i); (ii)(ii) using SAM blocks for the first half of the stage and MSA blocks for the second half, which can be formulated as (SAM×N2+MSA×N2)(SAM\times\frac{N}{2}+MSA\times\frac{N}{2}), depicted in Fig. 6(ii).

To assess the efficacy of these hybrid stacking strategies, we evaluated their top-1 accuracy on the ImageNet-1K, as shown in Table 9. Moreover, as depicted in Fig. 7, we calculate the relative receptive field of the MSA blocks in the penultimate stage, followed by the approach presented in . It is noteworthy that there is a slight downward trend in the onset of the relative receptive field in the early layers. This decline can be attributed to the impact of the SAM on the early MSA blocks, which emphasize neighboring tokens. We refer to this phenomenon as the adaptation period. As the network becomes deeper, we can see a smooth and steady upward trend in the receptive field, indicating that our proposed evolutionary hybrid network effectively simulates the transition from local to global dependency capture.

Experiments

To ensure a fair comparison under similar parameters and computation costs, we construct a range of SMT variants. We validate our SMTs on ImageNet-1K image classification, MS COCO object detection, and ADE20K semantic segmentation. Besides, extensive ablation studies provide a close look at different components of the SMT. (The detailed model settings are presented in Appendix A)

We conduct an evaluation of our proposed model and compare it with various networks on ImageNet-1K classification . To ensure a fair comparison, we follow the same training recipes as previous works . Specifically, we train the models for 300 epochs with an image size of 224×224224\times 224 and report the top-1 validation accuracy. The batch size used is 1024, and we employ the AdamW optimizer with a weight decay of 0.05 and a learning rate of 1×10−31\times 10^{-3}. In addition, we investigate the effectiveness of SMTs when pretrained on ImageNet-22K.(Further details regarding the training process can be found in Appendix B)

Results

Tab. 1 presents a comparison of our proposed SMT with various models, and the results demonstrate that our models outperform various architectures with fewer parameters and lower computation costs. Specifically, concerning the tiny-sized model, SMT achieves an impressive top-1 accuracy of 82.2%, surpassing PVTv2-b1 and Shunted-T by significant margins of 3.5% and 2.4%, respectively. Furthermore, when compared to small-sized and base-sized models, SMT maintains its leading position. Notably, SMT-B achieves a top-1 accuracy of 84.3% with only 32M parameters and 7.7GFLOPs of computation, outperforming many larger models such as Swin-B , ConvNeXt-B , and FocalNet-B , which have over 70M parameters and 15GFLOPs of computation. Additionally, to evaluate the scalability of the SMT, we have also created smaller and larger models, and the experimental results are presented in the Appendix C.

We also report the ImageNet-22K pre-training results here in Tab. 2. When compared to the previously best results, our models achieve significantly better accuracy with a reduced number of parameters and FLOPs. SMT-L attains an 88.1% top-1 accuracy, surpassing InternImage-XL by 0.1% while utilizing significantly fewer parameters (80.5M vs. 335M) and exhibiting lower FLOPs (54.6G vs. 163G). This highly encouraging outcome underscores the impressive scalability capabilities of SMT.

2 Object Detection and Instance Segmentation

We make comparisons on object detection with COCO 2017 . We use SMT-S/B pretrained on ImageNet-1K as the foundation for three well-known object detectors: Mask R-CNN , Cascade Mask R-CNN , and RetinaNet . To demonstrate a consistent comparison, two training schedules (1×1\times schedule with 12 epochs and 3×3\times schedule with 36 epochs) are adopted in Mask R-CNN. In 3×3\times schedule, we use a multi-scale training strategy by randomly resizing the shorter side of an image to between . We take AdamW optimizer with a weight decay of 0.05 and an initial learning rate of 2×10−42\times 10^{-4}. Both models are trained with batch size 16. To further showcase the versatility of SMT, we conducted a performance evaluation of SMT with three other prominent object detection frameworks, namely Sparse RCNN , ATSS , and DINO . We initialize the backbone with weights pre-trained on ImageNet-1K and fine-tune the model using a 3×\times schedule for Sparse RCNN and ATSS.

Results

Tab. 3 presents the superior performance of SMT over other networks with Mask R-CNN under various model sizes. Specifically, SMT demonstrates a significant improvement in box mAP of 5.6 and 4.2 over Swin Transformer in 1×\times schedule under small and base model sizes, respectively. Notably, with 3×\times schedule and multi-scale training, SMT still consistently outperforms various backbones. For instance segmentation, the results also demonstrate that our SMT achieves higher mask mAP in comparison to previous SOTA networks. In particular, for small and base models in the 1×\times schedule, we achieve 1.5 and 0.9 points higher than FocalNet, respectively. Furthermore, to assess the generality of SMT, we trained two additional detection models, Cascade Mask R-CNN and RetinaNet , using SMT-S as the backbone. The results, presented in Tab. 4, show clear improvements over various backbones in both box and mask mAPs. The resulting box mAPs for Sparse R-CNN, ATSS and DINO are presented in Tab. 5, which indicate that SMT outperforms other networks consistently across all detection frameworks, highlighting its exceptional performance in downstream tasks.

3 Semantic Segmentation on ADE20K

We evaluate the SMT for semantic segmentation using the ADE20K dataset. To conduct the evaluation, we use UperNet as the segmentation method and closely followed the training settings proposed by . Specifically, we train UperNet for 160k iterations with an input resolution of 512×512512\times 512. We employ the AdamW optimizer with a weight decay of 0.01, and set the learning rate to 6×10−56\times 10^{-5}.

Results

The results are presented in Tab. 6, which shows that our SMT outperforms Swin, FocalNet, and Shunted Transformer significantly under all settings. Specifically, SMT-B achieves 1.5 and 0.9 mIoU gains compared to Swin-B and a 0.6 and 0.1 mIoU improvement over Focal-B at single- and multi-scale, respectively, while consuming significantly fewer FLOPs and reducing the model size by more than 50%. Even for the SMT with a small model size, it achieves comparable accuracy with the previous SOTA models which have a larger model size.

4 Ablation Study

Table 7 shows the impact of the number of convolution heads in the Multi-Head Mixed Convolution (MHMC) on our model’s performance. The experimental results indicate that while increasing the number of diverse convolutional kernels is advantageous for modeling multi-scale features and expanding the receptive field, adding more heads introduces larger convolutions that may negatively affect network inference speed and reduce throughput. Notably, we observed that the top-1 accuracy on ImageNet-1K peaks when the number of heads is 4, and increasing the number of heads does not improve the model’s performance. This findings suggest that introducing excessive distinct convolutions or using a single convolution is not suitable for our SMT, emphasizing the importance of choosing the appropriate number of convolution heads to model a specific degree of multi-scale spatial features.

Different aggregation strategies

After applying the MHMC, we introduce an aggregation module to achieve information fusion. Table 8 presents a comparison of different aggregation strategies, including a single linear layer, two linear layers, and an Invert BottleNeck (IBN) . Our proposed scale-aware aggregation (SAA) consistently outperforms the other fusion modules, demonstrating the effectiveness of SAA in modeling multi-scale features with fewer parameters and lower computational costs. Notably, as the size of the model increases, our SAA can exhibit more substantial benefits while utilizing a small number of parameters and low computational resources.

Different hybrid stacking strategies

In Sec. 3.3, we propose two hybrid stacking strategies to enhance the modeling of the transition from local to global dependencies. The results shown in Table 9 indicate that the first strategy which sequentially stacks one scale-aware modulation block and one multi-head self-attention block is better, achieving a performance gain of 0.3% compared to the other strategy. Furthermore, the strategy stacking all MSA blocks achieves comparable performance as well, which means retaining the MSA block in the last two stages is crucial.

Component Analysis

In this section, we investigate the individual contributions of each component by conducting an ablation study on SMT. Initially, we employ a single-head convolution module and no aggregation module to construct the modulation. Based on this, we build an attention-free network, which can achieve 80% top-1 accuracy on the ImageNet-1K dataset. The effects of all the proposed methods on the model’s performance are given in Tab. 10, which can be summarized as followings.

Multi-Head Mixed Convolution (MHMC) To enhance the model’s ability to capture multi-scale spatial features and expand its receptive field, we replaced the single-head convolution with our proposed MHMC. This module proves to be effective for modulation, resulting in a 0.8% gain in accuracy.

Scale-Aware Aggregation (SAA) We replace the single linear layer with our proposed scale-aware aggregation. The SAA enables effective aggregation of the multi-scale features captured by MHMC. Building on the previous modification, the replacement leads to a 1.6% increase in performance.

Evolutionary Hybrid Network (EHN) We incorporate the self-attention module in the last two stages of our model, while also implementing our proposed hybrid stacking strategy in the penultimate stage, which improves the modeling of the transition from local to global dependencies as the network becomes deeper, resulting in a significant gain of 2.2% in performance based on the aforementioned modifications.

Conclusion

In this paper, we introduce a new hybrid ConvNet and vision Transformer backbone, namely Scale-Aware Modulation Transformer (SMT), which can effectively simulate the transition from local to global dependencies as the network becomes deeper, resulting in superior performance. To satisfy the requirement of foundation models, we propose a new Scale-Aware Modulation that includes a potent multi-head mixed convolution module and a lightweight scale-aware aggregation module. Extensive experiments demonstrate the efficacy of SMT as a backbone for various downstream tasks, achieving comparable or better performance than well-designed ConvNets and vision Transformers, with fewer parameters and FLOPs. We anticipate that the exceptional performance of SMT on diverse vision problems will encourage its adoption as a promising new generic backbone for efficient visual modeling.

Acknowledgement

This research is supported in part by NSFC (Grant No.: 61936003), Alibaba DAMO Innovative Research Foundation (20210925), Zhuhai Industry Core, Key Technology Research Project (no. 2220004002350) and National Key Research and Development Program of China (2022YFC3301703). We thank the support from the Alibaba-South China University of Technology Joint Graduate Education Program.

References

Appendix A Detailed Architecture Specifications

Tab. 13 provides a detailed overview of the architecture specifications for all models, with an assumed input image size of 224×224224\times 224. The stem of the model is denoted as ”conv n×nn\times n, 64-d, BN; conv 2×22\times 2, 64-d, LN”, representing two convolution layers with a stride of 2 to obtain a more informative token sequence with a length of H4×W4\frac{H}{4}\times\frac{W}{4}. Here, ”BN” and ”LN” indicate Batch Normalization and Layer Normalization , respectively, while ”64-d” denotes the convolution layer with an output dimension of 64. The multi-head mixed convolution module with 4 heads (conv 3×33\times 3, conv 5×55\times 5, conv 7×77\times 7, conv 9×99\times 9) is denoted as ”sam. head. 4”, while ”msa. head. 8” represents the multi-head self-attention module with 8 heads. Additionally, ”sam. ep_r. 2” indicates a Scale-Aware Aggregation module with twice as much expanding ratio.

Appendix B Detailed Experimental Settings

We trained all models on the ImageNet-1K dataset for 300 epochs, using an image size of 224×224224\times 224 . Following Swin , we utilized a standardized set of data augmentations , including Random Augmentation, Mixup , CutMix , and Random Erasing . To regularize our models, we applied Label Smoothing and DropPath techniques. The initial learning rate for all models was set to 2×10−32\times 10^{-3} after 5 warm-up epochs, beginning with a rate of 1×10−61\times 10^{-6}. To optimize our models, we employed the AdamW algorithm and a cosine learning rate scheduler . The weight decay was set to 0.05 and the gradient clipping norm to 5.0. For our mini, tiny, small, base, and large models, we used stochastic depth drop rates of 0.1, 0.1, 0.2, 0.3, and 0.5, respectively. For more details, please refer to the Tab. 11 provided.

B.2 Image classification pretrained on ImageNet-22K

We trained the SMT-L model for 90 epochs using a batch size of 4096 and an input resolution of 224×224. The initial learning rate was set to 1×10−31\times 10^{-3} after a warm-up period of 5 epochs. The stochastic depth drop rates were set to 0.1. Following pretraining, we performed fine-tuning on the ImageNet-1K dataset for 30 epochs. The initial learning rate was set to 2×10−52\times 10^{-5}, and we utilized a cosine learning rate scheduler and AdamW optimizer. The stochastic depth drop rate remained at 0.1 during fine-tuning, while both CutMix and Mixup augmentation techniques were disabled.

B.3 Object Detection and Instance Segmentation

In transferring SMT to object detection and instance segmentation on COCO , we have considered six common frameworks: Mask R-CNN , Cascade Mask RCNN , RetinaNet , Sparse R-CNN , ATSS , and DINO . For DINO, the model is fine-tuned for 12 epochs, utilizing 4 scale features. For optimization, we adopt the AdamW optimizer with an initial learning rate of 0.0002 and a batch size of 16. When training models of different sizes, we adjust the training settings according to the settings used in image classification. The detailed hyper-parameters used in training models are presented in Tab. 12.

B.4 Semantic Segmentation

For ADE20K, we utilized the AdamW optimizer with an initial learning rate of 0.00006, a weight decay of 0.01, and a batch size of 16 for all models trained for 160K iterations. In terms of testing, we reported the results using both single-scale (SS) and multi-scale (MS) testing in the main comparisons. For multi-scale testing, we experimented with resolutions ranging from 0.5 to 1.75 times that of the training resolution. To set the path drop rates in different models, we used the same hyper-parameters as those used for object detection and instance segmentation.

Appendix C More Experiments

This section demonstrates how we scaled our SMT to create both smaller (SMT-M) and larger (SMT-L) models. Their detailed architecture settings are provided in Tab. 13, along with previous variants. We then evaluated their performance on the ImageNet-1K dataset.

As shown in Tab. 14, SMT-M achieves competitive results with a top-1 accuracy of 78.4%, despite having only 6.5M parameters and 1.3 GFLOPs of computation. On the other side, SMT-L shows an example to scale our SMT to larger models, which outperforms other state-of-the-art networks with similar parameters and computation costs, achieving a top-1 accuracy of 84.6%. These results confirm the strong scalability of the SMT architecture, which can be applied to create models of varying sizes, demonstrating its immense potential.

Appendix D Additional Network Analysis

In Fig. 8, we present the learned scale-aware modulation (SAM) value maps in two variants of SMT-T: evolutionary SMT, which employs an evolutionary hybrid stacking strategy, and general SMT, which only employs SAM in the penultimate stage. In evolutionary SMT-T, comprising a total of 8 layers in the penultimate stage, we select the layers ()() containing SAM block and compare them with the corresponding layers in general SMT. Through visualization, we can observe some noteworthy patterns. In general SMT, the model primarily concentrates on local details in the shallow layers and on semantic information in the deeper layers. However, in evolutionary SMT, the focus region does not significantly shift as the network depth increases. Furthermore, it captures local details more effectively than general SMT in the shallow layers, while preserving detailed and semantic information about the target object at deeper layers. These results indicate that our evolutionary hybrid stacking strategy facilitates SAM blocks in capturing multi-granularity features while allowing multi-head self-attention (MSA) blocks to concentrate on capturing global semantic information. Accordingly, each block within each layer is more aptly tailored to its computational characteristics, leading to enhanced performance in diverse visual tasks.

Appendix E Additional Visual Examples

We present supplementary visualization of modulation value maps within our SMT. Specifically, we randomly select validation images from the ImageNet-1K dataset and generate visual maps for modulation at different stages, as illustrated in Fig 9. The visualizations reveal that the scale-aware modulation is critical in strengthening semantically relevant low-frequency signals and accurately localizing the most discriminative regions within images. By exploiting this robust object localization capability, we can allocate more effort towards modulating these regions, resulting in more precise predictions. We firmly believe that both our multi-head mixed convolution module and scale-aware aggregation module have the potential to further enhance the modulation mechanism.