Inception Transformer
Chenyang Si, Weihao Yu, Pan Zhou, Yichen Zhou, Xinchao Wang, Shuicheng Yan
Introduction
Transformer has taken the natural language processing (NLP) domain by storm, achieving surprisingly high performance in many NLP tasks, e.g., machine translation and question-answering . This is largely attributed to its strong capability of modeling long-range dependencies in the data with self-attention mechanism. Its success has led researchers to investigate its adaptation to the computer vision field, and Vision Transformer (ViT) is a pioneer. This architecture is directly inherited from NLP , but applied to image classification with raw image patches as input. Later, many ViT variants have been developed to boost performance or scale to a wider range of vision tasks, e.g., object detection and segmentation .
ViT and its variants are highly capable of capturing low-frequencies in the visual data , mainly including global shapes and structures of a scene or object, but are not very powerful for learning high-frequencies, mainly including local edges and textures. This can be intuitively explained: self-attention, the main operation used in ViTs to exchange information among non-overlap patch tokens, is a global operation and much more capable of capturing global information (low frequencies) in the data than local information (high frequencies). As shown in Fig. 1(a) and 1(b), the Fourier spectrum and relative log amplitudes of the Fourier show that ViT tends to well capture low-frequency signals but few high-frequency signals. This observation also accords with the empirical results in , which shows ViT presents the characteristics of low-pass filters. This low-frequency preferability impairs the performance of ViTs, as 1) low-frequency information filling in all the layers may deteriorate high-frequency components, e.g., local textures, and weakens modeling capability of ViTs; 2) high-frequency information is also discriminative and can benefit many tasks, e.g., (fine-grained) classification. Actually, human visual system extracts visual elementary features at different frequencies : low frequency provides global information about a visual stimulus, and high frequency conveys local spatial changes in the image (e.g., local edges/textures). Hence, it is necessary to develop a new ViT architecture for capturing both high and low frequencies in the visual data.
CNNs are the most fundamental backbone for general vision tasks. Unlike ViTs, they cover more local information through local convolution within the receptive fields, thus effectively extracting high-frequency representations . Recent studies have integrated CNNs and ViTs considering their complementary advantages. Some methods stack convolution and attention layers in a serial manner to inject the local information into global context. Unfortunately, this serial manner only models one type of dependency, either global or local, in one layer, and discards the global information during locality modeling, or vice versa. Other works adopt parallel attention and convolution to learn global and local dependencies of the input at the same time. However, it is found in that part of the channels are for processing local information and the other for global modeling, meaning current parallel structures have information redundancy if processing all channels in each branch.
To address this issue, we propose a simple and efficient Inception Transformer (iFormer), as shown in Fig. 2, which grafts the merit of CNNs for capturing high-frequencies to ViTs. The key component in iFormer is an Inception token mixer as shown in Fig. 3. This Inception mixer aims to augment the perception capability of ViTs in the frequency spectrum by capturing both high and low frequencies in the data. To this end, the Inception mixer first splits the input feature along the channel dimension, and then feeds the split components into high-frequency mixer and low-frequency mixer respectively. Here the high-frequency mixer consists of a max-pooling operation and a parallel convolution operation, while the low-frequency mixer is implemented by a vanilla self-attention in ViTs. In this way, our iFormer can effectively capture particular frequency information on the corresponding channel, and thus learn more comprehensive features within a wide frequency range compared with vanilla ViTs, which can be clearly observed in Fig. 1(a) and 1(b).
Moreover, we find that lower layers often need more local information, while higher layers desire more global information, which also accords with the observations in . This is because, like in human visual system, the details in high frequency components help lower layers to capture visual elementary features and also to gradually gather local information for having a global understanding of the input. Inspired by this, we design a frequency ramp structure. In particular, from lower to higher layers, we gradually feed more channel dimensions to low-frequency mixer and fewer channel dimensions to high-frequency mixer. This structure can trade-off high-frequency and low-frequency components across all layers. Its effectiveness has been verified by experimental results in Sec. 4.
Experimental results show that iFormer surpasses state-of-the-art ViTs and CNNs on several vision tasks, including image classification, object detection and segmentation. For example, as shown in Fig. 1(c), with different model sizes, iFormer makes consistent improvements over popular frameworks on ImageNet-1K , e.g., DeiT , Swin and ConvNeXt . Meanwhile, iFormer outperforms recent frameworks on COCO detection and ADE20K segmentation.
Related work
Transformers are firstly proposed for machine translation tasks and then become popular in other tasks like natural language understanding and generation in NLP domain, as well as image classification , object detection and semantic segmentation in computer vision. The attention module in Transformers has an outstanding ability to capture global dependency, but it makes the models produce similar representations across layers . Moreover, self-attention mainly captures low-frequency information and tends to neglect high-frequency components related to the detailed information .
CNNs are the de-facto model for vision tasks due to their outstanding ability to model local dependency as well as extract high-frequency . With these advantages, CNNs are rapidly introduced into Transformers in a serial or parallel manner . For serial methods, convolutions are applied at different positions of the Transformer. CvT and PVT-v2 replace the hard patch embedding with a layer of overlapping convolution. LV-ViT , LeViT and ViTC further stack several layers of convolutions as the stem for models, which is found helpful in training and achieving better performance. Besides the stem, ViT-hybrid , CoAtNet , Hybrid-MS and UniFormer design early stages with convolution layers. However, the combination of convolution and attention in a serial order means each layer can only process either high or low frequency and neglects the other part. To enable each layer to process different frequencies, we adopt the parallel manner to combine convolution and attention in a token mixer.
Compared with serial methods, there are not many works combining attention and convolution in a parallel manner in literature. CoaT and ViTAE introduce convolution as a branch parallel to attention and utilize elementwise sum to merge the output of the two branches. However, Raghu et al. find that some channels tend to extract local dependency while others are for modeling global information , indicating redundancy for the current parallel mechanism to process all channels in different branches. In contrast, we split channels into branches of high and low frequencies. GLiT also adopt parallel manner but it directly concatenate the features from convolution and attention branches as the mixer output, lacking the fusion of features in different frequencies. Instead, we design a explicit fusion module to merge the outputs from low- and high-frequency branches.
Method
In MSA, the attention-based mixer exchanges information between all patch tokens so that it strongly focuses on aggregating the global dependency across all layers. However, excessive propagation of global information would strengthen the low-frequency representation. It can be seen from the visualization of Fourier spectrum in Fig. 1(a) that low-frequency information dominates the representations of ViT . This actually impairs the performance of ViTs, as it may deteriorate the high-frequency components, e.g., local textures, and weakens the modeling capability of ViTs . In the visual data, high-frequency information is also discriminative and can benefit many tasks . Hence, to address the issue, we propose a simple and efficient Inception Transformer, as shown in Fig. 2, with two key novelties, i.e., Inception mixer and frequency ramp structure.
2 Inception token mixer
We propose an Inception mixer to graft the powerful capability of CNNs for extracting high-frequency representation to Transformers. Its detailed architecture is depicted in Fig. 3. We use the name of “Inception" since the token mixer is highly inspired by the Inception module with multiple branches. Instead of directly feeding image tokens into the MSA mixer, the Inception mixer first splits the input feature along the channel dimension, and then respectively feeds the split components into high-frequency mixer and low-frequency mixer. Here the high-frequency mixer consists of a max-pooling operation and a parallel convolution operation, while the low-frequency mixer is implemented by a self-attention.
where and denote the outputs of high-frequency mixers.
Finally, the outputs of low- and high-frequency mixers are concatenated along the channel dimension:
The upsample operation in Eq. (7) selects the value of the nearest point for each position to be interpolated regardless of any other points, which results in excessive smoothness between adjacent tokens. We design a fusion module to elegantly overcome this issue, i.e., a depthwise convolution exchanging information between patches, while keeping a cross-channel linear layer that works per location like in previous Transformers. The final output can be expressed as
Like the vanilla Transformer, our iFormer is equipped with a feed-forward network (FFN), and differently it also incorporates the above Inception token mixer (ITM); LayerNorm (LN) is applied before ITM and FFN. Hence the Inception Transformer block is formally defined as
We use the vanilla multi-head self-attention to communicate information among all tokens for the low-frequency mixer. Despite the strong capability of the attention for learning global representation, the large resolution of feature maps would bring large computation cost in lower layers. We therefore simply utilize an average pooling layer to reduce the spatial scale of before the attention operation and an upsample layer to recover the original spatial dimension after the attention. This design largely reduces the computational overhead and makes the attention operation focus on embedding global information. This branch can be defined as
where is the output of low-frequency mixer. Note that the kernel size and stride for the pooling and upsample layers are set to 2 only at the first two stages.
3 Frequency ramp structure
In the general visual frameworks, bottom layers play more roles in capturing high-frequency details while top layers more in modeling low-frequency global information, i.e., the hierarchical representations of ResNet . Like humans, by capturing the details in high frequency components, lower layers can capture visual elementary features, and also gradually gather local information to achieve a global understanding of the input. We are inspired to design a frequency ramp structure which gradually splits more channel dimensions from lower to higher layers to low-frequency mixer and thus leave fewer channel dimensions to high-frequency mixer. Specifically, as shown in Fig. 2, our backbone has four stages with different channel and spatial dimensions. For each blocks, we define a channel ratio to better balance the high-frequency and low frequency components, i.e., and , where . In the proposed frequency ramp structure, gradually decreases from shallow to deep layers, while gradually increases. Hence, with the flexible frequency ramp structure, iFormer can effectively trade-off high- and low-frequency components across all layers. The configuration of different iFormer models will be described in the appendix.
Experiments
We evaluate our iFormer on several vision benchmark tasks, i.e., image classification, object detection and semantic segmentation, by comparing it with representative ViTs, CNNs and their hybrid variants. Ablation analysis is also conducted to show the contribution of each novelty in our method. More results will be reported in the appendix.
Setup. For image classification, we evaluate iFormer on the ImageNet dataset . We train the iFormer model with the standard procedure in . Specifically, we use AdamW optimizer with an initial learning rate via cosine decay , a momentum of 0.9, and a weight decay of 0.05. We set the training epoch number as 300 and the input size as 224 224. We adopt the same data augmentations and regularization methods in DeiT for fair comparison.
We also use LayerScale to train deep models. Like previous studies , we further fine tune iFormer on the input size of , with the weight decay of , learning rate of , batch size of 512. For fairness, we adopt Timm to implement and train iFormer.
Results. Table 1 summarizes the image classification accuracy of all compared methods on ImageNet. For the small model size (20M), our iFormer surpasses both the SoTA ViTs and hybrid ViTs, although some ViTs, e.g., Swin , Focal and CSwin , actually already introduce convolution-like inductive bias into their architectures, and hybrid ViTs directly integrate convolution into ViTs. Specifically, our iFormer-S respectively gains and top-1 accuracy advantage over SoTA ViTs ( i.e., CSwin-T) and hybrid ViTs ( i.e., UniFormer-S), while enjoying the same or smaller model size.
For the medium model size (50M), iFormer-B achieves 84.6% top-1 accuracy, and improves over the SoTA ViTs and hybrid ViTs with similar model sizes by significant margins 1.0% and 0.7% respectively. For CNNs, similar to comparison results on medium model size, our iFormer-B outperforms ConvNeXt-S by . As for the large mode (100M), one can observe similar results on small and medium model sizes.
Table 2 reports the fine-tuning accuracy on the larger resolution, i.e., . One can observe that iFormer consistently outperforms the counterparts by a significant margin across different computation settings. These results clearly demonstrate the advantages of iFormer on image classifications.
2 Results on object detection and instance segmentation
Setup. We evaluate iFormer on the COCO object detection and instance segmentation tasks , where the models are trained on 118K images and evaluated on validation set with 5K images. Here, we use iFormer as the backbone in Mask R-CNN . In the training phase, we use iFormer pretrained on ImageNet to initialize the detector, and adopt AdamW to train with an initial learning rate of , a batch size of 16, and 1 training schedule with 12 epochs. For training, the input images are resized to be 800 pixels on the shorter side an no more than 1,333 pixels on the longer side. For the test image, its shorter side is fixed to 800 pixels. All experiments are implemented on mmdetection codebase.
Results. Table 3 reports the box mAP (APb) and mask mAP (APm) of the compared models. Under similar computation configurations, iFormers outperforms all previous backbones. Specifically, compared with popular ResNet backbones, our iFormer-S brings points of APb and points APm improvements over ResNet50. Compared with various Transformer backbones, our iFormers still maintain the performance superiority over their results. For example, our iFormer-B surpasses UniFormer-B , Swin-S by points of APb and points of APb respectively.
3 Results on semantic segmentation
Setup. We further evaluate the generality of iFormer through a challenging scene parsing benchmark on semantic segmentation, i.e., ADE20K . The dataset contains 20K training images and 2K validation images. We adopt iFormer pretrained on ImageNet as the backbone of the Semantic FPN framework. Following PVT and UniFormer , we use AdamW with an initial learning rate of with cosine learning rate schedule to train 80k iterations. All experiments are implemented on mmsegmentation codebase.
Results. In Table 4, we report the mIoU results of different backbones. On the Semantic FPN framework, our iFormer consistently outperforms previous backbones on this task, including CNNs and (hybrid) ViTs. For instance, iFormer-S achieves mIoU, surpassing UniFormer-S by mIoU, while using less computation complexity. Moreover, compared with UniFormer-B , our iFormer-S still achieves mIoU improvement with only parameters and nearly FLOPs.
4 Ablation study and visualization
In this section, we conduct experiments to better understand iFormer. All the models are trained for 100 epochs on ImageNet, with the same training setting as described in Sec. 4.1.
Inception token mixer. The Inception mixer is proposed to augment the perception capability of ViTs in the frequency spectrum. To evaluate the effects of the components in the Inception mixer, we increasingly remove each branch from the full model and then report the results in Table 5, where ✓and ✗denote whether or not the corresponding branch is enabled. Observably, combining attention with convolution and max-pooling can achieve better accuracy than the attention-only mixer, while using less computation complexity, which implies the effectiveness of Inception Token Mixer. To further explore this scheme, Fig. 4 visualizes the Fourier spectrum of the Attention, MaxPool and DwConv branches in Inception mixer. We can see the attention mixer has higher concentrations on low frequencies; with the high-frequency mixer, i.e., convolution and max-pooling, the model is encouraged to learn high frequency information. Overall, these results prove the effectiveness of the Inception mixer for expanding the perception capability of the Transformer in the frequency spectrum.
Frequency ramp structure. Previous investigations show requirement of more local information at lower layers of the Transformer and more global information at higher layers. We accordingly assume that a frequency ramp structure, i.e., decreasing dimensions at high-frequency components and increasing dimensions at low-frequency components from lower to higher layers, has a better trade-off between high-frequency and low-frequency components across all layers. In order to justify this hypothesis, we investigate the effects of the channel ratio ( and ) in Table 5. It can be clearly seen that the model with outperforms the other two models, which is consistent with the previous investigations. Hence, this indicates the rationality of the frequency ramp structure and its potential for leaning discriminating vision representations.
Visualization. We visualize the Grad-CAM activation maps of iFormer-S as well as Swin-T models trained on ImageNet-1K in Fig. 5. It can be seen that compared with Swin, iFormer can more accurately and completely locate the objects. For example, in the hummingbird image, iFormer skips the branch and accurately attends to the whole bird including the tail.
Conclusion
In this paper, we present an Inception Transformer (iFormer), a novel and general Transformer backbone. iFormer adopts a channel splitting mechanism to simply and efficiently couple convolution/max-pooling and self-attention, giving more concentrations on high frequencies and expanding the perception capability of the Transformer in the frequency spectrum. Based on the flexible Inception token mixer, we further design a frequency ramp structure, enabling effective trade-off between high-frequency and low-frequency components across all layers. Extensive experiments show that iFormer outperforms representative vision Transformers on image classification, object detection and semantic segmentation, demonstrating the great potential of our iFormer to serve as a general-purpose backbone for computer vision. We hope this study will provide valuable insights for the community to design efficient and effective Transformer architectures.
One obvious limitation of the proposed iFormer is that it requires manually defined channel ratio in the frequency ramp structure i.e., and for each iFormer block, which needs rich experience to define better on different tasks. it is not trained on large scale datasets, e.g., ImageNet-21K , due to computational constraint, which will be explored in further. Also, iFormer requires manually defined channel ratio in the frequency ramp structure i.e., and for each iFormer block, which needs rich experience to define better on different tasks. A straightforward solution would be to use neural architecture search.
Acknowledgement
Weihao Yu would like to thank TRC program and GCP research credits for the support of partial computational resources.
References
Appendix A Appendix
Potential Impacts. The study introduces a general vision Transformer, i.e., iFormer, which can be used on different vision tasks, e.g., image classification, object detection and semantic segmentation. iFormer has no direct negative societal impact. However, we will realize that iFormer as a general-purpose backbone can be used for harmful applications such as illegal face recognition.
We further evaluate the generality of iFormer on semantic segmentation with the Upernet framework. Following the training settings in Swin , the model is trained for 160K iterations with a batch size of 16. For training, we use AdamW optimizer with an initial learning rate . All experiments are implemented on mmsegmentation codebase.
Table 6 shows the mIoU and MS mIoU results of different backbones based on the UperNet framework. From these results, it can be seen that our iFormer achieves 48.4 mIoU and 48.8 MS mIoU, consistently surpassing previous backbones on this task. For instance, our iFormer-S outperforms the Swin-T by 3.9 mIoU while using fewer parameters. Compered with UniFormer-S , iFormer-S still achieves 0.8 mIoU improvement.
A.2 Visualization
Frequency ramp structure plays an important role in iFormer, which is designed to learn hierarchical representations, i.e., more high-frequency signals at lower layers of the Transformer and more low-frequency signals at higher layers. In order to justify this hypothesis, we visualize the Fourier spectrum of feature maps for different iFormer layers in Fig 6. We can see that iFormer captures more high-frequency components at 6- layer and more low-frequency information at 18- layer. Moreover, the high-frequency information gradually decreases from 6- layer to 18- layer, and low-frequency information does the opposite. Hence, these results indicate our iFormer can effectively trade-off high- and low-frequency components across all layers.
A.2.2 CAM
We futher show more examples of the Grad-CAM activation maps of iFormer-S as well as Swin-T models trained on ImageNet-1K in Fig. 7. These examples indicate that compared with Swin, iFormer can more accurately and completely attend to the key objects. Taking the hog picture as example, iFormer locates the hog accurately but Swin also locates irrelevant part.
A.3 Configurations of iFormers
In this work, three variants of iFormer are used for a fair comparison under computation configurations, i.e., iFormer-S, iFormer-B and iFormer-L. Table 7 shows their detailed configurations. Following Swin , iFormer adopts 4-stage architecture with , , , input sizes, where and are the width and height of the input image. In each iFormer block, and are used to balance the high-frequency and low frequency components. As shown in Table 7, gradually decreases from shallow to deep layers, while gradually increases. iFormer block uses depthwise convolution and max-pooling as high-frequency mixers. We set the kernel sizes of depthwise convolution and max-pooling to .