Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition

Qibin Hou, Cheng-Ze Lu, Ming-Ming Cheng, Jiashi Feng

Introduction

The prodigious progress in visual recognition in the 2010s was mostly dedicated to convolutional neural networks (ConvNets), typified by VGGNet , Inception series , and ResNet series , etc. These recognition models mostly aggregate responses with large receptive fields by stacking multiple building blocks and adopting the pyramid network architecture but neglect the importance of explicitly modeling the global contextual information. SENet series break through the traditional design of CNNs and introduce attention-based mechanisms into CNNs to capture long-range dependencies, attaining surprisingly good performance.

Since 2020, Vision Transformers (ViTs) further promoted the development of visual recognition models and show better results on the ImageNet classification and downstream tasks than the state-of-the-art ConvNets . This is because compared to convolutions that provide local connectivity, the self-attention mechanism in Transformers is able to model global pairwise dependencies, providing a more efficient way to encode spatial information as demonstrated in . Nevertheless, the computational cost caused by the self-attention when processing high-resolution images is considerable.

Recently, an interesting work, named ConvNeXt , reveals that by simply modernizing the standard ResNet and using the similar design and training recipe as Transformers, ConvNets can behave even better than some popular ViTs . RepLKNet also shows the potential of leveraging large-kernel convolutions for visual recognition. These explorations encourage many researchers to rethink the design of ConvNets by leveraging either large-kernel convolutions , or high-order spatial interactions , or sparse convolutional kernels , etc. Till now, how to more efficiently take advantage of convolutions to construct powerful ConvNet architectures is still a hot research topic.

In this paper, we are also interested in investigating new ways to more efficiently make use of spatial convolutions. Different from the ConvNeXt work that aims to adjust the training recipe or the position of spatial convolutions in building blocks, we compare the different ways ViTs and ConvNets use to encode spatial information. As shown in the left part of Fig. 1, self-attention computes the output of each pixel by a weighted summation of all other positions. This process can also be mimicked by computing the Hadamard product between the output of a large-kernel convolution and the value representations, which we call convolutional modulation as depicted in the right part of Fig. 1. The difference is that the convolutional kernels are static while the attention matrix generated by self-attention can adapt to the input. Experiments show that using convolutions to generate weight matrix yields great results as well.

Simply replacing the self-attention in ViTs with the proposed convolutional modulation operation yields the proposed network, termed Conv2Former. The meaning behind it is that we aim to use convolutions to construct a Transformer-style ConvNet, in which the convolutional features are used as weights to modulate the value representations. In contrast to the classic ViTs with self-attention, our method, like many classic ConvNets, is fully convolutional and hence its computations increase linearly rather than quadratically as in Transformers with the image resolution being higher. This makes our method more friendly to downstream tasks, like object detection and high resolution semantic segmentation. More interestingly, our Conv2Former can benefit more from convolutions with larger kernels, like 11×\times11 and 21×\times21. This is different from the conclusions made in previous ConvNets , which demonstrate using standard depthwise convolutions with kernel sizes larger than 9×\times9 brings nearly no performance gain but computational burden. We also show that our method performs better than the recent works using super large kernel convolutions .

We evaluate Conv2Former on popular vision tasks, including ImageNet classification , COCO object detection/instance segmentation , and ADE20k semantic segmentation . To validate the capability of Conv2Former on larger datasets, we also pretrain our model on the ImageNet-22k dataset and evaluate the performance on downstream tasks. Experiments show that Conv2Former performs better than popular ConvNets, like ConvNeXt and EfficientNetV2 . We hope our work could provide informative design choices for future visual recognition models.

Related Work

From ConvNets to the popular Vision Transformers, the architectures have always been keeping renovating . In this section, we briefly describe some representative visual recognition models.

The success of early visual recognition models is mostly dedicated to the development of ConvNets, typified by AlexNet , VGGNet , and GoogLeNet . These models, suffering from the gradient vanishing problem, mostly contain less than 20-layer convolutions. Later, the emerging of ResNets advances the conventional ConvNets by introducing shortcut connections, which make training very deep models possible. Inceptions and ResNeXt further enrich the design principles of ConvNets and propose to use building blocks with multiple parallel paths of specialized-filter convolutions. Instead of tuning network architectures, SENet and its follow-ups aim to improve ConvNets with lightweight attention modules that can explicitly model the inter-dependencies among channels. EfficientNets and MobileNetV3 take advantage of neural architecture search to search for efficient network architectures. RegNet presents a new network design paradigm for ConvNets by exploring network design spaces.

Very recently, some works aim to show the advantages of introducing large-kernel convolutions . A typical example should be VAN that utilizes a depthwise convolution and a dilated one to decompose large-kernel convolutions. Our Conv2Former is different from VAN in that we do not aim to decompose large-kernel convolutions but show self-attention can be reduced to the convolutional modulation operation, which yields good recognition performance as well. There are also some works leveraging different training or optimization methods or finetuning techniques to advance EfficientNet.

2 Vision Transformers

Transformers, originally designed for natural language processing tasks , have been widely used in visual recognition. The most typical work should be Vision Transformer (ViT) which shows the great potential of Transformers for processing large-scale data in image classification. DeiT improves the original ViT by using strong data augmentation methods and knowledge distillation and gets rid of the dependence of ViTs on large-scale data. Motivated by the success of pyramid architecture in ConvNets, some works design pyramid structures using Transformers to take advantage of multi-scale features. Some works propose to introduce local dependencies into ViTs, showing great performance in visual recognition. Besides, there are also some works exploring the scaling capability of ViTs in visual recognition. Specially, Yuan et al. show that a two-stage ViT outperforms the state-of-the-art CNNs on ImageNet for the first time.

3 Other Models

Some recent works show that mixing both Transformers and convolutions is a promising way to develop stronger visual recognition models especially for those aiming at efficient network design. A typical example should be MobileViT , which provides an efficient way to fuse both convolutions and Transformers. EfficientViT , EdgeNeXt , and MobileFormer bring back convolutions to Transformers and show great performance in both image classification and downstream tasks. Moreover, there are also hybrid networks that introduce different attention mechanisms into ConvNet for global context encoding . In addition, designing MLP-like architectures is also a popular research topic for visual recognition .

Model Design

In this section, we describe the architecture of our proposed Conv2Former and provide some useful suggestions in model design and layer adjustment.

Overall architecture. The overall architecture has been shown in Fig. 2. Similarly to the ConvNeXt and Swin Transformer network , our Conv2Former also adopts a pyramid architecture. There are four stages in total, each of which has a different feature map resolution. Between two consecutive stages, a patch embedding block is used to reduce the resolution, which is often a 2×22\times 2 convolution with stride 2. Different stages have different numbers of convolutional blocks. We build five Conv2Former variants, namely Conv2Former-N, Conv2Former-T, Conv2Former-S, Conv2Former-B, Conv2Former-L. Details are summarized in Tab. 1.

Stage configuration. When the number of learnable parameters is fixed, how to arrange the width and depth of the network has an impact on the model performance . The original ResNet-50 sets the number of blocks in each stage to (3,4,6,3)(3,4,6,3). ConvNeXt-T changes the block numbers to (3,3,9,3)(3,3,9,3) following the principle used in Swin-T and uses the stage compute ratio of 1:1:9:11:1:9:1 for larger models. Differently, we slightly adjust the ratios as shown in Tab. 1. We observe that for a tiny-sized model (with less than 30M parameters) deeper networks perform better. A brief comparison among four different tiny-sized models can be found in Tab. 2.

2 Convolutional Modulation Block

Our convolutional block used in each stage shares a similar structure to Transformers, which mainly contains a self-attention layer for spatial encoding and an FFN for channel mixing. Differently, we replace the self-attention layer with a simple convolutional modulation layer.

where A\mathbf{A} measures the relationships between each pair of input tokens, which can be written as

Advantages. A diagrammatic comparison among the residual block, self-attention, and the proposed modulation block can be found in Fig. 3. Compared to self-attention, our method utilizes convolutions to build relationships, which are more memory-efficient than self-attention especially when processing high-resolution images. Compared to the classic residual blocks , our method can also adapt to the input content due to the modulation operation.

3 Micro Design

Larger kernel than 7×\times7. How to make use of spatial convolutions is important for ConvNet design. Since VGGNet and ResNets , 3×33\times 3 convolutions have been a standard choice for building ConvNets. Later, the emerging of depthwise separable convolution changes this situation. ConvNeXt shows that enlarging the kernel size of ConvNets from 3 to 7 can improve the classification performance. However, further increasing the kernel size nearly brings no performance gain but computational burden without re-parameterization .

We argue that the reason making ConvNeXt benefit little from larger kernel sizes than 7×77\times 7 is the way to use spatial convolutions. For Conv2Former, we observe a consistent performance gain as the kernel size increases from 5×55\times 5 to 21×2121\times 21. This phenomenon not only happens for Conv2Former-T (82.8→83.482.8\rightarrow 83.4) but also holds for Conv2Former-B with 80M+ parameters (84.1→84.584.1\rightarrow 84.5). Considering the model efficiency, we set the kernel size to 11×1111\times 11 by default.

Weighting strategy. As shown in Fig. 3(d), we consider the outputs of depthwise convolutions as weights to modulate the features after the linear projection. It is worth noting that we use neither activation nor normalization layers (e.g., Sigmoid or LpL_{p} normalization) before the Hadamard product. This is an essential factor to attain good performance. For example, adding a Sigmoid function as done in SENet decreases the performance by more than 0.5%.

We want to stress that FocalNet adopts a similar weighting strategy as ours but its motivation is different. FocalNet aims to extract multi-level features via 3×33\times 3 depthwise convolutions and global average pooling for hierarchical context aggregation. Differently, we attempt to simplify the self-attention operation by leveraging simple large kernel convolutions and investigate an efficient way to make use of large kernel spatial convolutions for ConvNets. Our method is much simpler than FocalNet and experiments demonstrate the advantages of Conv2Former over FocalNet.

Normalization and activations. For normalization layers, we follow the original ViT and ConvNeXt and adopt the Layer Normalization instead of the widely-used batch normalization . For activation layers, we use GELU . We found that the combination of Layer Normalization and GELU brings 0.1%-0.2% performance gain.

Experiments

Datasets. We evaluate the classification performance of the proposed Conv2Former on the widely-used ImageNet-1k dataset , which contains around 1.2M training images and 1,000 different categories. We report the results on the validation set that has in total 50k images. Like some other popular models , we also test the scaling ability of the proposed Conv2Former using the large-scale ImageNet-22k dataset for pretraining, which has around 14M images and 21,841 classes. After pretraining, we use the ImageNet-1k dataset for finetuning and report results on the ImageNet-1k validation set as well.

Training settings. We implement our model based on PyTorch . During training, we use the AdamW optimizer with a linear learning rate scaling strategy lr=LRbase×batch_size/1024lr=\text{LR}_{\text{base}}\times\text{batch}\_\text{size}/{1024}. The initial learning rate LRbase\text{LR}_{\text{base}} is set to 0.001 and weight decay rate is set to 5×10−25\times 10^{-2} as suggested in previous work . Throughout the experiments on ImageNet, we randomly crop the image size to 224×224224\times 224 and adopt some common data augmentation methods, such as MixUp and CutMix . Stochastic Depth , Random Erasing , Label Smoothing , RandAug , and Layer Scale of initial value 1e-6 are used as well. We train all the models for 300 epochs. For experiments on the ImageNet-22k, we first pretrain our model on this dataset for 90 epochs and then finetuning on the ImageNet-1k dataset for 30 epochs following ConvNeXt .

2 Comparison with Other Methods

We compare our Conv2Former with some popular network architectures, including Swin Transformer , ConvNeXt , NFNet , DeiT , RegNet , FocalNet , EfficientNets , CoAtNet , RepLKNet , and MOAT . Note that some of them are hybrid models of CNNs and Transformers.

ImageNet-1k. We first train our Conv2Former on the ImageNet-1k dataset and show the results in Tab. 3. For tiny-sized models (<30<30M), our Conv2Former has 1.1% and 1.7% performance gains compared to ConvNeXt-T and SwinT-T, respectively. Even our Conv2Former-N with 15M parameters and 2.2G FLOPs performs the same as SwinT-T with 28M parameters and 4.5G FLOPs. For the base models, the performance gain decreases but there are still 0.6% and 0.9% improvement over ConvNeXt-B and SwinT-B. Compared to other popular models, our Conv2Former also perform better than those with similar model sizes. Notably, our Conv2Former-B even behaves better than EfficientNet-B7 (84.4% v.s. 84.3%), whose computations are two times larger than ours (37G v.s. 15G).

ImageNet-22k. We pretrain our Conv2Former on the large ImageNet-22k dataset and then finetune on the ImageNet-1k dataset. This experiment can reflect the data scaling capability of our Conv2Former. For all experiments, we follow the settings used in to train and finetune the models. The results have been listed in Tab. 4. Compared to the different variants of ConvNeXt, our Conv2Formers all perform better when the model sizes are similar. Typically, our Conv2Former-B performs better than ConvNeXt-B and the MOAT-2 network, which consumes more computations than ours. In addition, we can see that when finetuning on a larger resolution 384×384384\times 384 our Conv2Former-L attains better result than hybrid models, like CoAtNet and MOAT. Our Conv2Former-L achieves the best result 87.7%.

Discussion. Employing large-kernel convolutions is a straightforward way to assist CNNs in building long-range relationships. However, directly using large-kernel convolutions (>7×7>7\times 7) in existing CNN-based architectures makes the recognition models difficult to optimize . Recently, there are a few works aiming to develop new techniques to evoke the utilization of large-kernel convolutions in CNNs. In Tab. 5, we show the results by the recent state-of-the-art ConvNets with different kernel sizes. We can see that without any other training techniques, like re-parameterization or using sparse weights, our Conv2Former with kernel size 7×77\times 7 already performs better than other methods under the base model setting. Using a larger kernel size 11×1111\times 11 yields a better performance gain. These results reflect the advantage of our convolutional modulation block.

3 Method Analysis

In this subsection, we provide a series of method analysis on the proposed convolution modulation operation.

Kernel size. The ConvNeXt work shows that there is no performance gain when the kernel size of depthwise convolutions is more than 7×77\times 7. Here, we investigate how would the model performance change when larger kernel sizes are used. We select 6 different kernels for the depthwise convolutions, i.e., {5×5,7×7,9×9,11×11,15×15,21×21}\{5\times 5,7\times 7,9\times 9,11\times 11,15\times 15,21\times 21\} and show the results based on two model variants, Conv2Former-T and Conv2Former-B. The results can be found in Fig. 4(a). The performance gain seems to saturates until the kernel size is increased to 21×2121\times 21. This result is quite different from that made by ConvNeXt who concludes that using larger than 7×77\times 7 kernels brings no clear performance gain. This indicates that using the convolutional features as weights as formulated in Eqn. (3) can more efficiently take advantage of large kernels than traditional ways .

Hadamard product is better than summation. As shown in Fig. 3(d), we use the convolutional features extracted by the depthwise convolutions to modulate the weights of the right linear branch via the Hadamard product operation. In our experiments, we have also attempt to leverage the element-wise summation to fuse the two branches. Fig. 4(b) shows the comparison results on our Conv2Former at different model sizes. The Hadamard product performs better than element-wise summation, indicating convolutional modulation is more efficient than summation in encoding spatial information. We can also observe that small models benefit more from Hadamard product.

Weighting strategy. Other than the aforementioned two fusion strategies, we also attempt to use other ways to fuse the feature maps, including adding a Sigmoid function after A\mathbf{A}, applying L1L_{1} normalization to A\mathbf{A}, and linearly normalizing the values of A\mathbf{A} to (0,1](0,1]. The results are summarized in Tab. 6. We can see that the Hadamard product leads to better results than all other operations. More interestingly, when adjusting the values of A\mathbf{A} to positive values using either the Sigmoid function or linear normalization to (0,1](0,1], the performance drops more. This is different from the traditional attention mechanisms, like SE and CA that leverage the Sigmoid function before reweighing. We leave this for future research.

4 Results on Isotropic Models to ViTs

Different from the classic CNNs that adopt hierarchical architectures, the vanilla ViT due to the heavy self-attention layer utilizes a plain architecture that contains a patch embedding layer and a stack of Transformers with the same sequence length. This plain architecture has been widely used in recent works on Transformers. Here, we follow ConvNeXt and also attempt to investigate the performance of Conv2Former under the ViT-style architecture settings. Similar to ConvNeXt, we set the number of blocks to 18 for both Conv2Former-IS and Conv2Former-IB and adjust the channel numbers to match the model size. We use two versions of the patch embedding module: a 16×1616\times 16 convolution with stride 16 and three convolutions as done in .

Tab. 7 shows the results. We take the DeiT-S and DeiT-B model as baselines. For brevity, we add a letter ‘I’ in the model names, representing that the corresponding models use the isotropic architecture as the original ViT. We can see that for small-sized models with around 22M parameters, our Conv2Former-IS performs much better than DeiT-S and ConvNeXt-IS. The performance gain is around 1.5%. When scaling up the model size to 80M+, our Conv2Former-IB achieves a top-1 accuracy score of 82.7%, which is also 0.7% better than ConvNeXt-IB and 0.9% better than DeiT-B. In addition, using three convolutions for patch embedding can further improve the result.

5 Results on Downstream Tasks

In this subsection, we evaluate our method on two downstream tasks, including object detection on COCO and semantic segmentation ADE20k .

Results on COCO. MSCOCO is a large dataset for object detection, which contains 80 categories. Following previous works , we conduct experiments using two popular object detectors, Mask R-CNN and Cascade Mask R-CNN and report both the object detection and instance segmentation results. For training, we follow the experiment settings used in ConvNeXt , including multi-scale training, AdamW optimizer with a 3×\times learning schedule, GIoU loss , etc. Readers can refer to for more detailed experimental settings. We use the MMDetection toolbox to run all the object detection experiments.

The results can be found in Tab. 8. For tiny-sized models, our Conv2Former-T achieves about 2% AP improvement over SwinT-T and ConvNeXt-T when using the Mask R-CNN framework in object detection. For instance segmentation, the performance gain is also more than 1%. When using the Cascade Mask R-CNN framework, we can observe more than 1% performance gain than SwinT-T and ConvNeXt-T. When scaling up the models, the improvement is also clear.

Results on ADE20k. ADE20k is a popular semantic segmentation dataset. It contains 150 classes and a variety of scenes with 1,038 image-level labels. Following , we train the models using the training set and report results on the validation set. For tiny-, small-, base-sized models, we randomly crop the input image to 512×512512\times 512, and for the large-sized model, we use a crop size of 640×640640\times 640. We use the UperNet as our decoder.

Results are summarized in Tab. 9. For models at different scales, our Conv2Former can outperform both the Swin Transformer and ConvNeXt. Notably, there is a 1.3% mIoU improvement compared to ConvNeXt at the tiny scale and the improvement is 1.1% at the base scale. When we further increase the model size, our Conv2Former-L with UperNet achieves an mIoU score of 54.3%, which is also clearly better than Swin-L and ConvNeXt-L.

Conclusions and Discussions

This paper present Conv2Former, a new convolutional network architecture for visual recognition. The core of our Conv2Former is the convolutional modulation operation that simplifies the self-attention mechanism by using only convolutions and Hadamard product. We show that our convolutional modulation operation is a more efficient way to take advantage of large-kernel convolutions. Our experiments in ImageNet classification, object detection, and semantic segmentation also show that our proposed Conv2Former performs better than previous CNN-based models and most of the Transformer-based models.

Discussion. Recent state-of-the-art visual recognition models heavily rely on convolutions for low-level feature encoding. We believe there is still a large room to improve CNN-based models for visual recognition. For instance, how to more efficiently take advantage of large-kernel convolutions (≥7×7\geq 7\times 7), how to use convolutions with fixed-sized kernels to more effectively capture large receptive fields, and how to more effectively introduce lightweight attention mechanisms to CNNs all deserve further investigation.

Limitations. This paper aims at studing how to more effectively make use of large-kernel convolutions. We only pay attention to the design of CNN-based models. How to combine the proposed convolutional modulation block with Transformers warrants future study.

Appendix A More Experimental Settings

For all experiments, we use the cosine learning rate decay schedule for training. Most of other hyper-parameters used in our Conv2Former have been described in the main paper. Compared to ConvNeXt , we do not use layer-wise lr decay and EMA as we found they do not help in our Conv2Former training. Here, we show the stochastic depth rate we use for different variants of our Conv2Former. The stochastic depth rates we use for different model variants (pre)training can be found in Tab. 10.

Those for finetuning on ImageNet-1k can be found in Tab. 11.

A.2 COCO Detection

When training on the COCO datasets, we following the experimental settings as in , except the stochastic depth rates which are listed in Tab. 12.

A.3 ADE20k Semantic Segmentation

For semantic segmentation experiments on ADE20k , we following the settings used in except the stochastic depth rates that are summarized in Tab. 13.

References