Conformer: Local Features Coupling Global Representations for Visual Recognition
Zhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie, Yaowei Wang, Jianbin Jiao, Qixiang Ye
Introduction
Convolutional neural networks (CNNs) have significantly advanced computer vision tasks such as image classification, object detection, and instance segmentation. This largely attributes to the convolution operation, which collects local features in a hierarchical fashion as powerful image representations. Despite of the advantage upon local feature extraction, CNNs experience difficulty to capture global representations, long-distance relationships among visual elements, which are often critical for high-level computer visual tasks. An intuitive solution is enlarging the receptive field, which however could require more intensive yet damaging pooling operations.
Recently, the transformer architecture has been introduced to visual tasks . The ViT method constructs a sequence of tokens by splitting each image to patches with positional embeddings and applies cascaded transformer blocks to extract parameterized vectors as visual representations. Thanks to the self-attention mechanism and Multilayer Perceptron (MLP) structure, the visual transformer reflects complex spatial transforms and long-distance feature dependencies, which constitute global representations. Unfortunately, visual transformers are observed ignoring local feature details which decreases the discriminability between background and foreground, Figs. 1(c) and (g). Improved visual transformers have proposed a tokenization module or leveraged CNN feature maps as input tokens to capture feature neighboring information. Nevertheless, the problem about how to precisely embed local features and global representations to each other remains.
In this paper, we propose a dual network structure, termed Conformer, with the aim to couple CNN-based local features with transformer-based global representations for enhanced representation learning. Conformer consists of a CNN branch and a transformer branch which respectively follow the design of ResNet and ViT . The two branches form a comprehensive combination of local convolution blocks, self-attention modules, and MLP units. During training, the cross entropy losses are used to supervise both the CNN and transformer branches to couple CNN-style and transformer-style features.
Considering the feature misalignment between CNN and transformer features, the Feature Coupling Unit (FCU) is designed as the bridge. On the one hand, to fuse the two-style features, FCU leverages 11 convolution to align the channel dimensions, down/up sampling strategies to align feature resolutions, LayerNorm and BatchNorm to align feature values. On the other hand, since CNN and transformer branches tend to capture features of different levels (, local vs. global), FCU is inserted into every block to consecutively eliminate the semantic divergence between them, in an interactive fashion. Such a fusion procedure can greatly enhance the global perception capability of local features and the local details of global representations.
The ability of Conformer in coupling local features and global representations is demonstrated in Fig. 1. While conventional CNNs (, ResNet-101) tend to retain discriminative local regions (, the peacock’s head or tail), the CNN branch of Conformer can activate the full object extent, Figs. 1(b) and (f). When solely using the visual transformers, for the weak local features (, blurred object boundaries), it is difficult to distinguish the object from the background, Figs. 1(c) and (g). The coupling of local features and global representations significantly enhances the discriminability of transformer-based features, Figs. 1(d) and (h).
We propose a dual network structure, termed Conformer, which retains local features and global representations to the maximum extent.
We propose the Feature Coupling Unit (FCU), to fuse convolutional local features with transformer-based global representations in an interactive fashion.
Under comparable parameter complexity, Conformer outperforms CNNs and visual transformers by significant margins. Conformer inherits the structure and generalization advantages of both CNNs and visual transformers, demonstrating the great potential to be a general backbone network.
Related Work
CNNs with Global Cues. In the deep learning era, CNNs can be regarded as a hierarchical ensemble of local features with different reception fields. Unfortunately, most CNNs are good at extracting local features but experience difficulty to capture global cues.
To alleviate such a limitation, one solution is to define larger receptive fields by introducing deeper architectures and/or more pooling operations . The dilated convolution methods increased the sampling step size, while deformable convolution learned the sampling positions. SENet and GENet proposed to use global Avgpooling to aggregate global context and then used it to reweight feature channels, while CBAM respectively used global Maxpooling and global Avgpooling to refine features independently in the spatial and channel dimensions.
The other solution is the global attention mechanism , which has demonstrated great advantage in capturing long-distance dependencies in natural language processing . Inspired by the non-local means method , the non-local operation was introduced to CNNs in a self-attention manner so that the response at each position is a weighted sum of the features at all (global) positions. Attention augmented convolutional networks concatenated convolutional feature maps with self-attentional feature maps to augment convolution operations for capturing long-range interactions. Relation Networks proposed an object attention module, which processes a set of objects simultaneously through interaction between their appearance feature and geometry.
Despite of the progress, existing solutions that introduce global cues to CNNs have obvious disadvantages. For the first solution, larger receptive fields require more intensive pooling operations, which implies lower spatial resolution. For the second solution, if convolutional operations are not properly fused with attention mechanisms, local feature details could deteriorate.
Visual Transformers. As a pioneered work, ViT validated the feasibility of pure transformer architectures for computer vision tasks. To leverage the long-distance dependencies, transformer blocks acted as independent architectures or were introduced to CNNs for image classification , object detection , semantic segmentation , image enhancement and image generation . However, the self-attention mechanism in visual transformers often ignores local feature details. To solve, DeiT proposed using a distillation token to transfer CNN-based features to visual transformer while T2T-ViT proposed using a tokenization module to recursively reorganize the image to tokens considering neighboring pixels. The DETR method fed local features extracted by CNN to the transformer encoder-decoder to model the global relationships between features in a serial fashion.
Different from existing works, Conformer defines the first concurrent network structure which fuses features in an interactive fashion. Such a structure not only naturally inherits the structure advantages of both CNN and transformers but also retains the representation capability of local features and global representations to the maximum extent.
Conformer
Local features and global representations are important counterparts, which have been extensively studied in the long history of visual descriptors. Local features and their descriptors , which are compact vector representations of local image neighborhoods, have been the building blocks of many computer vision algorithms. Global representations include, but not limited to, contour representations, shape descriptors, and object typologies at long-distance . In the deep learning era, CNN collects local features in a hierarchical manner via convolutional operations and retains the local cues as feature maps. Visual transformer is believed to aggregate global representations among the compressed patch embeddings in a soft fashion by the cascaded self-attention modules.
In order to take advantage of local features and global representations, we design a concurrent network structure, as shown in Fig. 2(c), termed Conformer. Considering the complementarity of the two-style features, within Conformer, we consecutively feed the global context from the transformer branch to feature maps, to reinforce the global perception capability of the CNN branch. Similarly, local features from the CNN branch are progressively fed back to patch embeddings, to enrich the local details of the transformer branch. Such a process constitutes the interaction.
In special, Conformer is composed of a stem module, dual branches, FCUs to bridge dual branches, and two classifiers (a fc layer) for the dual branches. The stem module, which is a 77 convolution with stride 2 followed by a 33 max pooling with stride 2, is used to extract initial local features (, edge and texture information), which are then fed to the dual branches. The CNN branch and transformer branch are composed of (, 12) repeated convolution and transformer blocks, respectively, as described in Tab. 1. Such a concurrent structure implies that CNN and transformer branch can respectively preserve the local features and global representations to the maximum extent. FCU is proposed as a bridge module to fuse local features in the CNN branch with global representations in the transformer branch, Fig. 2(b). FCU is applied from the second block because the initialized features of the two branches are the same. Along the branches, FCU progressively fuses feature maps and patch embeddings in an interactive fashion.
Finally, for the CNN branch, all the features are pooled and fed to one classifier. For the transformer branch, the class token is taken out and fed to the other classifier. During training, we use two cross entropy losses to separately supervise the two classifiers. The importance of the loss functions are empirically set to be same. During inference, the outputs of the two classifiers are simply summarized as the prediction results.
2 Network Structure
CNN Branch. As shown in Fig. 2(b), the CNN branch adopts feature pyramid structure, where the resolution of feature maps decreases with network depth while the channel number increases. We split the whole branch into 4 stages, as described in Tab. 1(CNN Branch). Each stage is composed of multiple convolution blocks and each convolution block contains bottlenecks. Following the definition in ResNet , a bottleneck contains a 11 down-projection convolution, a 33 spatial convolution, a 11 up-projection convolution, and a residual connection between the input and output of the bottleneck. In experiments, is set to be 1 in the first convolution block and satisfies in the subsequent convolution blocks.
Visual transformers project an image patch into a vector through a single step, causing the lost of local details. While in CNNs, convolution kernels slide over feature maps with overlap, which provides the possibility to preserve fine-detailed local features. Consequently, the CNN branch is able to consecutively provide local feature details for the transformer branch.
Transformer Branch. Following ViT , this branch contains repeated transformer blocks. As shown in Fig. 2(b), each transformer block consists of a multi-head self-attention module and an MLP block (contains a up-projection fc layer and a down-projection fc layer). LayerNorms are applied before each layer and residual connections in both the self-attention layer and MLP block. For tokenization, we compress the feature maps generated by the stem module into 1414 patch embeddings without overlap, by a linear projection layer, which is a 44 convolution with stride 4. A class token is then pretended to the patch embeddings for classification. Considering that the CNN branch (33 convolution) encodes both local features and spatial location information , the positional embeddings are no longer required. This facilities increasing image resolution for downstream vision tasks.
Feature Coupling Unit. Given the feature maps in the CNN branch and patch embeddings in the transformer branch, how to eliminate the misalignment between them is an important issue. To solve, we propose the FCU to consecutively couple local features with global representations in an interactive manner.
On the one hand, we must realize that the feature dimensinalities of CNN and transformer are inconsistent. The CNN feature maps have the dimensinality (, , are channels, height and width respectively), while the shape of the patch embeddings is , where , 1, and respectively represent the number of image patches, class token and embedding dimensions. When fed to the transformer branch, feature maps first require to get through 11 convolution to align the channel numbers of the patch embeddings. A down-sampling module (Fig. 2(a)) is then used to complete the spatial dimension alignment. Finally, the feature maps are added with patch embeddings, as shown in Fig. 2(b). When fed back from the transformer branch to the CNN branch, the patch embeddings require to be up-sampled (Fig. 2(a)) to align the spatial scale. The channel dimension is then aligned with that of CNN feature maps through the 11 convolution, and added to the feature maps. Meanwhile, LayerNorm and BatchNorm modules are used to regularize features.
On the other hand, there is a significant semantic gap between feature maps and patch embeddings, , feature maps are collected from the local convolutional operators while patch embeddings are aggregated with the global self-attention mechanisms. FCU is therefore applied in each block (except the first) to progressively fill the semantic gap.
3 Analysis and Discussion
Structure Analysis. By considering the FCU as a short connection, we can abstract the proposed dual structure into the special serial residual structure, as shown in Fig. 3(a). Under different residual connection units, Conformer can implement different depths combinations of bottlenecks (as in ResNet, Fig. 3(b)) and transformer blocks (as in ViT, Fig. 3(d)), implying that Conformer inherits the structural advantages of both CNNs and visual transformers. Furthermore, it achieves different permutations of bottlenecks and transformer blocks at different depths, including but not limited to Figs. 3(c) and (e). This greatly enhances the representation capacity of the network.
Feature Analysis. We visualize the feature maps in Fig. 1, class activation maps and attention maps in Fig. 4. Compared with ResNet , with the coupled global representations, the CNN branch of Conformer tends to activate larger regions rather than local areas, suggesting enhanced long-distance feature dependencies, which are significantly demonstrated in Figs. 1(f) and 4(a). Thanks to the fine-detailed local features progressively provided by the CNN branch, the patch embeddings of the transformer branch in the Conformer retain important detailed local features (Figs. 1(d) and (h)), which are deteriorated by the visual transformers (Figs. 1(c) and (g)). Furthermore, the attention area in Fig. 4(b) is more complete while the background is significantly suppressed, implying the higher discriminative capacity of the learned feature representations by Conformer.
Experiments
By tuning the parameters of the CNN and transformer branches, we have the model variants, termed Conformer-Ti, -S, and -B, respectively. The details of Conformer-S are described in Tab. 1, and those of Conformer-Ti/B are in the Appendix. Conformer-S/32 splits the feature maps to 77 patches, , the patch size is 3232 in the transformer branch.
2 Image Classification
Experimental Setting. Conformer is trained on the ImageNet-1k training set with 1.3M images and tested upon the validation set. The Top-1 accuracy is reported in Tab. 2. To make the transformer converge to a reasonable performance, we follow the data augmentation and regularization techniques in DeiT . These techniques include Mixup , CutMix , Erasing , Rand-Augment and Stochastic Depth ). The model is trained for 300 epochs with the AdamW optimizer , batchsize 1024 and weight decay 0.05. The initial learning rate is set to 0.001 and decay in a cosine schedule.
Performance. Under similar parameters and computational budgets, Tab. 2, Conformers outperform both CNN and visual transformers. For example, Conformer-S (with 37.7M parameters and 10.6G MACs) respectively outperforms ResNet-152 (with 60.2M parameters and 11.6G MACs) by 4.1%((83.4% vs. 78.3%) and DeiT-B (with 86.6M parameters and 17.6G MACs) by 1.6% (83.4% vs. 81.8%). Conformer-B, with comparable parameters and moderate MAC cost, outperforms DeiT-B by 2.3% (84.1% vs. 81.8%). Beyond its superior performance, Conformer converges faster than the visual transformers.
3 Object Detection and Instance Segmentation
4 Ablation Studies
Number of Parameters. The parameters of the proposed Conformer are combinations of the CNN and transformer branches. The parameter proportion of the two branches is a hyper-parameter to be experimentally determined. In Tab. 4, we evaluate performance of the two branches under different parameter settings. For the CNN branch, we tune the parameters of the CNN branch by changing the channels and the number of bottlenecks, which respectively control the width and depth of the CNN branch. For the transformer branch, we tune the parameters by changing the numbers of embedding dimensions and heads. From Tab. 4, one can see that the accuracy is improved by increasing either parameters of the CNN or the transformer branch. More CNN parameters bring greater improvement while the computational cost overhead is lower.
Dual Structure. Conformer is a dual model, which is totally different from the serial hybrid ViT (CNN Transformer) . In Tab. 5, ResNet-26/50d & DeiT-S is a hybrid model which consists of ResNet-26/50d and DeiT-S , where DeiT-S forms tokens upon the feature maps extracted by ResNet-26/50d. With comparable computational cost overhead, Conformer-S/32 outperforms the serial hybrid model although ResNet-26/50d can retain more local information within the stem stage.
Positional Embeddings. Considering that the CNN branch (33 convolution) encodes both local features and spatial location information, the positional embeddings are assumed no longer required for Conformer. In Tab. 6, when the positional embedding is removed, the accuracy of DeiT-S decreases 2.4%, while that of Conformer-S decreases marginally (0.1%).
Sampling Strategies. In FCU, to make CNN-based feature maps coupling with Transformer-based patch embeddings, up/down-sampling operations are used to align the spatial scale. In Tab. 7, we compare different up/down-sampling strategies including Maxpooling, Avgpooling, convolution and attention-based samplingRefer to Appendix for detailed attention-based sampling.. Compared with Max/Avgpooling sampling, convolution and attention-based sampling methods use more parameters and computation cost but achieve comparable accuracy. We thereby choose the Avgpooling strategy.
Comparison with Ensemble Models. Conformer is compared with the ensemble models combining the outputs of CNN and transformer. For fair comparison, we use the same data augmentation and regularization strategies and the same training epochs (300) to train ResNet-101 , and combine it with the DeiT-S model to form an ensemble model, and report the accuracy in Tab. 8. The accuracies of the CNN branch, the transformer branch, and the Conformer-S respectively reach 83.3%, 83.1%, and 83.4%. In contrast, the ensemble model (DeiT-S+ResNet-101) archives 81.8%, which is 1.6% lower than that of Conformer-S (83.4%), although it uses significantly more parameters and MACs.
5 Generalization Capability
Rotation Invariance. To verify the generalization capability of the model in terms of rotation, we rotate test images by , , , , and and evaluate the performance of models trained under same data augmentation settings. As shown in Fig. 5(a), all models report comparable performance for images without rotation (). For the rotated test images, the performance of ResNet-101 drops significantly. In contrast, Conformer-S reports higher performance, which implies stronger rotation invariance.
Scale Invariance. In Fig. 5(b), we compare the scale adaptation ability of Conformer with those of visual transformers (DeiT-S) and CNN (ResNet). We interpolate the positional embeddings of DeiT-S to adapt it to input images of different resolutions during inference. When the size of input images reduces from 224 to 112, DeiT-S’s performance drops by 25% and that of ResNet-50/152 drops by 15%. In contrast, the performance of Conformer drops only by 10%, demonstrating higher scale invariance of the learned feature representations.
Conclusion
We propose Conformer, the first dual backbone to combining CNN with visual transformer. Within Conformer, we leverage the convolution operators to extract local features and the self-attention mechanisms to capture global representations. We design the Feature Coupling Unit (FCU) to fuse local features and global representations, enhancing the ability of visual representations in an interactive fashion. Experiments show that Conformer, with comparable parameters and computation budgets, outperforms both conventional CNNs and visual transformers, in striking contrast with the state-of-the-arts. On downstream tasks, Conformer has shown the great potential to be a simple yet effective backbone network.
References
Appendix A Model Architectures
The architectures of Conformer-Ti/B are detailed in Tab. 12. Compared with Conformer-S, Conformer-Ti reduces channel number of the CNN branch by 1/4, and Conformer-B increases channel number in the CNN branch, head number of the multi-head attention module and the embedding dimensions in the transformer branch by 1.5.
Appendix B Attention-based Sampling
We also design a down-sampling-up-sampling strategy based on the cross attention between feature maps and patch embeddings.
Let , and respectively denote the height, width, channel of feature maps in a block (we omit the batch dimension here for simplicity), and respectively represent the number of patch embeddings (termed ) and channel dimension in the transformer branch. We split the feature maps into patches (, 1414), termed . The dimension of each patch is . After aligning the channel dimension by 11 convolution, the shape of each patch is .
For down sampling, the fusion between patch in (denoted ) and patch in (denoted ) is formulated as
For up sampling, we re-use the attention weights in Eq. 1 and formulate the process as
Appendix C Inference Time
Classification. Following DeiT , we evaluate and compare the throughput of various methods in Fig. 6. One can see that our Conformer outperforms EfficientNet under comparable throughput.
Object detection and instance segmentation. Similarly, we measure Frame Per Second (FPS) as the inference speed and show the comparison in the Tab. 9. Combining Tab.3 in the paper and Tab. 9 here, compared with ResNet-101 , Conformer-S/32 has the comparable parameters, GFLOPs and inference speed, but can outperform ResNet-101 by a significant margin on both object detection and instance segmentation tasks, which further demonstrates the potential to be a general backbone network.
Appendix D Residual Structure
As shown in Fig. 3 in the paper, by considering FCUs as short connection we abstract Conformer with a dual structure to a serial structure with residual connections. In other words, under different residual connections, Conformer can degenerate to different sub-structures. We test some sub-structures and report the corresponding performance in Tab. 10. From Tab. 10, one can see that the proposed residual structure outperforms other sub-structures.
Appendix E Fusion Interval
In the paper, we proposed a Feature Coupling Unit to interact the local features and global representations in each block to progressively align the features to fill the semantic gap. To validate whether fusion should be done in each block, we conduct experiments on fusion intervals and report the performance on ImageNet in Tab. 11. From Tab. 11, one can see that smaller fusion intervals report higher performance, implying that frequent interaction facilities the representation learning.
Appendix F Convergence speed
For the convolution operations introduced, Fig. 7, both the CNN branch and the transformer branch of Conformer-S significantly outperforms DeiT during the first 50 epochs. This demonstrates the inductive bias of convolution facilities the convergence of visual transformers.