Contextual Transformer Networks for Visual Recognition
Yehao Li, Ting Yao, Yingwei Pan, Tao Mei
Introduction
Convolutional Neural Networks (CNN) demonstrates high capability of learning discriminative visual representations, and convincingly generalizes well to a series of Computer Vision (CV) tasks, e.g., image recognition, object detection, and semantic segmentation. The de-facto recipe of CNN architecture design is based on discrete convolutional operators (e.g., 33 or 55 convolution), which effectively impose spatial locality and translation equivariance. However, the limited receptive field of convolution adversely hinders the modeling of global/long-range dependencies, and such long-range interaction subserves numerous CV tasks . Recently, Natural Language Processing (NLP) field has witnessed the rise of Transformer with self-attention in powerful language modeling architectures that triggers long-range interaction in a scalable manner. Inspired by this, there has been a steady momentum of breakthroughs that push the limits of CV tasks by integrating CNN-based architecture with Transformer-style modules. For example, ViT and DETR directly process the image patches or CNN outputs using self-attention as in Transformer. present a stand-alone design of local self-attention module, which can completely replace the spatial convolutions in ResNet architectures. Nevertheless, previous designs mainly hinge on the independent pairwise query-key interaction for measuring attention matrix as in conventional self-attention block (Figure 1 (a)), thereby ignoring the rich contexts among neighbor keys.
In this work, we ask a simple question - is there an elegant way to enhance Transformer-style architecture by exploiting the richness of context among input keys over 2D feature map? For this purpose, we present a unique design of Transformer-style block, named Contextual Transformer (CoT), as shown in Figure 1 (b). Such design unifies both context mining among keys and self-attention learning over 2D feature map in a single architecture, and thus avoids introducing additional branch for context mining. Technically, in CoT block, we first contextualize the representation of keys by performing a 33 convolution over all the neighbor keys within the 33 grid. The contextualized key feature can be treated as a static representation of inputs, that reflects the static context among local neighbors. After that, we feed the concatenation of the contextualized key feature and input query into two consecutive convolutions, aiming to produce the attention matrix. This process naturally exploits the mutual relations among each query and all keys for self-attention learning with the guidance of the static context. The learnt attention matrix is further utilized to aggregate all the input values, and thus achieves the dynamic contextual representation of inputs to depict the dynamic context. We take the combination of the static and dynamic contextual representation as the final output of CoT block. In summary, our launching point is to simultaneously capture the above two kinds of spatial contexts among input keys, i.e., the static context via 33 convolution and the dynamic context based on contextualized self-attention, to boost visual representation learning.
Our CoT can be viewed as a unified building block, and is an alternative to standard convolutions in existing ResNet architectures without increasing the parameter and FLOP budgets. By directly replacing each 33 convolution in a ResNet structure with CoT block, we present a new Contextual Transformer Networks (dubbed as CoTNet) for image representation learning. Through extensive experiments over a series of CV tasks, we demonstrate that our CoTNet outperforms several state-of-the-art backbones. Notably, for image recognition on ImageNet, CoTNet obtains a 0.9% absolute reduce of the top-1 error rate against ResNeSt (101 layers). For object detection and instance segmentation on COCO, CoTNet absolutely improves ResNeSt with 1.5% and 0.7% mAP, respectively.
Related Work
Sparked by the breakthrough performance on ImageNet dataset via AlexNet , Convolutional Networks (ConvNet) has become a dominant architecture in CV field. One mainstream of ConvNet design follows the primary rule in LeNet , i.e., stacking low-to-high convolutions in series by going deeper: 8-layer AlexNet, 16-layer VGG , 22-layer GoogleNet , and 152-layer ResNet . After that, a series of innovations have been proposed for ConvNet architecture design to strengthen the capacity of visual representation. For example, inspired by split-transform-merge strategy in Inception modules, ResNeXt upgrades ResNet with aggregated residual transformations in the same topology. DenseNet additionally enables the cross-layer connections to boost the capacity of ConvNet. Instead of exploiting spatial dependencies in ConvNet , SENet captures the interdependencies between channels to perform channel-wise feature recalibration. further scales up an auto-searched ConvNet to obtain a family of EfficientNet networks, which achieve superior accuracy and efficiency.
2 Self-attention in Vision
Taking the inspiration from self-attention in Transformer that continuously achieves the impressive performances in various NLP tasks, the research community starts to pay more attention to self-attention in vision scenario. The original self-attention mechanism in NLP domain is devised to capture long-range dependency in sequence modeling. In vision domain, a simple migration of self-attention mechanism from NLP to CV is to directly perform self-attention over feature vectors across different spatial locations within an image. In particular, one of the early attempts of exploring self-attention in ConvNet is the non-local operation that severs as an additional building block to employ self-attention over the outputs of convolutions. further augments convolutional operators with global multi-head self-attention mechanism to facilitate image classification and object detection. Instead of using global self-attention over the whole feature map that scale poorly, employ self-attention within local patch (e.g., 33 grid). Such design of local self-attention effectively limits the parameter and computation consumed by the network, and thus can fully replace convolutions across the entirety of deep architecture. Recently, by reshaping raw images into a 1D sequence, a sequence Transformer is adopted to auto-regressively predict pixels for self-supervised representation learning. Next, directly apply a pure Transformer to the sequences of local features or image patches for object detection and image recognition. Most recently, designs a powerful backbone by replacing the final three 33 convolutions in a ResNet with global self-attention layers.
3 Summary
Here we also focus on exploring self-attention for the architecture design of vision backbone. Most of existing techniques directly capitalize on the conventional self-attention and thus ignore the explicit modeling of rich contexts among neighbor keys. In contrast, our Contextual Transformer block unifies both context mining among keys and self-attention learning over feature map in a single architecture with favorable parameter budget.
Our Approach
In this section, we first provide a brief review of the conventional self-attention widely adopted in vision backbones. Next, a novel Transformer-style building block, named Contextual Transformer (CoT), is introduced for image representation learning. This design goes beyond conventional self-attention mechanism by additionally exploiting the contextual information among input keys to facilitate self-attention learning, and finally improves the representational properties of deep networks. After replacing 33 convolutions with CoT block across the whole deep architecture, two kinds of Contextual Transformer Networks, i.e., CoTNet and CoTNeXt deriving from ResNet and ResNeXt , respectively, are further elaborated.
where is the head number, and \mathbin{\hbox{\hskip 5.0pt\hskip-5.0pt\hbox{\hbox{}}\hskip-5.0pt\hskip-2.5pt\raisebox{0.1736pt}{\hbox{\rule{0.0pt}{0.0pt}\rule{0.0pt}{0.0pt}\hbox{}}}\hskip-2.5pt\hskip 5.0pt}} denotes the local matrix multiplication operation that measures the pairwise relations between each query and the corresponding keys within the local grid in space. Thus, each feature at -th spatial location of is a -dimensional vector, that consists of local query-key relation maps (size: ) for all heads. The local relation matrix is further enriched with the position information of each grid:
Note that the local attention matrix of each head is only utilized for aggregating evenly divided feature map of along channel dimension, and the final output is the concatenation of aggregated feature maps for all heads.
2 Contextual Transformer Block
Conventional self-attention nicely triggers the feature interactions across different spatial locations depending on the inputs themselves. Nevertheless, in the conventional self-attention mechanism, all the pairwise query-key relations are independently learnt over isolated query-key pairs, without exploring the rich contexts in between. That severely limits the capacity of self-attention learning over 2D feature map for visual representation learning. To alleviate this issue, we construct a new Transformer-style building block, i.e., Contextual Transformer (CoT) block in Figure 2 (b), that integrates both contextual information mining and self-attention learning into a unified architecture. Our launching point is to fully exploit the contextual information among neighbour keys to boost self-attention learning in an efficient manner, and strengthen the representative capacity of the output aggregated feature map.
In other words, for each head, the local attention matrix at each spatial location of is learnt based on the query feature and the contextualized key feature, rather than the isolated query-key pairs. Such way enhances self-attention learning with the additional guidance of the mined static context . Next, depending on the contextualized attention matrix , we calculate the attended feature map by aggregating all values as in typical self-attention:
In view that the attended feature map captures the dynamic feature interactions among inputs, we name as the dynamic contextual representation of inputs. The final output of our CoT block () is thus measured as the fusion of the static context and dynamic context through attention mechanism .
3 Contextual Transformer Networks
The design of our CoT is a unified self-attention building block, and acts as an alternative to standard convolutions in ConvNet. As a result, it is feasible to replace convolutions with their CoT counterparts for strengthening vision backbones with contextualized self-attention. Here we present how to integrate CoT blocks into existing state-of-the-art ResNet architectures (e.g., ResNet and ResNeXt ) without increasing parameter budget significantly. Table 1 and Table 2 shows two different constructions of our Contextual Transformer Networks (CoTNet) based on the ResNet-50/ResNeXt-50 backbone, called CoTNet-50 and CoTNeXt-50, respectively. Please note that our CoTNet is flexible to generalize to deeper networks (e.g., ResNet-101).
CoTNet-50. Specifically, CoTNet-50 is built by directly replacing all the 33 convolutions (in the stages of res2, res3, res4, and res5) in ResNet-50 with CoT blocks. As our CoT blocks are computationally similar with the typical convolutions, CoTNet-50 has similar (even slightly smaller) parameter number and FLOPs with ResNet-50.
CoTNeXt-50. Similarly, for the construction of CoTNeXt-50, we first replace all the 33 convolution kernels in group convolutions of ResNeXt-50 with CoT blocks. Compared to typical convolutions, the depth of the kernels within group convolutions is significantly decreased when the number of groups (i.e., in Table 2) is increased. In ResNeXt-50, the computational cost of group convolutions is thus reduced by a factor of . Therefore, in order to achieve the similar parameter number and FLOPs with ResNeXt-50, we additionally reduce the scale of input feature map of CoTNeXt-50 from 324d to 248d. Finally, CoTNeXt-50 requires only 1.2 more parameters and 1.01 more FLOPs than ResNeXt-50.
4 Connections with Previous Vision Backbones
In this section, we discuss the detailed relations and differences between our Contextual Transformer and the previous most related vision backbones.
Blueprint Separable Convolution approximates the conventional convolution with a 11 pointwise convolution plus a depthwise convolution, aiming to reduce the redundancies along depth axis. In general, such design has some commonalities with the transformer-style block (e.g., the typical self-attention and our CoT block). This is due to that the transformer-style block also utilizes 11 pointwise convolution to transform the inputs into values, and the followed aggregation computation with local attention matrix is performed in a similar depthwise manner. Besides, for each head, the aggregation computation in transformer-style block adopts channel sharing strategy for efficient implementation without any significant accuracy drop. Here the utilized channel sharing strategy can also be interpreted as the tied block convolution , which shares the same filters over equal blocks of channels.
Dynamic Region-Aware Convolution introduces a filter generator module (consisting of two consecutive 11) to learn specialized filters for region features at different spatial locations. It therefore shares a similar spirit with the attention matrix generator in our CoT block that achieves dynamic local attention matrix for each spatial location. Nevertheless, the filter generator module in produces the specialized filters based on the primary input feature map. In contrast, our attention matrix generator fully exploits the complex feature interactions between contextualized keys and queries for self-attention learning.
Bottleneck Transformer is the contemporary work, which also aims to augment ConvNet with self-attention mechanism by replacing 33 convolution with Transformer-style module. Specifically, it adopts global multi-head self-attention layers, which are computationally more expensive than local self-attention in our CoT block. Therefore, with regard to the same ResNet backbone, BoT50 in only replaces the final three 33 convolutions with Bottleneck Transformer blocks, while our CoT block can completely replace 33 convolutions across the whole deep architecture. In addition, our CoT block goes beyond typical local self-attention in by exploiting the rich contexts among input keys to strengthen self-attention learning.
Experiments
In this section, we verify and analyze the effectiveness of our Contextual Transformer Networks (CoTNet) as a backbone via empirical evaluations over multiple mainstream CV applications, ranging from image recognition, object detection, to instance segmentation. Specifically, we first undertake experiments for image recognition task on ImageNet benchmark by training our CoTNet from scratch. Next, after pre-training CoTNet on ImageNet, we further evaluate the generalization capability of the pre-trained CoTNet when transferred to downstream tasks of object detection and instance segmentation on COCO dataset .
Setup. We conduct image recognition task on the ImageNet dataset, which consists of 1.28 million training images and 50,000 validation images derived from 1,000 classes. Both of the top-1 and top-5 accuracies on the validation set are reported for evaluation. For this task, we adopt two different training setups in the experiments, i.e., the default training setup and advanced training setup.
The default training setup is the widely adopted setting in classic vision backbones (e.g., ResNet , ResNeXt , and SENet ), that trains networks for around 100 epochs with standard preprocessing. Specifically, each input image is cropped into 224224, and only the standard data augmentation (i.e., random crops and horizontal flip with 50% probability) is performed. All the hyperparameters are set as in official implementations without any additional tuning. Similarly, our CoTNet is trained in an end-to-end manner, through backpropagation using SGD with momentum 0.9 and label smoothing 0.1. We set the batch size as that enables applicable implementations on an 8-GPU machine. For the first five epochs, the learning rate is scaled linearly from 0 to , which is further decayed via cosine schedule . As in , we adopt exponential moving average with weight 0.9999 during training.
For fair comparison with state-of-the-art backbones (e.g., ResNeSt , EfficientNet and LambdaNetworks ), we additionally involve the advanced training setup with longer training epochs and improved data augmentation & regularization. In this setup, we train our CoTNet with 350 epochs, coupled with the additional data augmentation of RandAugment and mixup , and the regularization of dropout and DropConnect .
Performance Comparison. We compare with several state-of-the-art vision backbones with two different training settings (i.e., default and advanced training setups) on ImageNet dataset. The performance comparisons are summarized in Tables 3 and 4 for each kind of training setup, respectively. Note that we construct several variants of our CoTNet and CoTNeXt with two kinds of depthes (i.e., 50-layer and 101-layer), yielding CoTNet-50/101 and CoTNeXt-50/101. In advanced training setup, as in LambdaResNet , we additionally include an upgraded version of our CoTNet, i.e., SE-CoTNetD-101, where the 33 convolutions in the res4 and res5 stages are replaced with CoT blocks under SE-ResNetD-50 backbone. Moreover, in default training setup, we also report the performances of our models with the use of exponential moving average for fair comparison against LambdaResNet.
As shown in Table 3, under the same depth (50-layer or 101-layer), the results across both top-1 and top-5 accuracy consistently indicate that our CoTNet-50/101 and CoTNeXt-50/101 obtain better performances against existing vision backbones with favorable parameter budget, including both ConvNets (e.g., ResNet-50/101 and ResNeXt-50/101) and attention-based models (e.g., Stand-Alone and AA-ResNet-50/101). The results generally highlight the key advantage of exploiting contextual information among keys in self-attention learning for visual recognition task. Specifically, under the same 50-layer backbones, by exploiting local self-attention in the deep architecture, LR-Net-50 and Stand-Alone exhibit better performance than ResNet-50, which ignores long-range feature interactions. Next, AA-ResNet-50 and LambdaResNet-50 enable the exploration of global self-attention over the whole feature map, and thereby boost up the performances. However, the performances of AA-ResNet-50 and LambdaResNet-50 are still lower than the stronger ConvNet (SE-ResNeXt-50) that strengthens the capacity of visual representation with channel-wise feature re-calibration. Furthermore, by fully replacing 33 convolutions with CoT blocks across the entirety of deep architecture in ResNet-50/ResNeXt-50, CoTNet-50 and CoTNeXt-50 outperform SE-ResNeXt-50. This confirms that unifying both context mining among keys and self-attention learning into a single architecture is an effective way to enhance representation learning and thus boost visual recognition. When additionally using exponential moving average as in LambdaResNet, the top-1 accuracy of CoTNeXt-50/101 will be further improved to 80.2% and 81.3% respectively, which is to-date the best published performance on ImageNet in default training setup.
Similar observations are also attained in advanced training setup, as summarized in Table 4. Note that here we group all the baselines with similar top-1/top-5 accuracy or network depth. In general, our CoTNet-50 & CoTNeXt-50 or CoTNet-101 & CoTNeXt-101 perform consistently better than other vision backbones across both metrics for each group. In particular, the top-1 accuracy of our CoTNeXt-50 and CoTNeXt-101 can achieve 82.1% and 83.2%, making the absolute improvement over the best competitor ResNeSt-50 or ResNeSt-101/LambdaResNet-10 by 1.0% and 0.9%, respectively. More specifically, the attention-based backbones (BoTNet-S1-50 and BoTNet-S1-59) exhibit better performances than ResNet-50 and ResNet-101, by replacing the final three 33 convolutions in ResNet with global self-attention layers. LambdaResNet-101 further boosts up the performances by leveraging the computationally efficient global self-attention layers (i.e., Lambda layer) to replace the convolutional layers. Nevertheless, LambdaResNet-101 is inferior to CoTNeXt-101 which capitalizes on the contextual information among input keys to guide self-attention learning. Even under the heavy setting with deeper networks, our SE-CoTNetD-152 (320) still manages to outperform the superior backbones of BoTNet-S1-128 (320) and EfficientNet-B7 (600), sharing the similar (even smaller) FLOPs with BoTNet-S1-128 (320).
Inference Time vs. Accuracy. Here we evaluate our CoTNet models with regard to both inference time and top-1 accuracy for image recognition task. Figure 3 and Figure 4 show the inference time-accuracy curve under both default and advanced training setups for our CoTNet and the state-of-the-art vision backbones. As shown in the two figures, we can see that our CoTNet models consistently obtain better top-1 accuracy with less inference time than other vision backbones across both training setups. In a word, our CoTNet models seek better inference time-accuracy trade-offs than existing vision backbones. More remarkably, compared to the high-quality backbone of EfficientNet-B6, our SE-CoTNetD-152 (320) achieves 0.6% higher top-1 accuracy, while runs 2.75 faster at inference.
Ablation Study. In this section, we investigate how each design in our CoT block influences the overall performance of CoTNet-50. In CoT block, we first mine the static context among keys via a 33 convolution. Conditioned on the concatenation of query and contextualized key, we can also obtain the dynamic context via self-attention. CoT block dynamically fuses the static and dynamic contexts as the final outputs. Here we include one variant of CoT block by directly summating the two kinds of contexts, named as Linear Fusion.
Table 5 details the performances across different ways on the exploration of contextual information in CoTNet-50 backbone. Solely using static context (Static Context) for image recognition achieves 77.1% top-1 accuracy, which can be interpreted as one kind of ConvNet without self-attention. Next, by directly exploiting the dynamic context via self-attention, Dynamic Context exhibits better performance. The linear fusion of static and dynamic contexts leads to a boost of 78.7%, which basically validates the complementarity of the two contexts. CoT block is further benefited from the dynamic fusion via attention, and the top-1 accuracy of CoT finally reaches 79.2%.
Effect of Replacement Settings. In order to show the relationship between performance and the number of stages replaced with our CoT blocks, we progressively replace the stages with our CoT blocks in ResNet-50 backbone (res2res3res4res5), and compare the performances. The results shown in Table 6 indicate that increasing the number of stages replaced with CoT blocks can generally lead to performance improvement, and meanwhile the parameter number & FLOPs are slightly decreased. When taking a close look on the throughputs and accuracies of different replacement settings, the replacement of CoT blocks in the last two stages (res4 and res5) contributes to the most performance boost. The additional replacement of CoT blocks in the fist stages (res1 and res2) can only lead to a marginal performance improvement (0.2% top-1 accuracy in total), while requiring 1.34 inference time. Therefore, in order to seek a better speed-accuracy trade-off, we follow and construct an upgraded version of our CoTNet, named SE-CoTNetD-50, where only the 33 convolutions in the res4 and res5 stages are replaced with CoT blocks under SE-ResNetD-50 backbone. Note that the SE-ResNetD-50 backbone is a variant of ResNet-50 with two widely adopted architecture changes (ResNet-D and Squeeze-and-Excitation in all bottleneck blocks ). As shown in Table 6, compared to the SE-ResNetD-50 counterpart, our SE-CoTNetD-50 achieves better performances at a virtually negligible decrease in throughput.
2 Object Detection
Setup. We next evaluate the pre-trained CoTNet for the downstream task of object detection on COCO dataset. For this task, we adopt Faster-RCNN and Cascade-RCNN as the base object detectors, and directly replace the vanilla ResNet backbone with our CoTNet. Following the standard setting in , we train all models on COCO-2017 training set (118K images) and evaluate them on COCO-2017 validation set (5K images). The standard AP metric of single scale is adopted for evaluation. During training, for each input image, the size of the shorter side is sampled from the range of . All models are trained with FPN and synchronized batch normalization . We utilize the 1x learning rate schedule for training. For fair comparison with other vision backbones in this task, we set all the hyperparameters and detection heads as in .
Performance Comparison. Table 7 summarizes the performance comparisons on COCO dataset for object detection with Faster-RCNN and Cascade-RCNN in different pre-trained backbones. We group the vision backbones with same network depth (50-layer/101-layer). From observation, our pre-trained CoTNet models (CoTNet-50/101 and CoTNeXt-50/101) exhibit a clear performance boost against the ConvNets backbones (ResNet-50/101 and ResNeSt-50/101) for each network depth across all IoU thresholds and object sizes. The results basically demonstrate the advantage of integrating self-attention learning with contextual information mining in CoTNet, even when transferred to the downstream task of object detection.
3 Instance Segmentation
Setup. Here we evaluate the pre-trained CoTNet in another downstream task of instance segmentation on COCO dataset. This task goes beyond the box-level understanding in object detection by additionally predicting the object mask for each detected object, pursuing the pixel-level understanding of visual content. Specifically, Mask-RCNN and Cascade-Mask-RCNN are utilized as the base models for instance segmentation. In the experiments, we replace the vanilla ResNet backbone in Mask-RCNN with our CoTNet. Similarly, all models are trained with FPN and synchronized batch normalization. We adopt the 1x learning rate schedule during training, and all the other hyperparameters are set as in . For evaluation, we report the standard COCO metrics including both bounding box and mask AP ( and ).
Performance Comparison. Table 8 details the performances of Mask-RCNN with different pre-trained vision backbones for the downstream task of instance segmentation on COCO dataset. Similar to the observations for object detection downstream task, our pre-trained CoTNet models yields consistent gains against both ConvNets backbones (ResNet-50/101 and ResNeSt-50/101) and attention-based model (BoTNet-50/101) over the most IoU thresholds. This generally highlights the generalizability of our CoTNet in the challenging instance segmentation task. In particular, BoTNet-50 achieves better performances than the best ConvNets (ResNeSt-50). This might attribute to the additional modeling of global self-attention in BoTNet plus the more advanced fine-tuning setup with larger input size (10241024) and longer training epochs (36). However, by uniquely exploiting the contextual information among neighbor keys for self-attention learning, our CoTNet-50 manages to lead the performance boosts over the most metrics, even when fine-tuned with smaller input size and less epoches (12). The results again confirm the merit of simultaneously performing context mining and self-attention learning in our CoTNet for visual representation learning.
Conclusions
In this work, we propose a new Transformer-style architecture, termed Contextual Transformer (CoT) block, which exploits the contextual information among input keys to guide self-attention learning. CoT block first captures the static context among neighbor keys, which is further leveraged to trigger self-attention that mines the dynamic context. Such way elegantly unifies context mining and self-attention learning into a single architecture, thereby strengthening the capacity of visual representation. Our CoT block can readily replace standard convolutions in existing ResNet architectures, meanwhile retaining the favorable parameter budget. To verify our claim, we construct Contextual Transformer Networks (CoTNet) by replacing the 33 convolutions in ResNet architectures (e.g., ResNet or ResNeXt). The CoTNet architectures learnt on ImageNet validate our proposal and analysis. Experiments conducted on COCO in the context of object detection and instance segmentation also demonstrate the generalization of the visual representation pre-trained by our CoTNet.