Augmented Shortcuts for Vision Transformers
Yehui Tang, Kai Han, Chang Xu, An Xiao, Yiping Deng, Chao Xu, Yunhe Wang
Introduction
Originating from the natural language processing filed, transformer models have recently made great progress in various computer vision tasks such as image classification , object detection and image processing . Wherein, the ViT model divides the input images as visual sequences and obtains an 88.36% top-1 accuracy which is competitive to the SOTA convolutional neural network (CNN) models (e.g., EfficientNet ). Compared to CNNs which are usually customized for vision tasks with prior knowledge (e.g., translation equivalence and locality), vision transformers introduce less inductive bias and have a larger potential to achieve better performance on different visual tasks. Besides, considering the high performance of transformers in various fields (e.g., natural language processing and computer vision), we may only need to support transformer for processing different tasks which can significantly simplify the software/hardware design.
Besides the self-attention layers in vision transformers, a shortcut is often included to directly connect multiple layers with identity projection . The introduce of shortcut is motivated by the architecture designs in CNNs, and has been demonstrated to be beneficial for a stable convergence and better performance. There is a series of works on interpreting and understanding the role of shortcut. For example, Balduzzi et al. analyze the gradients of deep networks and show that shortcut connection can effectively alleviate the problem of gradient vanishing and exploding. Veit et al. reckon ResNet as the ensemble of a collection of paths with different lengths, and the gradient vanishing problem is addressed by short paths. From a theoretical perspective, Liu et al. prove that shortcut connection can avoid network parameters trapped by spurious local optimum and help them converge to a global optimum. Apart from CNNs, shortcut connection is consequently widely-used in other deep neural networks such as transformer , LSTM and AlphaGo Zero .
In vision transformers, the shortcut connection bypasses the multihead self-attention (MSA) and multilayer perceptron (MLP) modules, which plays a critical role towards the success of vision transformers. A transformer without shortcut suffer extremely low performance (Table 1). Empirically, removing the shortcut results in features from different patches becoming indistinguishable as the network going deeper (shown in Figure 4(a)), and such features have limited representation capacity for the downstream prediction. We name this phenomenon as feature collapse. Fortunately, adding shortcut in transformers can alleviate the phenomenon (Figure 4(b)) and make the generated features diverse. However, the conventional shortcut simply copies the input feature to the output, limiting its ability for enhancing the feature’s diversity.
In this paper, we introduce a novel augmented shortcut scheme for improving feature diversity in vision transformers (Figure 2). Besides the conventional identity shortcut, we propose to parallel the MSA module with multiple parameterized shortcuts, which provide more alternative paths to bypass the attention mechanism. In particular, an augmented shortcut connection is constructed as the sequence of a linear projection with learnable parameters and a nonlinear activation function. We theoretically demonstrate that the transformer model equipped with augmented shortcuts can avoid feature collapse and produce more diverse features. For the efficiency reason, we further replace the original dense matrices with block-circulant matrices, which have lower computational complexity in the Fourier frequency domain while high representation ability in the spatial domain. Owing to the compactness of circulant projection, the parameters and computational costs introduced by the augmented shortcuts are negligible compared to those of the MSA and MLP modules. We empirically evaluate the effectiveness of augmented shortcut with the ViT model and its SOTA variants (PVT , T2T ) on ImageNet dataset. Equipped with our augmented shortcut, the performances (top-1 accuracies) of these models can be enhanced by about 1% with comparable computational costs (e.g., Figure 2).
Preliminaries and Motivation
Besides MSA and MLP, the shortcut connection also plays a vital role to achieve high performance. Empirically, removing shortcut severely harms the accuracies of ViT models (as shown in Table 1). To explore the reason behind, we propose to analyze the intermediate features in the ViT-like models. Specially, we calculate the cosine similarity between different patch features and show the similarity matrices of low (Layer 1), middle (Layer 6) and deep layers (Layer 12) in Figure 4. Features from different patches in a layer quickly become indistinguishable as the network depth increasing. We call this phenomenon feature collapse, which greatly restrict the representation capacity and then prevent high performance.
Given a model stacked by the MSA modules, the diversity of feature in the -th layer can be bounded by that of input data , i.e.,
where is number of heads, is feature dimension and is a constant related to the norms of weight matrices in the MSA module.
and are usually smaller than 1, so the feature diversity will decrease rapidly as the network depth increases . We also empirically show how varies w.r.t. the network depth in the ViT models (Figure 4), and the empirical results are accordant to the Theorem 1.
Fortunately, adding a shortcut connection parallel to the MSA module can empirically preserve the feature diversity especially in deep layers (see Figure 4(b) and Figure 4). The MSA module with shortcut can be formulated as:
where the identity projection (i.e., ) is parallel to the MSA module. Intuitively, the shortcut connection bypasses the MSA module and provides another alternative path, where features can be directly delivered to the next layer without interference of other patches. Adding the shortcut connections can also theoretically improve the bound of feature diversity (as discussed in Section 3.1). The success of shortcut shows that bypassing the attention layers with extra paths is an effective way to enhance feature diversity and improve the performance of transformer-like models.
However, in general vision transformer models (e.g., ViT , PVT , T2T ), there is only a single shortcut connection with identity projection for each MSA module, which only copies the input features to the outputs. This simple formulation may not have enough representation capacity to improve the feature diversity maximally. In the following chapters, we aim to refine the existing shortcut connections in vision transformers and explore efficient but powerful augmented shortcuts to produce visual features with higher diversity.
Approach
We propose augmented shortcuts to alleviate the feature collapse by paralleling the original identity shortcut with more parameterized projections. The MSA module equipped with augmented shortcuts can be formulated as:
where is the -th augmented shortcut connection of the -th layer and denotes its parameters. Besides the original shortcut, the augmented shortcuts provide more alternative paths to bypass the attention mechanism. Different from the identity projection directly copying the input patches to the corresponding outputs, the parameterized projection can transform input features into another feature space. Actually, projections will make different transformations on the input feature as long as their weight matrices are different, and thus paralleling more augmented shortcuts has potential to enrich the feature space.
A simple formulation for is the sequence of a linear projection and an activation function i.e.,
Recall that in a transformer-like model without shortcut, the upper bound of feature diversity decreases dramatically as the increase of network depth(Theorem 1). In the following, we analyze how the diversity changes w.r.t. the layer in the model stacked by the AugMSA modules, and we has the follow theorem.
Given a model stacked by the AugMSA modules, the diversity of feature in the -th layer can be bounded by that of input data , i.e.,
where . is the weight matrix in the -th augmented shortcut of the -th layer, and is the Lipschitz constant of activation function .
Compared with Theorem 1, the augmented shortcuts introduce an extra term , which will increase doubly exponentially as is usually larger than 1. This tends to suppress the diversity decay incurred by attention mechanism. The term () is determined by the norms of weight matrices of the augmented shortcuts in the -th layer, and then bound of diversity in the -th layer can be affected by all the augmented shortcuts in the previous layers. For the ShortcutMSA module (Eq. 6) with only a identity shortcut, we have . Adding more augmented shortcuts can increase the magnitude of , which further improves the bound. Detailed proof for Theorem 2 is represented in the supplemental material.
Considering that shortcut connections exist in both MSA and MLP modules, the proposed augmented shortcuts can also be embedded into MLP similarly, i.e.,
where is the input feature of the MLP module in the -th layer and denotes the parameters in augmented shortcuts. Paralleling the MLP module with the augmented shortcuts can further improve the diversity, which is analyzed detailedly in the supplemental material. Stacking the AugMSA and AugMLP modules constructs the Aug-ViT model, whose feature has stronger diversity as shown in Figure 4(c) and Figure 4.
2 Efficient Implementation via Circulant Projection
where is the size of sub-matrices and . Each sub-matrix is a circulant matrix generated by circulating the elements in a -dimension vector , i.e.,
Complexity Analysis. For a matrix split into multiple sub-matrices, the numbers of parameters is only , which has linear complexity with matrix size . Note that the learnable parameters can be stored as its Fourier form, and both the FFT and IFFT operations in Eq. 13 are conducted one time. Thus the computational cost (i.e., FLOPs) of Eq. 13 is about (), where comes from the FFT and IFFT transformation and is consumed by the element-wisely operation in the Fourier domain The constants are approximated by considering the relation between complex operation and real operation, as well as the symmetry property . The size of matrix in a transformer is usually large while the number of sub-matrices is small, and thus the parameters and computation costs incurred by the augmented shortcuts can be negligible. For example, in the ViT-S model with , is set to 4 in our implementation. The augmented shortcuts only adds 0.07 M parameters and 0.08 G FLOPs, which is negligible considering the ViT-S model has 22.1 M parameters and 4.6 G FLOPs.
Experiments
In this section, we conduct extensive experiments to demonstrate the effectiveness of the proposed augmented shortcuts. The vision transformer and its SOTA variants are equipped with the augmented shortcuts for performance improving. We first compare the performances of different models on the ImageNet dataset for the image classification task. Then ablation studies are conducted to analyze the algorithm. To validate its generalization ability, we further test the proposed Aug-ViT model on the object detection and transfer learning tasks.
Dataset. ImageNet (ILSVRC-2012) dataset contains 1.3 M training images and 50k validation images from 1000 classes, which is a widely used image classification benchmark.
Implementation details. We use the same training strategy of DeiT for a fair comparison. Specifically, the model is trained with AdamW optimizer for 300 epochs with batchsize 1024. The learning rate is initialized to and then decayed with the cosine schedule. Label smoothing , DropPath and repeated augmentation are also implemented following DeiT . The data augmentation strategy contains Rand-Augment , Mixup and CutMix . Besides the original shortcuts, two augmented shortcuts are added, where the hyper-parameter for partitioning matrices is set to 4 empirically. The models are trained from scratch on ImageNet and no extra data are used. All experiments are conducted with PyTorch on NVIDIA V100 GPUs.
Backbones. We apply the augmented shortcuts on multiple vision transformer models. ViT is the typical transformer model for the vision tasks, which splits an image to multiple patches. Deit adopts the same architecture with ViT but improves the training strategy for better performance. T2T and PVT are two recently proposed SOTA variants of ViT. T2T improves the process of producing patches by considering the structured information in images. PVT designs a pyramid-like structure by partitioning the model into multiple stages. In Table 2, ‘ViT(DeiT)-S’ ,‘ViT(DeiT)-B’ and ‘PVT-S’, ‘PVT-B’ denote the DeiT and PVT models with different size. ‘T2T-14’, ‘T2T-19’ and ‘T2T-24’ are the T2T models with different depths.
Experimental Results. The validation accuracies of different models on ImageNet are shown in Table 2. Note that DeiT adopts the same model architecture with ViT but achieves higher performance by adjusting the training strategy, which is used as the baseline model. We firstly equip the ViT models with the augmented shortcuts and get Aug-ViT, which show large superiority to the plain counterparts, i.e., more than 1% accuracy improvement without noticeable parameter and computational complexity increasing. For example, the proposed ’Aug-ViT-S’ achieves 80.9% top-1 accuracy with 4.6G FLOPs, which suppresses the baseline (’ViT(DeiT)-S’ with 4.6G FLOPs) by 1.2% top-1 accuracy while the computational cost is barely changed. For a large model with higher input image resolution (e.g., ViT-B with input resolution ) and high performance, equipping it with the augmented shortcuts can still further improve the performance (e.g., 83.1% 84.2%).
Besides the typical ViT model, the augmented shortcuts can be embedded into multiple vision transformers flexibly and improve their performance as well. For example, equipping the PVT-M with augmented shortcut can improve its accuracy form 81.2% to 82.3%. For the T2T-14 model, the performance improvement is even more obvious, i.e., 1.3% accuracy improvement from 82.3% to 83.6%.
Performance Improvement w.r.t. Network Depth. It is interesting to see that the augmented shortcuts improve the performance of deeper models more obviously. ‘T2T-14’, ‘T2T-19’ and ‘T2T-24’ compose of the same blocks but have different depths. For the baseline, increasing the depth of T2T model from 14 to 24, the accuracy is only improved by 0.7% (from 81.5% to 82.3%). While with the augmented shortcuts, the Aug-T2T model with 24 layers can achieve 83.6%, which achieves more obvious performance improvement than models with 14 layers. We conjecture that it is because deeper models tend to suffer more serve feature collapse suffer more serve feature collapse as features from different patches are aggregated with more attention layers.
2 Ablation Studies
To better understand the proposed augmented shortcuts for vision transformers, we conduct extensive experiments to investigate the impact of each component. All the ablation experiments are conducted based on ViT(DeiT)-S model on the ImageNet dataset.
The number of augmented shortcuts. The performance varies w.r.t. the number of augmented shortcuts as shown in Table 4. Besides the original identity shortcut, adding only one augmented shortcut can significantly improves the performance of the ViT model (e.g., 0.8% top-1 accuracy improvement compared to the baseline). Further increasing the number of augmented shortcuts will further improve the performance, and the improvement margin will be saturated gradually. We empirically find that two augmented shortcuts are enough to achieve obvious performance improvement.
Location for implementing the augmented shortcuts. As discussed before, the augmented shortcuts can be paralleled with both MSA and MLP modules to increase the feature diversity. Table 4 shows how the implementation location affects the final performance. Paralleling MSA with the augmented shortcuts significantly improves the performance (e.g., 0.8% top-1 accuracy), which we attribute it to the increasing of feature diversity. Enhancing MLP module also bring the performance improvement, which is accordant to our analysis in supplementary materials. Combining them together can achieve the highest performance (1.1% accuracy improvement), which we adopt in our implementation.
Efficiency of the block-circulant projection. In the block-wisely circulant projection, the hyper-parameter controls the number of sub-matrices partitioned by original matrix . Table 6 shows how the performance varies w.r.t. the parameter . A larger implies the matrix will be partitioned into more circulant matrix with smaller sizes, which brings more parameters and higher computational cost, as well as the performance improvement. The unstructured projection can also be used as augmented shortcuts and brings performance improvement, but it incurs obvious increasing of parameters and computational cost. Using the block-circulant projection with can achieve very similar performance with the unstructured projection but has much few parameters (e.g., 0.07M vs.7.1M), implying that the block-circulant projection is an efficient and effective formulation to transform features in the augmented shortcuts.
Formulation of the augmented shortcuts. The augmented shortcuts are implemented sequentially with the block-circulant projection and activation function (e.g., GeLU). Table 6 shows the impact of each component. Only using the activation function introduces no learnable parameters, which only achieves similar accuracy with the baseline model, implying that the learnable parameters are vital to produce diverse features. If only the linear circulant projection is kept, multiple augmented shortcut can be merged to a single one. Swapping the sequential order of the circulant projection and activation function has negligible influence on the final performance.
Feature Visualization. We intuitively show the features of different models in Figure 5, From top to bottom are features in low, middle and deep layers of the ViT-Small model. The input image is scaled to for better visualization and the patch embeddings are reshaped to their spatial positions to construct the features maps. Without shortcut, the feature maps in deep layers conveys no effective information (Figure 5 (a)), and adding a shortcut connection make the feature maps informative ((b)). Compared with them, the features in the Aug-ViT model are further enriched, especially for the deep layers.
3 Object Detection with Pure Transformer
We also validate the effectiveness of the augmented shortcuts on the objection detection task. A pure transformer detector can be constructed by combining the vision transformer backbone and the DETR head, and we equip the backbones with the augmented shortcuts. For a fair comparison, we follow the training strategy in PVT and fine-tune the models for 50 epochs on the COCO train2017 dataset. Random flip and random scale are used as the data augmentation strategy. The results on COCO val2017 are shown in Table 7. The detector equipped with augmented shortcut achieve better performance than the base model. For the DeiT-S model with 33.9 AP, the augmented shortcut can improve 1.8% AP and achieve 35.7%.
4 Transfer Learning
To validate the generalization ability of vision transformers equipped with the projection, we conduct experiments on the transfer learning tasks. Specifically, the models trained on ImageNet is further fine-tuned on the downstream tasks, containing superordinate-level image recognition dataset (CIFAR-10 , CIFAR-100 ) and fine-grained image recognition dataset( Oxford 102 Flowers and Oxford-IIIT Pets ). Following , we fine-tune the models on images with resolution , and use the same fine-tuning strategy as . Table 8 shows the performances of different models on the downstream tasks, and the model equipped with the augmented shortcut always achieves higher accuracies than the baseline on different tasks.
Conclusion
We presented augmented shortcuts for resolving the feature collapse issue in vision transformers. The augmented shortcuts are parallel with the original identity shortcuts, and each connection has its own learnable parameters to make diverse transformations on the input features. Efficient circulant projections are used to implement the augmented shortcuts, whose memory and computational cost are negligible compared with other components in the vision transformer. Similar to the widely used identity shortcuts, the augmented shortcuts do not depend on the specific architecture design either, which can be flexibly embedded into various variants of vision transformers (e.g., ViT , T2T , PVT ) for enhancing their performance on different tasks such as image classification, object detection and transfer learning. In the future, we plan to research designing deeper vision transformers with the help of augmented shortcuts.