Vicinity Vision Transformer
Weixuan Sun, Zhen Qin, Hui Deng, Jianyuan Wang, Yi Zhang, Kaihao Zhang, Nick Barnes, Stan Birchfield, Lingpeng Kong, Yiran Zhong
Introduction
Recent years have witnessed the success of the transformer structure in natural language processing and computer vision . However, transformers inherently suffer from quadratic computational complexity and a quadratic memory footprint. As a result, vision transformer networks have to adopt patch-wise image tokenization to reduce the sequence length. Despite the temporary relief provided by such a tokenization method, the quadratic complexity problem still exists. This limitation prohibits vision transformers from handling high-resolution images or fine-grained image patches.
Linear attention is a promising direction to solve this issue. This group of methods reorders the self-attention mechanism of transformers with Kernelization methods to reduce the quadratic complexity to linear . Numerous methods have been proposed for attention decomposition, such as approximating the softmax , and finding a new similarity metric . However, most of these methods are only verified on NLP tasks and suffer from a crucial performance drop in computer vision , when compared with conventional softmax attention.
We investigate this issue and point out that existing linear attention methods ignore an inductive bias in vision, i.e., 2D locality. Our intuition is based on the fact that convolutional neural networks have dominated computer vision tasks since the rise of deep networks. Most of them have such locality bias, that is, 2D neighbouring regions are more likely to be highly related. Recent transformer-based methods also adopt such an assumption and obtain improved results by attaching convolution-based 2D locality bias . Additionally, some efficient transformer backbones use window attention , neighbourhood attention , or deformable attention to enable lower complexity and 2D locality, but they still suffer limited receptive field and quadratic complexity within the sampled tokens. Therefore we hypothesize that the 2D locality bias is an important property and should be incorporated into linear transformers.
In this paper, we present Vicinity Attention, a new linear attention method that effectively enforces 2D locality. Our locality re-weighting mechanism is inspired by the recently proposed cosFormer , which assumes locality bias in language and uses a cosine re-weighting mechanism for 1D NLP tasks. However, directly applying the cosFormer to vision tasks will lead to unsatisfactory results since the 1D locality enforces stronger connections only on the 1D tokenized neighbouring image patches, so the cosFormer is not compatible with 2D distance. In that case, it will assign less weight to vertically connected patches as they are far away when we tokenized the patches to 1D tokens as shown in Fig. 3a. To solve this issue, we propose a 2D variant of cosFormer in this work. Specifically, we propose a 2D Manhattan distance decomposition to encode relative positions, and integrate it with the cosine function to encourage visual tokens to have a stronger connection to their neighbors in 2D (Fig. 3b-d). Since this distancing mechanism is decomposable in two directions, it can be seamlessly applied to linear attention.
Compared with vanilla transformer-based methods, linear attention complexity grows linearly with respect to sequence length but quadratically with feature dimension, which becomes a new computational bottleneck in our Vicinity Attention. To address this concern so as to further reduce computation, we propose a novel Vicinity Attention Block. First, we propose Feature Reduction Attention (FRA) to reduce the input feature dimension by half, so that the overall theoretical complexity can be reduced by a factor of four. Then a Feature Preserving Connection (FPC) is added to retrieve the original feature distribution and strengthen the representational ability. Such a block structure is seamlessly integrated with our linear Vicinity Attention and we experimentally validate that our Vicinity Vision Block further reduces computational complexity without degenerating the accuracy.
Finally, we build a general-purpose linear vision backbone, termed Vicinity Vision Transformer (VVT). Since our linear Vicinity Attention has a better efficiency advantage over vanilla self-attention, which enables feature maps with higher resolutions. We build VVT in a pyramid structure, which starts from high-resolution image patches and progressively shrinks to adapt to different vision tasks with multi-scale outputs.
Fig. 1 a compares our method with current state-of-the-art vision backbones on the ImageNet-1k benchmark. Our method outperforms all the competitors with only half the parameters. We also provide the growth rate of GFLOPs with different input resolutions for these methods in Fig. 1 b. Given its linear complexity in token numbers, VVT can efficiently process images with much larger resolutions. Further, our experiments validate that VVT does not compromise accuracy and achieves superior results over transformer-based methods as well as convolution-based methods of comparable model sizes.
Our main contributions are as follows: (1) We introduce a linear self-attention mechanism for vision called Vicinity Attention, which introduces a Manhattan distance-based 2D locality to linear vision transformers. (2) To further reduce computational complexity by targeting linear attention, we propose a novel attention block called the Vicinity Attention Block. It contains a feature reduction attention (FRA) to improve the efficiency and a feature preserving connection (FPC) to retain the feature extraction ability. (3) We correspondingly build the Vicinity Vision Transformer (VVT), which serves as a general-purpose vision backbone and can be easily applied to various vision tasks. Extensive experiments validate the effectiveness of VVT on various computer vision benchmarks.
Related Work
In computer vision tasks, CNN networks have achieved great successes, while recently transformers have gained strong emerging interest. In this section, we mainly discuss vision networks using self-attention.
Some works adopt self-attention mechanisms to replace some or all convolution layers in the CNN networks for image recognition. To further leverage the power of self-attention, proposes a non-local operation and adds it within ResNet , the non-local block can capture long-range dependencies and lead to improvements on vision tasks such as video classification, object detection, segmentation, and pose estimation. Recently, pure transformer-based vision networks were introduced. ViT splits images into local patches. Then it projects and flattens patches into embedding sequences, subsequently, it employs a pure transformer structure for image classification.
Variations of ViT like CVT , Swin , PVT , Twins and T2T were proposed. CVT introduces Convolutional Token Embedding and Convolutional Projection for Attention to include desirable properties of CNNs into ViT to improve both performance and efficiency. The Swin Transformer introduces non-overlapping windows and applies self-attention in each window, then the window partitions are shifted between adjacent layers. Swin transformer reduces computational complexity and is suitable for different vision tasks. PVT introduces a pure transformer network with a pyramid structure, and it further proposes spatial reduction attention which computes self-attention in reduced sequence length to save computation. Twins combines locally-grouped self-attention from Swin and sub-sampled attention from , which decreases the computational cost and enhances communications between sub-windows. DAT builds a general vision transformer backbone with deformable attention. More recently, adopts the quad-tree algorithm which progressively ignores less related image patches to improve efficiency. However, none of the above methods are able to directly process full-length self-attention on high-resolution inputs, requiring locally grouping or sub-sampling. Contrarily, our method can directly calculate self-attention on high-resolution input.
2 Efficient Transformers
Various works have been proposed to address the computational complexity problems of transformers in both computer vision and natural language processing. In addition to vision backbones such as CVT , Swin , PVT , and Twins introduced in the previous section, in this section we introduce other more efficient transformer algorithms.
Existing efficient transformer methods can be generally grouped into two categories: pattern-based and kernel-based. Pattern-based methods sparsify the attention matrix with handcrafted or learnable patterns. reduces the complexity by applying a combination of a strided pattern and a local pattern to the standard attention matrix. Beyond fixed diagonal windows and global windows, Longformer also extends sliding windows with dilation to enlarge the receptive field. Instead of fixed patterns, Reformer and adopt locally sensitive hashing to group tokens into different buckets. On the other hand, kernel-based methods aim to replace softmax self-attention with approximations or decomposable functions, which change the order of scale dot product calculation and reduce the complexity of self-attention from quadratic to linear. proposes a cosine re-weighting function to enforce locality in a 1D sequence and achieves linear complexity. Nevertheless, the aforementioned kernel-based methods only consider the linearization in a 1D sequence for natural language processing tasks, their performances on 2D vision tasks are not satisfactory. In contrast, we aim at enforcing 2D locality in linear complexity.
Two works that are similar to ours are and . applies softmax on and respectively, then changes the dot product order to access a linear complexity. However, it is not validated on a large-scale classification benchmark as a backbone network. SOFT uses a Gaussian kernel function to replace the dot-product similarity in self attention. It assumes that the similarity is symmetrical, which may not hold true in practice. Further, it does not consider the 2D locality mechanism. In this paper, we propose a linear self-attention that facilitates a 2D locality mechanism, and we validate the proposed method on various vision tasks.
Preliminary
where is the self-attention module and is a feed-forward module.
Linearization of self-attention aims to reduce the quadratic theoretical computation complexity to linear. It can be achieved by picking a decomposable similarity function to satisfy
where is a kernel function. Given such a kernel, we can write the outputs of the self-attention module as follows:
so that the operation is converted to an one. Its computational complexity grows linearly with respect to the sequence length . A number of linear attention approaches have been proposed for NLP tasks, which use different kernel functions to replace the quadratic softmax, as discussed in the Related Work. However, directly applying existing linear attention to vision suffers from a performance drop. Notably, all of the aforementioned linear attention approaches and existing linear vision transformers do not consider 2D locality in vision. To address this problem, we propose a linear self-attention that is aware of 2D position.
Locality is a widely used assumption in computer vision , i.e., neighbouring pixels should have a higher possibility to belong to the same object than distant pixels. In convolution-based networks, this assumption is inherently coupled into each layer throughout the whole model via convolutional kernels . However, it does not hold for standard transformer-based networks due to the self-attention mechanism . Several works show that with the same number of parameters, transformer-based networks cannot match the performance of the CNN counterparts, possibly due to the lack of locality bias . Recent state-of-the-art methods partially introduce locality bias to vision transformers at the architectural level. directly combine convolution with transformers. Some efficient transformer backbones use window attention , neighbourhood attention , or deformable attention to achieve lower complexity and better performance. These methods enable 2D locality, however, they still suffer limited receptive field and quadratic complexity within the sampled tokens.
In this paper, for the first time, we introduce locality bias into the linear self-attention. It can be smoothly integrated into existing vision transformer architectures for better performance. In fact, our method achieves better performance than previous state-of-the-art vision transformers in various computer vision benchmarks.
Our Method
In this section, we start from presenting the details of Vicinity Attention mechanism in Sec. 4.1 which enforces 2D locality in linear attention. Then in Sec. 4.2, we introduce a new attention structure named Vicinity Attention Block to address the computational bottleneck for linear attention targeting linear attention. In Sec. 4.3 we introduce the overall architecture of VVT and show the structure illustration in Fig. 5. Targeting general vision tasks, we integrate our Vicinity Attention into a four-stage pyramid structure. Each stage consists of a patch embedding module followed by several Vicinity Vision Blocks and feed forward residual blocks, Image inputs are progressively down-sampled to generate multi-scale outputs. We provide a family of VVT variants and detail their configurations in Table. I.
Enforcing 2D Locality Bias Locality bias has been discussed in language tasks with 1D sequence distance encoding, but not considered by existing linear attention in vision. In vision transformers, the embedding sequence is flattened from a 2D mask, hence it is essential to consider token positions in 2D before being flattened. To enforce the locality bias in linear transformers, we need a kernel function that can 1) put more emphasis on 2D neighboring patches and 2) can be decomposed with Eq. (2).
Given two tokens , from and respectively, the positions of these two tokens on the 2D feature maps before flattening are:
where is width of the feature map, denotes the row index, and denotes the column index. Following Eq. (2), we can define a re-weighted attention with a distance function to enforce the locality bias between two tokens as:
where produces the weight according to the distance so nearby patches can be emphasized. A naive choice might be directly using the Euclidean distance . However, since this term cannot be decomposed into two terms relating to and separately, it cannot be applied to the linear transformers.
Instead, since 2D Manhattan distance decouples relative position in two directions, we propose to use it as the distance function:
However, direct Manhattan distance decoupling is still hindered as the absolute operation cannot be decomposed.
Inspired by , the cosine function has two desirable features: 1) It cancels the absolute operation in Eq. (8) and applies non-linear emphasis on the nearby tokens; and 2) It can be decomposed into two terms relating to and separately, which fulfills the linear complexity. Thus, we propose to bind Manhattan distance and cosine function to achieve linear attention with 2D locality. First, given a feature map with size by , following Eq. (6) we redefine the 2D positions as:
Then, using the Manhattan distance, the self-attention calculation can be decomposed into four terms:
where .
Here the queries and keys are related to their own positions and , and hence we can fulfill Eq. (4) to reorder the dot product in self-attention. Following our locality decomposition, we get:
Our method provides three appealing properties: (1) Linear complexity: cosine re-weighting and Manhattan distance ensure decomposability, thus we avoid calculating . The overall computational complexity is which grows linearly with sequence length. (2) Locality bias: if two tokens are close to each other in 2D feature maps, they are encouraged to have a stronger relationship, as shown in Fig. 3. (3) Global context: it retains a global receptive field as standard self-attention.
2 Vicinity Attention Block
Compared to vanilla attention, the complexity of linear attention grows linearly with sequence length but quadratically with respect to feature embedding size. Thus, the computational bottleneck of linear attention is shifted from the input resolution to feature dimension, whilst it is not considered by existing linear vision transformers. This issue is amplified in Vicinity Attention as our theoretical computational complexity is . To achieve a better computational efficiency in linear attention, we redesign the structure of the multi-head self-attention (MSA) module to reduce the feature dimension without hindering the performance.
In the second preserving step, we propose to preserve the original feature distribution to compensate the feature compression with negligible computational overhead. Inspired by , we add a skip connection called Feature Preserving Connection (FPC) to capture the global context features of input . The FPC consists of an average pooling operation and two linear layers which have a complexity of . FPC retains the original feature distribution and can strengthen the representational ability. Finally, the outputs of FRA and FPC are fused in a residual manner to get final attention output.
In summary, our Vicinity Attention Block proposes to calculate self-attention in a compressed feature dimension while still preserving the original feature space. It reduces the theoretical complexity to . Our experiments validate that it notably reduces computation without degenerating accuracy.
In the following, we discuss the relationships between our VVT and several existing methods.
Relationship to cosFormer. The cosFormer was originally developed for natural language processing and achieves linear complexity in 1D. In this paper we develop a 2D variant of the cosFormer. The cosFormer attention adopts cosine re-weighting and considers 1D locality bias in NLP, but it shows poor results on vision tasks (Table VII). Our Vicinity Attention can be seen as a non-trivial extension of the cosFormer to vision. The combination of Manhattan distance and the cosine function composes a novel 2D positional encoding that enables the capture of long-range visual dependency with 2D locality and linear complexity. In addition, we propose a new vision attention block named the Vicinity Attention Block, which contains FRA and FPC to further reduce the complexity without diminishing performance, The Vicinity Attention Block is then used to build vision backbones with a pyramid structure and can process high-resolution images.
Relationship to existing window attention and neighborhood attention. 2D locality has been introduced into vision transformers by several methods such as window attention , deformable attention and neighborhood attention . A key difference between these and our Vicinity Attention lies in the receptive field. These existing methods apply hard locality, i.e., they consider self-attention only within the sampled tokens (from nearby windows or deformable operations) unsampled tokens are disregarded. Further, complexity within the sampled tokens is still quadratic. Differently, Vicinity Attention has a constant global receptive field via soft locality, i.e., all spatial locations are visible, but spatially nearby ones are emphasized via our linear attention.
Relationship to GCNet. The main difference between the Vicinity Attention Block and GCNet block is the attention generation method, which reflects the different motivations of the two blocks. The GCNet block adds a uniform context vector to every spatial location. This uniform context vector does not vary by query location and is motivated by the observation that different attention maps are mostly similar. Our Vicinity Attention Block performs a transformer operation that is modified to reduce complexity. That is, our attention maps vary by query location using the 2D locality constraint. We also have a channel attention-like connection called FPC which is an instantiation of the squeeze-excitation (SE) connection , but aims to reduce complexity while preserving representation ability. Results in Table VII show the superiority of the Vicinity Attention Block compared to the GCNet block.
3 The Vicinity Vision Transformer
In this section, we introduce the overall architecture of Vicinity Vision Transformer (VVT) as shown in Fig. 5.
The linear complexity of Vicinity Attention enables higher-resolution inputs, so we integrate the Vicinity Attention Block into a progressively shrinking pyramid structure that has four stages to generate feature maps at different scales. Each stage contains a patch embedding layer and multiple stacked Vicinity Transformer Blocks and Feed-forward blocks. In detail, we use a patch size of in the first stage. Given the input image of size , we first divide it into patches. Then we feed the patches into a patch embedding module to obtain a flattened embedding sequence with a size of . Here we adopt the overlapping patch embedding module and Convolutional Feed-Forward proposed by Wang et al. . The embedding sequence is subsequently fed into several successive transformer blocks with layers.
In the second stage, the feature sequence from the first stage is reshaped back to and down-sampled to an embedding sequence of size , and then processed by the transformer blocks of the second stage. We follow the same approach to obtain multi-scale output feature maps of the third and fourth stage with output sizes of and , respectively. Hierarchical feature maps can be easily leveraged to many downstream vision tasks. In Fig. 6, we show qualitative examples of Grad-CAM obtained from different stages of VVT and ViT respectively. As shown, our approach is able to produce fine-grained features whereas the ViT can only capture low-resolution features.
Table I details different architecture variants of VVT from VVT_Tiny to VVT_Large. We empirically choose our network settings following the common principles of vision backbones: (1) spatial resolution is decreased progressively with feature dimension increased. (2) stage 3 has the most of the computational cost. The architecture hyper-parameters of VVT are:
C: the input channel dimension of the attention block.
P: the patch size of the patch embedding.
R: the feature reduction ratio of the Vicinity Attention Block.
H: the number of heads in the Vicinity Attention Block.
E: the feature expansion ratio of the feed forward layer.
Experiments
To verify the effectiveness of our method, we conduct extensive experiments on the CIFAR-100 and ImageNet-1k datasets for image classification, and on the ADE20K dataset for semantic segmentation. Specifically, we first make a comparison with existing state-of-the-art methods, and then give an ablation study over the design of VVT.
ImageNet-1k The image classification experiments are conducted on the ImageNet-1k dataset, which contains 1.28 million training images and 50 thousand validation images from categories. We follow the training hyper-parameters of . In detail, the model is trained using an AdamW optimizer with a weight decay of and a momentum of . We use an initial learning rate of and decrease it by a cosine schedule , with epochs for warming up. We also adopt the same data augmentation strategy as in , including random cropping, random horizontal flipping, etc. All the models are trained on the training set for epochs with a crop size of . We use the top-1 accuracy as the evaluation metric.
The quantitative results are shown in Table II, which cover the popular transformer-based and CNN-based classification networks. For fair comparison, we compare the VVT variants respectively with the networks using similar parameters, split by solid lines. For example, VVT-L performs better than PVTv2-B5 but only uses around 70% parameters of the latter. In particular, our medium-size variant VVT-M surpasses the large models Twins-SVT-L and Swin-B with substantially fewer parameters (51.7% and 45.5% respectively). Moreover, compared to the state-of-the-art CNN-based networks such as RegNet and ResNeXt , VVT shows stronger performance while using a similar number of parameters.
It is worth noting that VVT directly computes self-attention on the high-resolution feature maps, while the competitors such as PVT , Swin Transformer , Twins and SOFT rely on sub-sampling or window self-attention to reduce computational cost. The GFLOPs of VVT could be further reduced if similar subsampling or windowing was used, we validate this in Fig. 8.
CIFAR-10 and CIFAR-100. It is known that vision transformers may suffer critical performance drops on small-scale datasets. To validate the performance of our method on small-scale datasets, we use well-known CIFAR-10 and CIFAR-100 as target datasets. Tables III and IV show our VVT-S and competitors’ results. All models are trained from scratch using an AdamW optimizer with a weight decay of 0.05 and a momentum of 0.9. We use an initial learning rate of and train for 300 epochs. All models are trained and tested on resolution for fair comparison. We see that our VVT outperforms all competing methods on both CIFAR-10 and CIFAR-100. We hypothesize that for small datasets, stronger inductive bias may make the model more data efficient. We can see ViT does not have inductive bias and performs poorly. Other models that have different types of locality bias achieve better results than ViT. Our VVT introduces self-attention with 2D locality and achieves the best results.
2 Semantic Segmentation
We utilize the challenging ADE20K dataset to evaluate our idea in semantic segmentation. It has semantic classes, with , and images for training, validation and testing. With the VVT models (pre-trained on ImageNet-1k) as the backbone, we provide the results using two segmentation methods in Table V and show qualitative samples in Fig. 7. Specifically, we pick Semantic FPN and UPerNet as the semantic segmentation architecture, respectively following the settings of PVT and Swin Transformer . The quantitative results show that our method is consistently superior to the competing CNN-based methods. For example, using Semantic FPN, VVT-S outperforms ResNet50 by 8.9%, and VVT-L is 7.7% better than ResNeXt101-32x4d , using mIoU as the metric. VVT also shows a more favourable segmentation result than the transformer-based competitors such as the Swin Transformer. A similar phenomenon is observed if taking UPerNet as the segmentation method. These results validate that the extracted features of VVT are also valuable for semantic segmentation, which benefits from the locality mechanism together with the global attention.
3 Ablation Study
Computational Overhead. We plot the GFLOPs growth rates with respect to the input image size on the right of Fig. 1. The growth rate of VVT (purple) is substantially lower than other transformer-based methods, and even slightly lower than ResNet . However, theoretical GFLOPs may arguably not reflect real computational overhead. To further demonstrate the computational efficiency of Vicinity Attention, we compare GPU memory footprints of Vicinity Attention against recent efficient vision transformer methods in Fig. 8. All results are obtained under the same settings including feature dimension, number of heads etc. Note that although VVT, SOFT , Performer , Quad and Linformer all have linear complexity, their actual memory consumption may vary according to their specific algorithms.
From Fig. 8, we can make the following observations: (i) Our method consumes relatively more memory than Performer and Linformer, as both methods are designed for 1D sequences and show inferior performance on vision tasks. (ii) Compared to efficient vision transformers, i.e., PVTv2, SOFT and Quadtree, the memory consumption of Vicinity Attention grows notably slower, allowing inputs with much higher resolutions. (iii) The spatial reduction (sr) is a computation reduction technique proposed by PVT . Specifically, it reduces the spatial dimensions of the key and value vectors in the self-attention module with convolution or pooling operations, thus the complexity of the self-attention computation is reduced. PVTv2, SOFT, and Quadtree all adopt the sr strategy to reduce computation, while the standard VVT directly calculates attention on the original sequence without sr. If we also integrate sr, the VVT_sr in Fig. 8 validates that the efficiency can be further improved with a spatial reduction ratio of 4. Finally, in Table. VIII, we show memory footprints of the models with various attention modules that correspond to Table VII. In summary, this ablation study validates that VVT successfully mitigates the computational issue in the vision transformers.
Effect of Locality Constraint. The Vicinity locality mechanism is introduced in the self-attention module of VVT. We show its effectiveness on the CIFAR-100 and the ImageNet-1k datasets in Table VI. Experiments are performed on the VVT-Tiny structure. We compare VVT with a variation without locality (denoted as VVT w/o locality) and the one with 1D locality based on cosFormer (denoted as VVT 1D locality). As shown in the table, our VVT achieves the best performance among all the others on the CIFAR-100 dataset. By comparing the VVT w/o locality and VVT 1D locality, we find that if we wrongly enforce the locality bias, the performance will drop critically. To further demonstrate its effectiveness on large-scale data, we also conduct the same ablation on ImageNet-1k. We observe that the model fails to converge with 1D locality which indicates that the 1D cosFormer cannot be trivially transplanted to 2D vision data. Finally, our VVT outperforms the VVT w/o locality by , showing the efficiency of our linear locality attention.
Furthermore, we show the qualitative examples of Grad-CAM from VVT and the competitors in Fig. 9. As shown, the locality mechanism encourages the tokens to assign higher attention to 2D neighbours, and hence the class-wise activations are more concentrated and accurately located on the target object regions of Grad-CAM. These results validate that the introduction of the 2D locality mechanism helps the VVT models generate more reliable object features, which tend to be beneficial in image classification and downstream tasks.
Comparison with Existing Attention Modules We compare with several existing self-attention methods: cosFormer, Performer, Linformer PVTv2 and vanilla self-attention . For all methods, we adopt the VVT-Tiny network structure and only replace the attention block with competing attention modules for a fair comparison. All experiments are conducted using the same network configuration and training setting. We report results on the CIFAR-100 and the ImageNet1K in Table VII. As shown, on CIFAR-100 our Vicinity attention outperforms both the vanilla attention and all the alternative efficient methods. On ImageNet1K, we observe that cosFormer fails to converge due to the false 1D locality. Compared to other efficient methods including GCNet , Performer , Linformer and PVTv2 , VVT still achieves better classification result. Further, vanilla softmax attention is higher than our VVT but at a cost of substantially higher computational complexity (see Table. VIII), which makes it implausible to extend to any larger models with pyramid structures. In Table. VIII, we display the quantitative memory footprints of the aforementioned models in Table VII. All models share the same network structure except the attention module with a batch size of 16, allowing for a fair comparison of different attention mechanisms. Our VVT has linear complexity and similar memory footprints to existing linear attention methods, i.e., cosFormer , Linformer and Performer , whereas our performances on both datasets are superior. Finally, VVT has better memory efficiency than PVTv2 and vanilla softmax attention .
Analysis of Vicinity Attention Block. In this section, we investigate the proposed Vicinity Attention Block from three perspectives: (1) we show that feature reduction attention (FRA) can significantly reduce computation without harming performance; (2) We validate that feature preserving connection (FPC) is effective to strengthen feature extraction ability; and (3) we report the effectiveness of FRA and FPC on an existing linear attention method. In summary, we demonstrate that FRA effectively reduces computational complexity and FPC preserves representation ability without causing degeneration of performance.
By FRA, we reduce the feature dimension to fulfill the assumption of , which forces the computational complexity to approach linear. In Fig. 10, we ablate the computational cost under different feature reduction (FR) ratios. We can observe an obvious GFLOPs decrease when increasing the FR ratio from 1 to 2 (reducing the feature dimension by half). However, when the FR ratio further increases, the reduction tends to saturate. This is because the computational cost of the self-attention module is already significantly reduced and the remaining model parts dominate the computational overhead. Additionally, we ablate the performance under different FR ratios on the CIFAR-100 dataset. As shown in Table IX. When the FR ratio is small, the model retains a similar performance (i.e., 82.89% vs. 82.92%), indicating that the result is not affected. When the FR ratio increases to a number like 8, a clear performance drop is observed. As a trade-off, we set the FR ratio as 2 for our formal model settings.
In the Vicinity Attention Block, FPC is added to compensate for reduced feature dimensions. Ablation of FPC is reported in Table X. The results indicate that FPC effectively retains the original feature distribution, and improves the feature extraction ability of the Vicinity attention module even when the feature dimension is reduced.
To further demonstrate the effectiveness of Vicinity Attention Block, we integrate FRA and FPC with an existing linear attention method, i.e., Performer. As shown in Table XI, we attach the Performer attention block on VVT-T as a baseline. After integrating FRA and FPC, the computation is reduced while we can still observe an accuracy improvement (i.e., 77.2% vs. 77.3%). It validates that our FRA and FPC yield performance gains in the existing linear attention approach.
Image Classification with Different Input Size Table XII lists the performance of VVT with higher input resolution i.e., 384. Obviously, larger input resolution leads to better top-1 accuracy but requires larger computation. Compared to competing methods such as ViT, Swin and CVT, VVT achieves better results with fewer parameters and a comparable computational overhead.
Analysis of Overlapping Patch Embedding and Convolutional Feed-Forward. In standard VVT, we adopt the overlapping patch embedding module and Convolutional Feed-Forward proposed by Wang et al. . Both modules may also contribute to locality bias and improve results . Table XIII ablates our Vicinity Attention without these two modules, compared to existing methods under the same setting. As shown, we observe that without the contribution of the overlapping patch embedding and convolutional feed-forward, VVT also outperforms competing methods in both CIFAR-10 and ImageNet-1k which shows the efficacy of our attention method.
Throughput Fig. 11 illustrates actual image throughputs of the proposed VVT compared to PVTv2 and Swin transformer. We run the test on one local 2080 Ti GPU with 10 Gb memory. Since the theoretical computational complexity of Vicinity attention is , in low resolution regime where sequence length does not dominate the complexity, VVT shows a relatively slow speed compared to PVTv2 and Swin. When the input resolution is gradually increased, the VVT starts demonstrating better inference speed, while competing methods exhaust GPU memory in the high-resolution regime. In summary, due to the better overall efficiency curve we consider slightly increased time consumption in the low-resolution regime is a price worth paying for our better accuracy.
Conclusion
We have introduced Vicinity Vision Transformer (VVT), a general-purpose vision backbone that produces hierarchical feature representations through a pyramid structure. At its core, we propose Vicinity Attention, which introduces 2D locality into linear self-attention modules. We validate that the locality mechanism improves results and helps generate better class activation. Then, targeting the computational bottleneck of the linear attention, we propose a novel Vicinity Attention Block to further reduce computation. In addition, because VVT has linear complexity, it facilitates processing feature maps or input images with higher resolutions. The proposed VVT has been validated with strong performance on both image classification and semantic segmentation. Future work will be aimed at using neural architecture search to further improve upon architectures specifically designed for linear vision transformers.