VOLO: Vision Outlooker for Visual Recognition
Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, Shuicheng Yan
Introduction
Modeling in visual recognition, which was long dominated by convolutional neural networks (CNNs), has recently been revolutionized by Vision Transformers (ViTs) . Different from CNNs that aggregate and transform features via local and dense convolutional kernels, ViTs directly model long-range dependencies of local patches (a.k.a. tokens) through the self-attention mechanism which is with greater flexibility in modeling visual contents. Despite the remarkable effectiveness on visual recognition , the performance of ViT models still lags behind that of the state-of-the-art CNN models. For instance, as shown in Table 2, the state-of-the-art transformer-based CaiT attains 86.5% top-1 accuracy on ImageNet, which however is still 0.3% lower compared with the 86.8% top-1 accuracy achieved by the CNN-based NFNet-F5 with SAM and augmult .
In this work we try to close such performance gap. We find one major factor limiting ViTs from outperforming CNNs is their low efficacy in encoding fine-level features and contexts into token representations, which are critical for achieving compelling visual recognition performance. Fine-level information can be encoded into tokens by finer-grained image tokenization, which however would lead to a token sequence of greater length that increases quadratically the complexity of the self-attention mechanism of ViTs.
In this work, we present a new simple and light-weight attention mechanism, termed Outlooker, to enrich the token representations with fine level information efficiently. The proposed Outlooker innovates the way of generating attention for token aggregation, and enables the model to efficiently encode fine-level information. In particular, it extrapolates the mechanism of aggregating surrounding tokens from the anchor token feature directly via efficient linear projections, thus getting rid of the expensive dot-product attention computation.
Based on the proposed Outlooker, we present VOLO, a simple yet powerful model architecture for visual recognition. VOLO achieves fine-level token representation encoding and global information aggregation with a two-stage architecture design. Specifically, given an input image of size , before using self-attention to build global dependencies at the coarse level (\eg, ), the VOLO tokenizes the image on smaller-size patches (\eg, ) and employs multiple Outlookers to encode token representations at the fine level (\eg, ). The obtained token representations are more expressive, thus significantly improving the model performance in image classification.
Experiments show that our proposed VOLO performs extremely well in ImageNet classification. Take a VOLO model with 26.6M learnable parameters as an example. It achieves 84.2% top-1 accuracy on ImageNet without using any extra data. Finetuning this model on the input resolution can further increase the accuracy to 85.2%. Moreover, when scaling up the model size to 296M parameters, it can reach a top-1 accuracy of 87.1% on ImageNet, 90.6% on ImageNet-ReaL, and 78.0% on ImageNet-V2, setting new SOTA performance for all the three classification benchmarks.
As depicted in Figure 1, compared to the previous state-of-the-art CNN-based model (NFNet-F6 with SAM ), and the transformer-based model (CaiT-M48 with KD), our best model VOLO-D5 leverages the least amount of learnable parameters but achieves the best accuracy. Moreover, as shown in Table 2, even compared with previous state-of-the-art models using stronger data augmentation and optimization methods (such as SAM and augmult ), our Outlooker still performs the best.
Our VOLO also achieves strong performance on the semantic segmentation task. We run experiments on two widely-used segmentation benchmarks: Cityscapes and ADE20K . Experiments show that our VOLO attains 84.3% mIoU score on the Cityscapes validation set, 0.3% better than the previous state-of-the-art result (by SegFormer-B5 ). On the ADE20K validation set, we achieve 54.3% mIoU score, largely improving the state-of-the-art result (53.5%) by Swin Transformer , which is pretrained on ImageNet-22k.
Method
Our model can be regarded as an architecture with two separate stages. The first stage consists of a stack of Outlookers that generates fine-level token representations. The second stage deploys a sequence of transformer blocks to aggregate global information. At the beginning of each stage, a patch embedding module is used to map the input to token representations with designed shapes.
Here, refers to LayerNorm .
Outlook attention is simple, efficient, and easy to implement. The main insights behind it are: 1) the feature at each spatial location is representative enough to generate attention weights for locally aggregating its neighboring features; 2) the dense and local spatial aggregation can encode fine-level information efficiently.
For each spatial location , outlook attention computes its similarity to all the neighbors within a local window of size centered at . Unlike self-attention that requires a Query-Key matrix multiplication for the computation of the attention (i.e., ), outlook attention simplifies this process via just a reshaping operation.
Dense aggregation Outlook attention aggregates the projected value representations densely. Summing up the different weighted values at the same location from different local windows yields the output
PyTorch-like outlook attention codes are summarized in Algorithm 1. Eqn. (3) and Eqn. (5) correspond to the Unfold and Fold operations, respectively. After outlook attention, a linear layer is often adopted as in self-attention.
1.2 Multi-Head Outlook Attention
1.3 Discussion
Our outlook attention inherits the merits of both convolutions and self-attention. It offers the following advantages. First of all, outlook attention encodes spatial information by measuring the similarity between pairs of token representations, which is more parameter-efficient for feature learning than convolutions, as studied in previous work . Second, outlook attention adopts a sliding window mechanism to locally encode token representations at fine level, and to some extent preserves the crucial positional information for vision tasks . Third, the way of generating attention weights is simple and efficient. Unlike self-attention that relies on a query-key matrix multiplication, our outlook weight can be directly produced by a simple reshaping operation, saving computation. To see this, we compare the computation for a self-attention (SA) and that for a local version of self-attention (LSA) when operating on tokens with a sliding window size :
Considering a normal case in which , , and , our outlook attention is more computationally efficient as .
2 Network Architecture Variants
We build the proposed VOLO based on the LV-ViT model which we find is a surprisingly strong baseline that achieves 86.2% ImageNet top-1 accuracy with 150M learnable parameters. The original LV-ViT model consists of a patch embedding module that maps an input image of size to tokens and a sequence of transformers that operate on the tokens. To leverage the fine-level token representations, in the first stage, we adjust the patch embedding module to make the image tokenize on small image patches of size instead of . A stack of Outlookers is used to generate more expressive token representations at the fine level. In the second stage, another patch embedding module is utilized to downsample the tokens. A sequence of transformers is then adopted to encode global information.
Based on the above network structure, we introduce five versions of the proposed VOLO: VOLO-D1, VOLO-D2, VOLO-D3, VOLO-D4, and VOLO-D5. Detailed hyper-parameter settings of all the five versions can be found in Table 2. In all versions, we keep the ratio of Outlooker and Transformer to around 1:3, which we have empirically found works the best in our experiments. We also add two class attention layers in the final stage to update the class embedding. The hidden dimension in Outlookers is set to half of that in Transformers.
Experiments
We evaluate our proposed VOLO on the ImageNet dataset. During training, we do not use any extra training data. Our code is based on PyTorch , the Token Labeling toolbox , and timm . We use the LV-ViT-S model with Token Labeling as our baseline.
Setup: We use the AdamW optimizer with a linear learning rate scaling strategy and weight decay rate as suggested by previous work , and are given in Table 3 for all VOLO models. Stochastic Depth is used. We train our models on the ImageNet dataset for 300 epochs. For data augmentation methods, we use CutOut , RandAug , and the Token Labeling objective with MixToken . We do not use MixUp or CutMix as they conflict with MixToken. We train all VOLO models on a machine node with 8 NVIDIA V100 or A100 GPUs except for VOLO-D5 which needs two nodes. For VOLO-D1 and VOLO-D2, 4 GPUs also suffice with batch size 512 (16G) or 1024 (32G). For finetuning on larger image resolutions, we set the batch size to 512, learning rate to 5e-6, weight decay to 1e-8 and run the models for 30 epochs. Other hyper-parameters are set the same as default. Finetuning requires 2-8 nodes depending on the model size.
Model Settings: The model settings for VOLO-D1 to VOLO-D5 are listed in Table 3. We find that larger models (with 100M+ parameters) suffer overfitting. To mitigate this issue, we set large stochastic depth rate for them. Moreover, the learning rate selection also has a slight impact on the performance. We find it is more beneficial to use larger initial learning rates for small-sized models. In addition, the crop ratio can also slightly influence the performance. Larger models prefer larger crop ratios.
We compare the proposed VOLO with the state-of-the-art models from the literature in Table 4. All results listed are based on using only ImageNet-1k images for training and no extra training data are used. “Top-1,” “Real Top-1,” and “V2 Top-1” refer to the top-1 accuracy using the original ImageNet validation labels, cleaned-up real labels , and ImageNetV2 labels , respectively. “Train size” and “Test size” represent resolutions used in training and finetuning (test for CNNs). We separate the results into five segments according to model size (number of parameters).
As can be seen, for different model sizes, our proposed VOLO consistently performs better than previous state-of-the-art models. Specially, taking the proposed VOLO-D1 with 26.6M parameters as an example, testing on a resolution of 224 already yields 84.2% top-1 accuracy on ImageNet. Finetuning on 384 resolution further improves the performance to 85.2%, which is clearly better than all the models with a comparable amount of training parameters. When the model size is scaled up to 296M, we can achieve 87.1% top-1 accuracy on ImageNet, setting a new record in case of no extra training data. To the best of our knowledge, our VOLO-D5 is the first reaching 87.1% top-1 accuracy on ImageNet without extra training data.
Our models also achieve the best results on the “Real Top-1” and “V2 Top-1” benchmarks. As shown in Table 4, our VOLO-D4 with merely 193M parameters performs much better than previous state-of-the-art models, such as CaiT-M48 and NFNet. Our models perform even better on the ImageNet-V2 benchmark. As can be seen, our VOLO-D3 can improve upon the previous best result by 0.8% (76.9% v.s. 77.7%) using only a quarter of the parameters of CaiT-M48 (86M v.s. 356M). Our largest VOLO-D5 can further boost the performance to 78%.
2 Performance of Outlooker
In this subsection, we demonstrate the importance of the proposed Outlooker in VOLO. We take the recent state-of-the-art vision transformer model, named LV-ViT-S, as our baseline. LV-ViT-S contains 16 transformers in total and receives 83.3% top-1 accuracy on ImageNet. Each token in LV-ViT-S corresponds to an image patch of size , and hence there are totally tokens for a input image. The experiment path from the LV-ViT-S baseline to our VOLO-D1 and the corresponding results can be found in Table 5.
As the goal of our proposed Outlooker is to encode expressive finer-level features, we first adjust the starting patch embedding module and change the patch size from to . We replace two transformers with our Outlooker at the fine level. As can be seen from the second row of Table 5, such a slight adjustment brings us 0.4% gain based on the baseline that already reaches 83.3% top-1 accuracy. Adding another two Outlookers further increases the performance to 83.9%. Finally, changing the head number in all the transformers from 6 to 12 and finetuning the resulting model at resolution allows us to yield a result of 85.2%, which, to the best of our knowledge, is the first time to attain 85+% accuracy within less than 30M parameters.
We also attempt to replace the proposed outlook attention with other methods for fine-level feature encoding, including local self-attention and spatial convolutions. For a fair comparison, we set the window size to for both local self-attention and convolutions. The results can be found in Table 6. As can be seen, under the same training recipe and architecture, our Outlooker performs better than both local self-attention and convolutions. In addition, we can also observe that local self-attention and convolutions can also lift the performance compared to the LV-ViT-S baseline, demonstrating that encoding fine-level token representations indeed helps.
3 Ablation Analysis
Model Scaling: We scale up the VOLO-D1 model to 4 different models (VOLO-D2 to VOLO-D5) in two different ways: 1) increasing the model size during training, including network depth, hidden dimension, expansion ratio in MLP, and head number in both Outlookers and Transformers, and 2) increasing the image resolution during finetuning and test. The specifications for all models have been shown in Table 2 and their corresponding results can be found in Table 7. We can observe that both aforementioned ways can largely improve the model performance. From VOLO-D1 to VOLO-D2, there is 1% improvement with doubled parameters. Further increasing the model size form VOLO-D2 to VOLO-D5 yields nearly another 1% accuracy gain. In addition, for all the five models, increasing the resolution during finetuning brings around 1% performance gain.
Number of Outlookers: We observe that the number of Outlookers used in our VOLO has an impact on the classification performance. Here, we investigate the influence of using different numbers of Outlookers in our VOLO. Note that all Outlookers act on finer-level token representations (). The results have been shown in the top part of Table 8. Without any Outlookers, the baseline with 16 transformers receives 83.3% accuracy. Increasing the number of Outlookers can improve the result but the performance saturates when using 4 Outlookers. Further adding Outlookers does not bring any performance gain. Thus, when scaling up the model, we approximately use a ratio of 1:3 for Outlooker and Transformers.
Head Number in Outlookers: In Transformers, the channel dimension in each head is inversely proportional with the head number given a fixed hidden dimension. Differently, in Outlookers, the channel dimension in each head is fixed when the kernel size is fixed (\ie, ). So, will Outlookers perform better if more heads are used? In the bottom part of Table 8, we show the results with different head numbers Outlookers. Experiments show that using more heads in Outlookers can slightly improve the performance with nearly no extra parameter increase but such increase stops when the head number is more than 6. Therefore, by default, we set the head number in Outlookers to 6 for 384 hidden dimension. When the hidden dimension is set to 768, we use 12 heads in Outlookers.
4 Semantic Segmentation
In this subsection, we use our VOLO as pretrained models to evaluate the performance in semantic segmentation. Our code is based on mmsegmentation . We report results on two widely-used segmentation benchmarks: Cityscapes and ADE20K . The UperNet segmentation framework is adopted. In training, we utilize the AdamW optimizer with an initial learning rate of 6e-5 and a weight decay of 0.01. We also use a linear learning schedule with a minimum learning rate of 5e-6. All models can be trained on a machine node with 8 A100 GPUs. For cityscapes, we set the batch size to 8 and the input resolution to . For ADE20K, the batch size is set to 16 and input resolution is used. As suggested by , we report results in terms of mean intersection-over-union (mIoU) for both datasets and mean pixel accuracy for ADE20K. In inference, we perform multi-scale test with interpolation rates of [0.75, 1.0, 1.25, 1.5, 1.75].
Cityscapes is one of the most popular datasets for semantic segmentation, which targets at street scene segmentation. It has 5K high-quality pixel-annotated images with resolution and contains 19 classes in total. As in most previous work, we split the whole dataset into three splits for training, validation and test, which contain 2,975, 500, and 1,525 images, respectively. We report results on the validation set. The comparison results can be found in Table 9. It is obvious that the proposed approach outperforms all other methods, including the recent state-of-the-art SegFormer-B5 model. Our VOLO-D4 with UperNet decoder head achieves the best result 84.3%, 0.3% better than the previous state-of-the-art result 84.0% made by SegFormer-B5. According to PaperWithCodehttps://paperswithcode.com/sota/semantic-segmentation-on-cityscapes-val, this is a new state-of-the-art result on Cityscapes validation set.
4.2 ADE20K
We also run experiments on the widely-used ADE20K dataset. ADE20K contains 25K images in total, including 20K images for training, 2K images for validation, and 3K images for test. It covers 150 different common foreground categories. We compare our segmentation results with previous state-of-the-art segmentation methods in Table 10. Without pretraining on large-scale datasets, such as ImageNet-22K, our VOLO-D1 with UperNet achieves an mIoU score of 50.5. When the VOLO-D5 is used as backbone, the mIoU score can be further improved to 54.3, a new state-of-the-art result on ADE20K with no extra pretraining data except for ImageNet-1k.
Related Work
As one of the most fundamental problems in computer vision, image classification has experienced remarkable progress since the introduction of deep neural network models. In what follows, we briefly review those successful models that are closely related to this work.
Earlier models attaining state-of-the-art performance for image classification are mostly CNN-based ones that simply stack a sequence of spatial convolutions and poolings, represented by AlexNet and VGGNet . ResNets advances the design of CNN architectures by introducing skip connections to enable training of very deep models. Inceptions and ResNeXt examine the design principles of the model building blocks and introduce multiple parallel paths of sets of specialized filters. SENet presents a squeeze-and-excitation module to explicitly model the inter-dependencies among channels. DPNs leverage both residual and dense connections for designing stronger building blocks. EfficientNet and NasNet take advantage of neural architecture search to search for powerful network architectures. Later state-of-the-art models mostly utilize different training or optimization methods or finetuning techniques to improve EfficientNet. Very recently, NFNet breaks the dominance of EfficientNet by designing a normalization-free architecture, making the first work attaining 86.5% top-1 accuracy on ImageNet using no extra data. CNNs, as the de-facto networks in visual recognition for years, have indeed been very successful but their focus is on how to learn more discriminative local features by designing better architectures. Essentially, they are short of the capability of explicitly building global relationships among representations that have been proven crucial .
Recent progress on image classification is mostly driven by attention-based models or specifically transformer-based models. Transformers make use of the self-attention mechanism, making modeling long-range dependencies possible. Transformers are originally designed for natural language tasks and have recently been demonstrated effective in image classification. Dosovitskiy et al. are among the first to show that purely transformer-based architectures (\ie, ViT) can also get state-of-the-art performance in image classification but require large-scale datasets, such as ImageNet-22k and JFT-300M (which is not publicly available) for pretraining. DeiT and T2T-ViT mitigate the problem of ViTs requiring large-scale datasets and propose data efficient ViTs. Since then, a surge of works on ViT continuously come into being with further improvements. Some of them introduce local dependency into vision transformers by modifying the patch embedding block or the transformer block or both, while others adopt a pyramid structure to reduce the overall computation while maintaining the models’ ability to capture low-level features. There are also some works aiming at solving the optimization and scaling problems of ViTs.
Our VOLO not only models long-range dependencies but also encodes fine-level features into token representations by the proposed Outlooker. Unlike the recent hybrid architectures (e.g., Hybrid-ViT and BoTNet ) that rely on convolutions for feature encoding, Outlooker proposes to use local pair-wise token similarities to encode fine-level features and spatial context into tokens features and hence is more effective and parameter-efficient. This also makes our model different from the Dynamic Convolution and Involution that generate input-dependent convolution kernels to encode the features.
Conclusions
We presented a new model, Vision Outlooker (VOLO). Extensive experiments for image classification and segmentation demonstrate VOLO outperforms CNN- and Transformer-based models, and establishes new SOTA results. We hope that the strong performance of VOLO on several computer vision tasks will encourage follow-up research on better fine-level feature learning. The performance superiority of VOLO comes from the new outlook attention mechanism that dynamically aggregates fine-level features in a dense manner, and we will continue our investigation in other applications, like natural language processing.
Acknowledgement
We gratefully acknowledge the support of NVIDIA AI Tech Center (NVAITC) to this research project, especially the great helps in GPU technology supports from Terry Jianxiong Yin (NVAITC) and Qingyi Tao (NVAITC).