Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios

Jiashi Li, Xin Xia, Wei Li, Huixia Li, Xing Wang, Xuefeng Xiao, Rui Wang, Min Zheng, Xin Pan

Introduction

Recently, vision Transformers (ViTs) have received increasing attention in industry and academia, and demonstrated much success in various computer vision tasks, such as image classification, object detection, semantic segmentation and etc. However, CNNs still dominate vision tasks from a real-world deployment perspective, because ViTs are usually much slower than classical CNNs, e.g. ResNets. There are some factors that limit the inference speed of the Transformer model, including quadratic complexity with respect to token length of the Multi-Head Self Attention (MHSA) mechanism, non-foldable LayerNorm and GELU layers, the complex model design causes frequent memory access and copying, etc.

Many works have struggled to free ViTs from high latency dilemma. For example, Swin Transformer and PVT try to design more efficient spatial attention mechanisms to alleviate the quadratic-increasing computation complexity of MHSA. The others consider combining efficient convolution blocks and powerful Transformer blocks to design CNN-Transformer hybrid architecture to obtain a better trade-off between accuracy and latency. Coincidentally, almost all existing hybrid architectures adopt convolution blocks in the shallow stages and just stack Transformer block in the last few stages. However, we observe that such a hybrid strategy is effortless to lead to performance saturation on downstream tasks (e.g. segmentation and detection). Furthermore, we found that both convolution blocks and Transformer blocks in existing works can not possess characteristics of efficiency and performance at the same time. Although the accuracy-latency trade-off has been improved when compared with Vision Transformer, the overall performance of the existing hybrid architecture is still far away from satisfactory.

To address the above issues, this work develops three important components to design efficient vision Transformer networks. Firstly, we introduce the Next Convolution Block (NCB), which is skilled at capturing short-term dependency information in visual data with a novel deployment-friendly Multi-Head Convolutional Attention (MHCA). Secondly, we build the Next Transformer Block (NTB), NTB is not only an expert in capturing long-term dependency information but also works as a lightweight and high-and-low-frequency signal mixer to enhance modeling capability. Finally, we design Next Hybrid Strategy (NHS) to stack NCB and NTB in a novel hybrid paradigm in each stage, which greatly reduce the proportion of the Transformer block and retaining the high precision of the vision Transformer network in various downstream tasks.

Based on the above-proposed approaches, we propose next generation vision Transformer for realistic industrial deployment scenarios (abbreviated as Next-ViT). In this paper, to present a fair comparison, we provide a view that treats the latency on the specific hardware as direct efficiency feedback. TensorRT and CoreML represent generic and easy-to-deploy solutions for server-side and mobile-side devices, respectively, that help provide convincing hardware-oriented performance guidance. With this direct and accurate guidance, we redraw the accuracy and latency trade-off diagram of several existing competitive models in Figure 1. As depicted in Figure 1(a)(d), Next-ViT achieves best latency/accuracy trade-off on ImageNet-1K classification task. More importantly, Next-ViT shows a more significant latency/accuracy trade-off superiority on downstream tasks. As shown in Figure 1(b)(c), on TensorRT, Next-ViT outperforms ResNet by 5.5 mAP (from 40.4 to 45.9) on COCO detection and 7.7% mIoU (from 38.8% to 46.5%) on ADE20K segmentation under similar latency. Next-ViT achieves comparable performance with CSWin, while the inference speed is increased by 3.6×. As depicted in Figure 1(e)(f), on CoreML, Next-ViT surpasses EfficientFormer by 4.6 mAP (from 42.6 to 47.2) on COCO detection and 3.5% mIoU (from 45.1% to 48.6%) on ADE20K segmentation under similar CoreML latency.

Our main contributions are summarized as follows:

We develop powerful convolution block and Transformer block, i.e. NCB and NTB, with deployment-friendly mechanisms. Next-ViT stacks NCB and NTB to build advanced CNN-Transformer hybrid architecture.

We design an innovative CNN-Transformer hybrid strategy from a new insight that boosts performance with high efficiency.

We present Next-ViT, a family of powerful vision Transformer architecture. Extensive experiments demonstrate the advantage of Next-ViT. It achieves SOTA latency/accuracy trade-off on image classification, object detection and semantic segmentation on TensorRT and CoreML.

Related Work

Convolutional Networks. Over the past decade, Convolutional Neural Networks (CNNs) have dominated vision architectures in a variety of computer vision tasks, including image classification, object detection, and semantic segmentation. ResNet uses residual connections to eliminate network degradation, ensuring that the network builds deeper and can capture high-level abstractions. DenseNet alternately enhances feature reuse and concatenates feature maps through dense connections. MobileNets introduce depthwise convolution and point-wise convolution to build models with small memory and low latency. ShuffleNet adopts group point-wise convolution and channel shuffle to reduce the computational cost further. ShuffleNetv2 propose that network architecture design should consider the direct metric such as speed, instead of the indirect metric like FLOPs. ConvNeXt reviews the design of the vision Transformers and proposes a pure CNN model that can compete favorably with SOTA hierarchical vision Transformers across multiple computer vision benchmarks, while retaining the simplicity and efficiency of standard CNNs.

Vision Transformers. Transformer is first proposed in the field of natural language processing (NLP). ViT splits the image into patches and treats these patches as words to perform self-attention, which shows that Transformer also achieves impressive performance on various vision tasks. DeiT introduces a teacher-student strategy specific to Transformers. T2T-ViT introduces a novel tokens-to-token (T2T) process to progressively tokenize images to tokens and structurally aggregate tokens. Swin Transformer proposes a general-purpose Transformer backbone, which constructs hierarchical feature maps and has linear computational complexity to image size. PiT incorporates a pooling layer into ViT, and shows that these advantages can be well harmonized to ViT through extensive experiments. Today, researchers pay more attention to efficiency, including efficient self-attention, training strategy, pyramid design, and etc.

Hybrid Models. Recent works have shown that combining convolution and Transformer as a hybrid architecture helps absorb the strengths of both architectures. BoTNet replaces the spatial convolutions with global self-attention in the final three bottleneck blocks of ResNet. CvT introduces the depthwise and pointwise convolution in front of self-attention. CMT proposes a new Transformer based hybrid network by taking advantage of Transformers to capture long-range dependencies and CNN to model local features. In MobileViT, introduces a light-weight and general- purpose vision Transformer for mobile devices. Mobile-Former combines with the proposed lightweight cross attention to model the bridge, which is not only computationally efficient, but also has more representation power. EfficientFormer complies with a dimension consistent design that smoothly leverages hardware-friendly 4D MetaBlocks and powerful 3D MHSA blocks. In this paper, we design a family of Next-ViT models that adapt more to the realistic industrial scenarios.

Methods

In this section, we first demonstrate the overview of the proposed Next-ViT. Then, we discuss some core designs within Next-ViT, including the Next Convolution Block (NCB), Next Transformer Block (NTB) and the Next Hybrid Strategy (NHS). Moreover, we provide the architecture specifications with different model sizes.

We present the Next-ViT as illustrated in Figure 2. By convention, Next-ViT follows the hierarchical pyramid architecture equipped with a patch embedding layer and a series of convolution or Transformer blocks in each stage. The spatial resolution will be progressively reduced by 32×\times while the channel dimension will be expanded across different stages. In this chapter, we first dive deeper into designing the core blocks for information interaction and respectively develop powerful NCB and NTB to model short-term and long-term dependencies in visual data. The fusion of local and global information is also performed in NTB which further boosts modeling capability. Finally, we systematically study the manners of integrating convolution and Transformer blocks. To overcome the inherent defects of existing methods, we introduce Next Hybrid Strategy which stacks innovative NCB and NTB to build our advanced CNN-Transformer hybrid architecture.

2 Next Convolution Block (NCB)

To present the superiority of the proposed NCB, we first revisit some classical structural designs of convolution and Transformer blocks as shown in Figure 3. BottleNeck block proposed by ResNet has dominance in visual neural networks for a long time by its inherent inductive biases and deployment-friendly characteristics in most hardware platforms. Unfortunately, the effectiveness of the BottleNeck block is inadequate compared to the Transformer block. ConvNeXt block modernizes the BottleNeck block by imitating designs of Transformer block. While ConvNeXt block partly improves network performance, its inference speed on TensorRT/CoreML is severely limited by inefficient components, such as 7×77\times 7 depthwise convolution, LayerNorm, and GELU. Transformer block has achieved excellent results in the various visual task and its intrinsic superiority is jointly endowed by the paradigm of MetaFormer and the attention-based token mixer module . However, the inference speed of Transformer block is much slower than BottleNeck block due to its complex attention mechanisms, which is unbearable in most realistic industrial scenarios.

To overcome the defeats of the above blocks, we introduce a Next Convolution Block (NCB), which maintains the deployment advantage of BottleNeck block while obtaining prominent performance as Transformer block. As shown in Figure 3(f), NCB follows the general architecture of MetaFormer, which is verified to be essential to the Transformer block. In the meantime, an efficient attention-based token mixer is equally important. We design a novel Multi-Head Convolutional Attention (MHCA) as an efficient token mixer with deployment-friendly convolution operation. Finally, we build NCB with MHCA and MLP layer in the paradigm of MetaFormer. Our proposed NCB can be formulated as follows:

To free the existing attention-based token mixer from the high latency dilemma, we design a novel attention mechanism with efficient convolution operation, i.e. Convolutional Attention (CA), for fast inference speed. In the meantime, inspired by the effective multi-head design in MHSA, we build our convolutional attention with multi-head paradigm which jointly attend to information from different representation subspaces at a different position for effective local representation learning. The definition of proposed Multi-Head Convolutional Attention (MHCA) can be summarized as follows:

Here, MHCA captures information from hh parallel representation subspaces. z=[z1,z2,...,zh]z=[z_{1},z_{2},...,z_{h}] indicates to divide the input feature zz into multi-head form in channel dimension. To promote the information interaction across the multiple heads, we also equip MHCA with a projection layer (WPW^{P}). CA is single-head convolutional attention which can be defined as:

where TmT_{m} and TnT_{n} are adjacent tokens in input feature zz. OO is an inner product operation with trainable parameter WW and input tokens T{m,n}T_{\{m,n\}}. CA is capable of learning affinity between different tokens in the local receptive field through iteratively optimizing trainable parameter WW. Concretely, the implementation of MHCA is carried out with a group convolution (multi-head convolution) and a point-wise convolution, as shown in Figure 3(f). We uniformly set head dim to 32 in all MHCA for fast inference speed with various date-type on TensorRT. Besides, we adopt efficient BatchNorm (BN) and ReLU activation function in NCB rather than LayerNorm (LN) and GELU in traditional Transformer blocks, which further accelerates inference speed. Experimental results in the ablation study show the superiority of NCB compared with existing blocks, e.g BottleNeck block, ConvNext block, LSA block and etc.

3 Next Transformer Block (NTB)

Although the local representation has been effectively learned via NCB, the capture of global information is urgent to be addressed. Transformer block has a strong ability to capture low-frequency signals which provide global information (e.g global shapes and structures). Nevertheless, relevant studies have observed that Transformer blocks may deteriorate high-frequency information, such as local textures information, to a certain extent. Signals in different frequency segments are indispensable in the human visual system and will be fused in some specific way to extract more essential and distinct features.

Motivated by these observations, we develop the Next Transformer Block (NTB) to capture multi-frequency signals in the lightweight mechanism. Furthermore, NTB works as an effective multi-frequency signals mixer to further enhance overall modeling capability. As shown in Figure 2, NTB firstly captures low-frequency signals with an Efficient Multi-Head Self Attention(E-MHSA) which can be depicted as:

where z=[z1,z2,...,zh]z=[z_{1},z_{2},...,z_{h}] denotes to divide the input feature zz into multi-head form in channel dimension. SA is a spatial reduction self-attention operator which is inspired by Linear SRA and performing as:

where Attention represents a standard attention calculating as Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(Q,K,V)=\text{softmax}(\frac{QK^{T}}{d_{k}})V, in which dkd_{k} denotes the scaling factor. WQ,WK,WVW^{Q},W^{K},W^{V} are linear layers for context encoding. Ps is an avg-pool operation with stride ss for downsampling the spatial dimension before the attention operation to reduce computational cost. Specifically, We observe the time consumption of the E-MHSA module is also greatly affected by its number of channels. NTB thus performs a channel dimension reduction before the E-MHSA module with point-wise convolutions to further accelerate inference. A shrinking ratio rr is introduced for channel reduction. We also utilize Batch Normalization in the E-MHSA module for extremely efficient deployment.

Furthermore, NTB is equipped with an MHCA module that cooperates with the E-MHSA module to capture multi-frequency signals. After that, output features from E-MHSA and MHCA are concatenated to mix high-low-frequency information. Finally, an MLP layer is borrowed at the end to extract more essential and distinct features. Briefly, the implementation of NTB can be formulated as follows:

4 Next Hybrid Strategy (NHS)

Some recent works have paid great efforts to combine CNN and Transformer for efficient deployment. As shown in Figure 4(b)(c), almost all of them monotonously adopt convolution blocks in the shallow stages and just stack Transformer blocks in the last one or two stages which presents effective results in the classification task. Unfortunately, we observe that these traditional hybrid strategies are effortless to reach performance saturation on downstream tasks (e.g. segmentation and detection). The reason is, the classification task just uses outputs from the last stage for prediction while downstream tasks (e.g. segmentation and detection) usually rely on features from each stage to gain better results. The traditional hybrid strategies, however, just stack Transformer blocks in the last few stages. The shallow stages thus fail to capture global information, e.g. global shapes and structures of object, which is vital to segmentation and detection tasks.

To overcome the defeats of existing hybrid strategies, we propose a Next Hybrid Strategy (NHS) from a new insight, which creatively stacks convolution block (NCB) and Transformer block (NTB) with (N+1)∗L(N+1)*L hybrid paradigm. NHS significantly promotes model performance in downstream tasks under controlling the proportion of Transformer block for efficient deployment. Firstly, in order to endow the shallow stages with the capability of capturing global information, we present a novel hybrid strategy in (NCB×N+NTB×1)(\text{NCB}\times N+\text{NTB}\times 1) pattern, which sequentially stack NN NCB and one NTB in each stage as shown in Figure 4(d). Specifically, the Transformer block(NTB) is placed at the end of each stage, which enables the model to learn global representation in the shallow layers. We conduct a series of experiments to verify the superiority of the proposed hybrid strategy. The performances of difference hybrid strategies are shown in Table 1. C denotes uniformly stacking convolution block(NCB) in one stage and T denotes consistently building one stage with Transformer block(NTB). Specially, HN\text{H}_{\text{N}} indicates stacking NCB and NTB with (NCB×N+NTB×1)(\text{NCB}\times N+\text{NTB}\times 1) pattern in the corresponding stage. All models in Table 1 are equipped with four stages. For example, C C C C represents consistently using convolution block in all of the four stages. For fair comparison, we build all the model under similar TensorRT latency. More implementation details are presented in Section 4. As shown in Table 1, the proposed hybrid strategy significantly promotes model performance compared with existing methods in the downstream tasks. C HN HN HN\text{C}\,\text{H}_{\text{N}}\,\text{H}_{\text{N}}\,\text{H}_{\text{N}} achieve the best overall performance. For example, C HN HN HN\text{C}\,\text{H}_{\text{N}}\,\text{H}_{\text{N}}\,\text{H}_{\text{N}} surpasses C C C T 0.8 mAP in detection and 0.8% mIoU in segmentation. Besides, The results of HN HN HN HN\text{H}_{\text{N}}\,\text{H}_{\text{N}}\,\text{H}_{\text{N}}\,\text{H}_{\text{N}} shows that placing Transformer block in the first stage will deteriorate the latency-accuracy trade-off of model.

We further verify the general effectiveness of C HN HN HN\text{C}\,\text{H}_{\text{N}}\,\text{H}_{\text{N}}\,\text{H}_{\text{N}} on large model by increasing the number of blocks in the third stage as ResNet . Experimental results of the first three rows in Table 2 show that the performance of large model is hard to promote and gradually reaches saturation. Such a phenomenon indicates that expanding model size by enlarging the NN of (NCB×N+NTB×1)(\text{NCB}\times N+\text{NTB}\times 1) pattern, i.e. simply adding more convolution block is not the best choice. It also implies that the value of NN in (NCB×N+NTB×1)(\text{NCB}\times N+\text{NTB}\times 1) pattern may seriously affect the model performance. We thus begin to explore the impact of the value of NN on the model performance through extensive experiments. As shown in Table 2 (middle), we build models with different configurations of NN on the third stage. To build model with similar latency for fair comparison, we stack LL groups of (NCB×N+NTB×1)(\text{NCB}\times N+\text{NTB}\times 1) pattern when the value of NN is small. Surprisingly, we found that stack NCB and NTB in (NCB×N+NTB×1)×L(\text{NCB}\times N+\text{NTB}\times 1)\times L pattern achieve better model performance compared to (NCB×N+NTB×1)(\text{NCB}\times N+\text{NTB}\times 1) pattern. It denotes that repeatedly combining low-frequency signal extractors and high-frequency signal extractors in a proper manner((NCB×N+NTB×1)(\text{NCB}\times N+\text{NTB}\times 1)) leads to higher quality representation learning. As shown in Table 2, model with N=4N=4 in the third stage achieving the best trade-off between performance and latency. We further build the larger model by enlarging LL of (NCB×4+NTB×1)×L(\text{NCB}\times 4+\text{NTB}\times 1)\times L pattern in the third stage. As shown in Table 2 (bottom), the performance of Base (L=4L=4) and Large (L=6L=6) model are significantly promote compared to small model, which verifies the general effectiveness of proposed (NCB×N+NTB×1)×L(\text{NCB}\times N+\text{NTB}\times 1)\times L pattern. We use N=4N=4 as the basic configurations in the rest of the paper.

We stack NCB and NTB with the above Next Hybrid Strategy to build Next-ViT, which can be formally defined as:

where i∈(1,2,3,4)i\in{(1,2,3,4)} denotes the stage index. Ψ\varPsi means NCB. Γ\varGamma means identity layer when i=1i=1, otherwise, NTB. Finally, ∮\oint indicates the operation of stacking the stages sequentially.

5 Next-ViT Architectures

To provide a fair comparison with existing SOTA networks, we present three typical variants, namely, Next-ViT-S/B/L. The architecture specifications are listed in Table 3, in which CC represents output channel and SS denotes stride of each stage. Additionally, the channel shrink ratio rr in NTB is uniformly set as 0.75 and the spatial reduction ratio ss in E-MHSA is in different stages. The expansion ratios of MLP layer are set as 3 for NCB and 2 for NTB, respectively. The head dim in E-MHSA and MHCA is set as 32. For normalization layer and activation functions, both NCB and NTB use BatchNorm and ReLU.

Experimental Results

We carry out the image classification experiment on the ImageNet-1K , which contains about 1.28M training images and 50K validation images from 1K categories. For a fair comparison, we follow the training settings of the recent vision Transformer with minor changes. Concretely, all of the Next-ViT variants are trained for 300 epochs on 8 V100 GPUs with a total batch size of 2048. The resolution of the input image is resized to 224 ×\times 224. We adopt the AdamW as the optimizer with weight decay 0.1. The learning rate is gradually decayed based on the cosine strategy with the initialization of 2e-3 and the use of a linear warm-up strategy with 20 epochs for all Next-ViT variants. Besides, we have also employed the increasing stochastic depth augmentation with the maximum drop-path rate of 0.1, 0.2, 0.2 for Next-ViT-S/B/L. Models with †\dagger are trained on large-scale dataset follow SSLD. For 384 ×\times 384 input size, we fine-tune the models for 30 epochs with the weight decay of 1e-8, learning rate of 1e-5, batch size of 1024. With the input size corresponding to the respective method, latency in Table 4 is uniformly measured based on the TensorRT-8.0.3 framework with a T4 GPU (batch size=8) and CoreML framework on an iPhone12 Pro Max with iOS 16.0 (batch size=1). Note that both the iPhone 12 and iPhone 12 Pro Max are equipped with the same A14 processor.

1.2 Comparison with State-of-the-art Models

As shown in Table 4, compared to the latest state-of-the-art methods (e.g. CNNs, ViTs and hybrid networks), we achieve the best trade-off between accuracy and latency. Specifically, compared with the famous CNNs such as ResNet101 , Next-ViT-S improves the accuracy by 1.7% with a similar latency on TensorRT and faster speed on CoreML(from 4.0ms to 3.5ms). Meanwhile, Next-ViT-L achieves the similar accuracy as EfficientNet-B5 and ConvNeXt-B while 4.0×\times and 1.4×\times faster on TensorRT, 3.2×\times and 44×\times faster on CoreML. In terms of the advanced ViTs, Next-ViT-S outperforms Twins-SVT-S by 0.8% with 1.3×\times faster inference speed on TensorRT. Next-ViT-B surpasses CSwin-T by 0.5% while the inference latency is compressed by 64% on TensorRT. Finally, compared with recent hybrid methods, Next-ViT-S beats CMT-XS by 0.7% with 1.8×\times and 1.4×\times faster speed on TensorRT and CoreML. Compared to EfficientFormer-L7 , Next-ViT-L predict with 20% fewer runtime on CoreML and 25% fewer runtime on TensorRT while the performance is improved from 83.3% to 83.6%. Next-ViT-L also obtains a 15% inference latency gain and achieves a better performance than TRT-ViT-D. These results demonstrate that the proposed Next-ViT design is an effective and promising paradigm.

2 ADE20K Semantic Segmentation

To further verify the capacity of our Next-ViT, we conduct the semantic segmentation experiment on ADE20K , which contains about 20K training images and 2K validation images from 150 categories. To make fair comparisons, we also follow the training conventions of the previous vision Transformers on the Semantic FPN and UperNet frameworks. Most of models are pre-trained on the ImageNet-1k and models with †\dagger are pre-trained on large-scale dataset. All the models are pretrained with resolution 224×\times224 and then trained on ADE20K with the input size of 512×\times512. For the Semantic FPN framework, we adopt the AdamW optimizer with both the learning rate and weight decay being 0.0001. Then we train the whole network for 40K iterations with a total batch size of 32 based on the stochastic depth of 0.2 for Next-ViT-S/B/L. For the training and testing on the UperNet framework, we also train the models for 160K iterations with the stochastic depth of 0.2. AdamW optimizer is used as well but with the learning rate 6×10−56\times 10^{-5}, total batch size 16, and weight decay 0.01. Then we test the mIoU based on both single-scale and multi-scale (MS) where the scale goes from 0.5 to 1.75 with an interval of 0.25. For detection and segmentation tasks, due to some modules in Mask R-CNN and Upernet are not easy to be deployed on TensorRT and CoreML, we only measure the latency of the backbone for a fair comparison, with the same test environments as classification. For simplicity, the input size of 512×\times512 is uniformly used to measure latency in Table 5 and Table 6.

2.2 Comparison with State-of-the-art Models

In Table 5, we make a comparison with CNNs, ViTs, and recent hybrid methods as well. Next-ViT-S surpasses ResNet101 and ResNeXt101-32x4d by 7.7% and 6.8% mIoU, respectively. Next-ViT-B beats CSwin-T by 0.4% mIoU and the inference speed is accelerated by 2.5×\times on TensorRT. Compared with the Uniformer-S/B , Next-ViT-B/L achieves 2.0% and 1.1% mIoU performance gain while 0.4×\times/1.3×\times faster on CoreML and 0.8×\times/1.6×\times faster on TensorRT. Next-ViT-B surpasses EfficientFormer-L7 by 3.5% mIoU with similar CoreML runtime and 38% fewer latency on TensorRT. In terms of the UperNet framework, Next-ViT-S surpasses recent SOTA CNN model ConvNeXt 2.3% MS mIoU while 1.0×\times and 18.0×\times faster on TensorRT and CoreML respectively. Compared to the CSWin-S, Next-ViT-L achieves 3.6×\times faster speed on TensorRT with similar performance. Extensive experiments reveal that our Next-ViT achieves excellent potential on segmentation tasks.

3 Object Detection and Instance Segmentation

Next, we evaluate Next-ViT on the objection detection and instance segmentation task based on the Mask R-CNN frameworks with COCO2017 . Specifically, all of our models are pre-trained on ImageNet-1K and then finetuned following the settings of the previous works . As for the 12 epochs (1×\times) experiment, we use the AdamW optimizer with the weight decay of 0.05. There are 500 iterations for a warm-up during the training, and the learning rate will decline by 10×\times at epochs 8 and 11. Based on the 36 epochs (3×\times) experiment with multi-scale (MS) training, models are trained with the resized images such that the shorter side ranges from 480 to 800 and the longer side is at most 1333. The learning rate will decline by 10×\times at epochs 27 and 33. The other settings are the same as 1×\times.

3.2 Comparison with State-of-the-art Models

Table 6 shows the evaluation results with the Mask R-CNN framework. Based on the 1×\times schedule, Next-ViT-S surpasses ResNet101 and ResNeSt50 by 5.5 APb and 3.3 APb. Next-ViT-L beats PVTv2-B4 by 0.5 APb and predict with 4.0×\times and 3.9×\times faster runtime on TensorRT and CoreML. Compared to the EfficientFormer-L7, Next-ViT-B improves APb from 42.6 to 47.2 with similar CoreML latency and 39% fewer TensorRT runtime. Next-ViT-B outperforms TRT-ViT-D by 1.9 APb but is still faster on both TensorRT and CoreML. Based on the 3×\times schedule, Next-ViT shows the same superiority as the 1×\times. Specifically, Next-ViT-S surpasses ResNet101 by 5.2 APb with similar latency. Compared to the Twins-SVT-S, Next-ViT-S achieves 1.2 APb higher performance but with 3.2×\times faster speed on TensorRT. Next-ViT-B outperforms CSwin-T by 0.5 APb but with 2.5×\times fewer prediction time. For the Next-ViT-L, it exhibits a similar performance on object detection and instance segmentation as CSwin but the inference speed is accelerated by 79%.

4 Ablation Study and Visualization

To understand our Next-ViT better, we ablate each critical design by evaluating its performance on ImageNet-1K classification and downstream tasks. We also visualize the Fourier spectrum and heat map of output features to show the intrinsic superiority of Next-ViT.

To verify the effectiveness of the proposed NCB, we replace NCB in Next-ViT with famous blocks, such as Bottleneck in ResNet , ConvNeXt block, LSA block in Twins , and etc. For a fair comparison, we consistently use NTB and NHS to build different models under similar latency on TensorRT.

As shown in Table 7, NCB achieves the best latency/accuracy trade-off on all of the three tasks, which verifies the advantage of the proposed NCB. For example, NCB outperforms the recent ConvNeXt block by 2.9% in classification, 4.5 APb in detection and 2.8% mIoU in segmentation.

4.2 Impact of Different Shrink Ratios in NTB

Furthermore, we explore the effect of shrink ratio rr of Next Transformer Block on the overall performance of Next-ViT. As stated in Table 8, decreasing the shrinking ratio rr, i.e. the number of channels in E-MHSA module, will reduce the model latency. Furthermore, model with r=0.75r=0.75 and r=0.5r=0.5 achieve better performance over model with pure Transformer (r=1r=1). This denotes that fusing multi-frequency signals in a proper manner will enhance the model ability of representation learning.

Specially, model with r=0.75r=0.75 achieve the best latency/accuracy trade-off. It outperforms baseline model (r=1.0)(r=1.0) with 0.4%, 0.5 APb and 1.0% mIoU on classification, detection and segmentation while is more lightweight. The above results indicate the effectiveness of the proposed NTB block.

4.3 Impact of Normalization and Activation

We further study the impact of different normalization layers and activation functions in Next-ViT. As shown in Table 9, both the LN and GELU bring negligible performance improvement but with significantly higher inference latency on TensorRT. On the other hand, BN and ReLU achieve the best latency/accuracy trade-off on overall tasks. Therefore, we uniformly use BN and ReLU in Next-ViT for efficient deployment in realistic industrial scenarios.

4.4 Visualization

To verify the superiority of our Next-ViT, we visualize the Fourier spectrum and heat maps of the output features from ResNet, Swin Transformer and Next-ViT in Figure 5 (a). The spectrum distribution of ResNet denotes that convolution blocks tend to capture high-frequency signals while difficult to focus on low-frequency information. On the other hand, ViT experts in capturing low-frequency signals but ignore high-frequency signals. Finally, Next-ViT is capable of simultaneously capturing high-quality and multi-frequency signals, which shows the effectiveness of NTB.

Furthermore, As shown in Figure 5 (b), we can see that Next-ViT can capture richer texture information and more accurate global information (e.g. edge shape) compared with ResNet and Swin, which shows the stronger modeling capability of Next-ViT.

Conclusion

In this paper, we present a family of Next-ViT that stacks efficient Next Convolution Block and Next Transformer Block in a novel strategy to build powerful CNN-Transformer hybrid architecture for efficient deployment on both mobile device and server GPU. Experimental results demonstrate that Next-ViT achieves a state-of-the-art latency/accuracy trade-off across diverse visual tasks, such as image classification, object detection and semantic segmentation. We believe that our work builds a stable bridge between academic research and industrial deployment in terms of visual neural network design. We hope that our work will provide new insights and promote more research in neural network architecture design for realistic industrial deployment.

References