EfficientFormer: Vision Transformers at MobileNet Speed

Yanyu Li, Geng Yuan, Yang Wen, Ju Hu, Georgios Evangelidis, Sergey Tulyakov, Yanzhi Wang, Jian Ren

Introduction

The transformer architecture , initially designed for Natural Language Processing (NLP) tasks, introduces the Multi-Head Self Attention (MHSA) mechanism that allows the network to model long-term dependencies and is easy to parallelize. In this context, Dosovitskiy et al. adapt the attention mechanism to 2D images and propose Vision Transformer (ViT): the input image is divided into non-overlapping patches, and the inter-patch representations are learned through MHSA without inductive bias. ViTs demonstrate promising results compared to convolutional neural networks (CNNs) on computer vision tasks. Following this success, several efforts explore the potential of ViT by improving training strategies , introducing architecture changes , redesigning attention mechanisms , and elevating the performance of various vision tasks such as classification , segmentation , and detection .

On the downside, transformer models are usually times slower than competitive CNNs . There are many factors that limit the inference speed of ViT, including the massive number of parameters, quadratic-increasing computation complexity with respect to token length, non-foldable normalization layers, and lack of compiler level optimizations (e.g., Winograd for CNN ). The high latency makes transformers impractical for real-world applications on resource-constrained hardware, such as augmented or virtual reality applications on mobile devices and wearables. As a result, lightweight CNNs remain the default choice for real-time inference.

To alleviate the latency bottleneck of transformers, many approaches have been proposed. For instance, some efforts consider designing new architectures or operations by changing the linear layers with convolutional layers (CONV) , combining self-attention with MobileNet blocks , or introducing sparse attention , to reduce the computational cost, while other efforts leverage network searching algorithm or pruning to improve efficiency. Although the computation-performance trade-off has been improved by existing works, the fundamental question that relates to the applicability of transformer models remains unanswered: Can powerful vision transformers run at MobileNet speed and become a default option for edge applications? This work provides a study towards the answer through the following contributions:

First, we revisit the design principles of ViT and its variants through latency analysis (Sec. 3). Following existing work , we utilize iPhone 12 as the testbed and publicly available CoreML as the compiler, since the mobile device is widely used and the results can be easily reproduced.

Second, based on our analysis, we identify inefficient designs and operators in ViT and propose a new dimension-consistent design paradigm for vision transformers (Sec. 4.1).

Third, starting from a supernet with the new design paradigm, we propose a simple yet effective latency-driven slimming method to obtain a new family of models, namely, EfficientFormers (Sec. 4.2). We directly optimize for inference speed instead of MACs or number of parameters .

Our fastest model, EfficientFormer-L1, achieves 79.2%79.2\% top-1 accuracy on ImageNet-1K classification task with only 1.61.6 ms inference time (averaged over 1,0001,000 runs), which runs as fast as MobileNetV2×1.4\times 1.4 and wields 4.5%4.5\% higher top-1 accuracy (more results in Fig. 1 and Tab. 1). The promising results demonstrate that latency is no longer an obstacle for the widespread adoption of vision transformers. Our largest model, EfficientFormer-L7, achieves 83.3%83.3\% accuracy with only 7.07.0 ms latency, outperforms ViT×\timesMobileNet hybrid designs (MobileViT-XS, 74.8%74.8\%, 7.27.2ms) by a large margin. Additionally, we observe superior performance by employing EfficientFormer as the backbone in image detection and segmentation benchmarks (Tab. 2). We provide a preliminary answer to the aforementioned question, ViTs can achieve ultra fast inference speed and wield powerful performance at the same time. We hope our EfficientFormer can serve as a strong baseline and inspire followup works on the edge deployment of vision transformers.

Related Work

Transformers are initially proposed to handle the learning of long sequences in NLP tasks . Dosovitskiy et al. and Carion et al. adapt the transformer architecture to classification and detection, respectively, and achieve competitive performance against CNN counterparts with stronger training techniques and larger-scale datasets. DeiT further improves the training pipeline with the aid of distillation, eliminating the need for large-scale pretraining . Inspired by the competitive performance and global receptive field of transformer models, follow-up works are proposed to refine the architecture , explore the relationship between CONV nets and ViT , and adapt ViT to different computer vision tasks . Other research efforts explore the essence of attention mechanism and propose insightful variants of token mixer, e.g., local attention , spatial MLP , and pooling-mixer .

Despite the success in most vision tasks, ViT-based models cannot compete with the well-studied lightweight CNNs when the inference speed is the major concern , especially on resource-constrained edge devices . To accelerate ViT, many approaches have been introduced with different methodologies, such as proposing new architectures or modules , re-thinking self-attention and sparse-attention mechanisms , and utilizing search algorithms that are widely explored in CNNs to find smaller and faster ViTs . Recently, LeViT proposes a CONV-clothing design to accelerate vision transformer. However, in order to perform MHSA, the 44D features need to be frequently reshaped into flat patches, which is still expensive to compute on edge resources (Fig. 2). Likewise, MobileViT introduces a hybrid architecture that combines lightweight MobileNet blocks (with point-wise and depth-wise CONV) and MHSA blocks; the former is placed at early stages in the network pipeline to extract low-level features, while the latter is placed in late stages to enjoy the global receptive field. Similar approach has been explored by several works as a straightforward strategy to reduce computation.

Different from existing works, we aim at pushing the latency-performance boundary of pure vision transformers instead of relying on hybrid designs, and directly optimize for mobile latency. Through our detailed analysis (Sec. 3), we propose a new design paradigm (Sec. 4.1), which can be further elevated through architecture search (Sec. 4.2).

On-Device Latency Analysis of Vision Transformers

Most existing approaches optimize the inference speed of transformers through computation complexity (MACs) or throughput (images/sec) obtained from server GPU . While such metrics do not reflect the real on-device latency. To have a clear understanding of which operations and design choices slow down the inference of ViTs on edge devices, we perform a comprehensive latency analysis over a number of models and operations, as shown in Fig. 2, whereby the following observations are drawn.

Observation 1: Patch embedding with large kernel and stride is a speed bottleneck on mobile devices.

Patch embedding is often implemented with a non-overlapping convolution layer that has large kernel size and stride . A common belief is that the computation cost of the patch embedding layer in a transformer network is unremarkable or negligible . However, our comparison in Fig. 2 between models with large kernel and stride for patch embedding, i.e., DeiT-S and PoolFormer-S24 , and the models without it, i.e., LeViT-256 and EfficientFormer, shows that patch embedding is instead a speed bottleneck on mobile devices.

Large-kernel convolutions are not well supported by most compilers and cannot be accelerated through existing algorithms like Winograd . Alternatively, the non-overlapping patch embedding can be replaced by a convolution stem with fast downsampling that consists of several hardware-efficient 3×33\times 3 convolutions (Fig. 3).

Observation 2: Consistent feature dimension is important for the choice of token mixer. MHSA is not necessarily a speed bottleneck.

Recent work extends ViT-based models to the MetaFormer architecture consisting of MLP blocks and unspecified token mixers. Selecting a token mixer is an essential design choice when building ViT-based models. The options are many—the conventional MHSA mixer with a global receptive field, more sophisticated shifted window attention , or a non-parametric operator like pooling .

We narrow the comparison to the two token mixers, pooling and MHSA, where we choose the former for its simplicity and efficiency, while the latter for better performance. More complicated token mixers like shifted window are currently not supported by most public mobile compilers and we leave them outside our scope. Furthermore, we do not use depth-wise convolution to replace pooling as we focus on building architecture without the aid of lightweight convolutions.

To understand the latency of the two token mixers, we perform the following two comparisons:

First, by comparing PoolFormer-s24 and LeViT-256 , we observe that the Reshape operation is a bottleneck for LeViT-256. The majority of LeViT-256 is implemented with CONV on 44D tensor, requiring frequent reshaping operations when forwarding features into MHSA since the attention has to be performed on patchified 33D tensor (discarding the extra dimension of attention heads). The extensive usage of Reshape limits the speed of LeViT on mobile devices (Fig. 2). On the other hand, pooling naturally suits the 44D tensor when the network primarily consists of CONV-based implementations, e.g., CONV 1×11\times 1 as MLP implementation and CONV stem for downsampling. As a result, PoolFormer exhibits faster inference speed.

Second, by comparing DeiT-Small and LeViT-256 , we find that MHSA does not bring significant overhead on mobiles if the feature dimensions are consistent and Reshape is not required. Though much more computation intensive, DeiT-Small with a consistent 3D feature can achieve comparable speed to the new ViT variant, i.e., LeViT-256.

In this work, we propose a dimension-consistent network (Sec. 4.1) with both 44D feature implementation and 33D MHSA, but the inefficient frequent Reshape operations are eliminated.

Observation 3: CONV-BN is more latency-favorable than LN (GN)-Linear and the accuracy drawback is generally acceptable.

Choosing the MLP implementation is another essential design choice. Usually, one of the two options is selected: layer normalization (LN) with 3D linear projection (proj.) and CONV 1×11\times 1 with batch normalization (BN). CONV-BN is more latency favorable because BN can be folded into the preceding convolution for inference speedup, while dynamic normalizations, such as LN and GN, still collects running statistics at the inference phase, thus contributing to latency. From the analysis of DeiT-Small and PoolFormer-S24 in Fig. 2 and previous work , the latency introduced by LN constitutes around 10%−20%10\%-20\% latency of the whole network.

Based on our ablation study in Appendix Tab. 3, CONV-BN only slightly downgrades performance compared to GN and achieves comparable results to channel-wise LN. In this work, we apply CONV-BN as much as possible (in all latent 44D features) for the latency gain with a negligible performance drop, while using LN for the 33D features, which aligns with the original MHSA design in ViT and yields better accuracy.

Observation 4: The latency of nonlinearity is hardware and compiler dependent.

Lastly, we study nonlinearity, including GeLU, ReLU, and HardSwish. Previous work suggests GeLU is not efficient on hardware and slows down inference. However, we observe GeLU is well supported by iPhone 12 and hardly slower than its counterpart, ReLU. On the contrary, HardSwish is surprisingly slow in our experiments and may not be well supported by the compiler (LeViT-256 latency with HardSwish is 44.544.5 ms while with GeLU 11.911.9 ms). We conclude that nonlinearity should be determined on a case-by-case basis given specific hardware and compiler at hand. We believe that most of the activations will be supported in the future. In this work, we employ GeLU activations.

Design of EfficientFormer

Based on the latency analysis, we propose the design of EfficientFormer, demonstrated in Fig. 3. The network consists of a patch embedding (PatchEmbed) and stack of meta transformer blocks, denoted as MB:

where X0\mathcal{X}_{0} is the input image with batch size as BB and spatial size as [H,W][H,W], Y\mathcal{Y} is the desired output, and mm is the total number of blocks (depth). MB consists of unspecified token mixer (TokenMixer) followed by a MLP block and can be expressed as follows:

where Xi∣i>0\mathcal{X}_{i|i>0} is the intermediate feature that forwarded into the ithi^{th} MB. We further define Stage (or S) as the stack of several MetaBlocks that processes the features with the same spatial size, such as N1×N_{1}\times in Fig. 3 denoting S1\texttt{S}_{1} has N1N_{1} MetaBlocks. The network includes 44 Stages. Among each Stage, there is an embedding operation to project embedding dimension and downsample token length, denoted as Embedding in Fig. 3. With the above architecture, EfficientFormer is a fully transformer-based model without integrating MobileNet structures. Next, we dive into the details of the network design, specifically, the architecture details and the search algorithm.

With the observations in Sec. 3, we propose a dimension consistent design which splits the network into a 4D partition where operators are implemented in CONV-net style (MB4D\texttt{MB}^{4D}), and a 3D partition where linear projections and attentions are performed over 3D tensor to enjoy the global modeling power of MHSA without sacrificing efficiency (MB3D\texttt{MB}^{3D}), as shown in Fig. 3. Specifically, the network starts with 4D partition, while 3D partition is applied in the last stages. Note that Fig. 3 is just an instance, the actual length of 4D and 3D partition is specified later through architecture search.

First, input images are processed by a CONV stem with two 3×33\times 3 convolutions with stride 2 as patch embedding,

where CjC_{j} is the channel number (width) of the jthj\hskip 1.42271ptth stage. Then the network starts with MB4D\texttt{MB}^{4D} with a simple Pool mixer to extract low level features,

where ConvB,G\texttt{Conv}_{B,G} refers to whether the convolution is followed by BN and GeLU, respectively. Note here we do not employ Group or Layer Normalization (LN) before the Pool mixer as in , since the 4D partition is CONV-BN based design, thus there exists a BN in front of each Pool mixer.

After processing all the MB4D\texttt{MB}^{4D} blocks, we perform a one-time reshaping to transform the features size and enter 3D partition. MB3D\texttt{MB}^{3D} follows conventional ViT structure, as in Fig. 3. Formally,

where LinearG\texttt{Linear}_{G} denotes the Linear followed by GeLU, and

where Q,K,VQ,K,V represents query, key, and values learned by the linear projection, and bb is parameterized attention bias as position encodings.

2 Latency Driven Slimming

Design of Supernet. Based on the dimension-consistent design, we build a supernet for searching efficient models of the network architecture shown in Fig. 3 (Fig. 3 shows an example of searched final network). In order to represent such a supernet, we define the MetaPath (MP), which is the collection of possible blocks:

where II represents identity path, jj denotes the jthj^{th} Stage, and ii denotes the ithi^{th} block. The supernet can be illustrated by replacing MB in Fig. 3 with MP.

As in Eqn. 7, in S1\texttt{S}_{1} and S2\texttt{S}_{2} of the supernet, each block can select from MB4D\texttt{MB}^{4D} or II, while in S3\texttt{S}_{3} and S4\texttt{S}_{4}, the block can be MB3D\texttt{MB}^{3D}, MB4D\texttt{MB}^{4D}, or II. We only enable MB3D\texttt{MB}^{3D} in the last two Stages for two reasons. First, since the computation of MHSA grows quadratically with respect to token length, integrating it in early Stages would largely increase the computation cost. Second, applying the global MHSA to the last Stages aligns with the intuition that early stages in the networks capture low-level features, while late layers learn long-term dependencies.

Searching Algorithm. Previous hardware-aware network searching methods generally rely on hardware deployment of each candidate in search space to obtain the latency, which is time consuming . In this work, we propose a simple, fast yet effective gradient-based search algorithm to obtain a candidate network that just needs to train the supernet for once. The algorithm has three major steps.

First, we train the supernet with Gumbel Softmax sampling to get the importance score for the blocks within each MP, which can be expressed as

where α\alpha evaluates the importance of each block in MP as it represents the probability to select a block, e.g., MB4D\texttt{MB}^{4D} or MB3D\texttt{MB}^{3D} for the ithi^{th} block. ϵ∼U(0,1)\epsilon\sim U(0,1) ensures exploration, τ\tau is the temperature, and nn represents the type of blocks in MP, i.e., n∈{4D,I}n\in\{4D,I\} for S1\texttt{S}_{1} and S2\texttt{S}_{2}, and n∈{4D,3D,I}n\in\{4D,3D,I\} for S3\texttt{S}_{3} and S4\texttt{S}_{4}. By using Eqn. 8, the derivatives with respect to network weights and α\alpha can be computed easily. The training follows the standard recipe (see Sec. 5.1) to obtain the trained weights and architecture parameter α\alpha.

Second, we build a latency lookup table by collecting the on-device latency of MB4D\texttt{MB}^{4D} and MB3D\texttt{MB}^{3D} with different widths (multiples of 1616).

Finally, we perform network slimming on the supernet obtained from the first step through latency evaluation using the lookup table. Note that a typical gradient-based searching algorithm simply select the block with largest α\alpha , which does not fit our scope as it cannot search the width CjC_{j}. In fact, constructing a multiple-width supernet is memory-consuming and even unrealistic given that each MP has several branches in our design. Instead of directly searching on the complex searching space, we perform a gradual slimming on the single-width supernet as follows.

We first define the importance score for MPi\texttt{MP}_{i} as αi4DαiI\frac{\alpha_{i}^{4D}}{\alpha^{I}_{i}} and αi3D+αi4DαiI\frac{\alpha^{3D}_{i}+\alpha^{4D}_{i}}{\alpha^{I}_{i}} for S1,2\texttt{S}_{1,2} and S3,4\texttt{S}_{3,4}, respectively. Similarly, the importance score for each Stage can be obtained by summing up the scores for all MP within the Stage. With the importance score, we define the action space that includes three options: 1) select II for the least important MP, 2) remove the first MB3D\texttt{MB}^{3D}, and 3) reduce the width of the least important Stage (by multiples of 1616). Then, we calculate the resulting latency of each action through lookup table, and evaluate the accuracy drop of each action. Lastly, we choose the action based on per-latency accuracy drop (−%ms\frac{-\%}{ms}). This process is performed iteratively until target latency is achieved. We show more details of the algorithm in Appendix.

Experiments and Discussion

We implement EfficientFormer through PyTorch 1.11 and Timm library , which is the common practice in recent arts . Our models are trained on a cluster with NVIDIA A100 and V100 GPUs. The inference speed on iPhone 12 (A14 bionic chip) is measured with iOS version 15 and averaged over 1,0001,000 runs, with all available computing resources (NPU), or CPU only. CoreMLTools is used to deploy the run-time model. In addition, we provide latency analysis on Nvidia A100 GPU with batch size 64 to exploit hardware roofline. The trained PyTorch models are deployed in ONNX format and are compiled with TensorRT. We report GPU runtime that excludes preprocessing. We provide the detailed network architecture and more ablation studies in Appendix Appendix.

All EfficientFormer models are trained from scratch on ImageNet-1K dataset to perform the image classification task. We employ standard image size (224×224224\times 224) for both training and testing. We follow the training recipe from DeiT but mainly report results with 300300 training epochs to have the comparison with other ViT-based models. We use AdamW optimizer , warm-up training with 55 epochs, and a cosine annealing learning rate schedule. The initial learning rate is set as 10−3×(batchsize/1024)10^{-3}\times(batch\hskip 2.84544ptsize/1024) and the minimum learning rate is 10−510^{-5}. The teacher model for distillation is RegNetY-16GF pretrained on ImageNet with 82.9%82.9\% top-1 accuracy. Results are demonstrated in Tab. 1 and Fig. 1

Comparison to CNNs. Compared with the widely used CNN-based models, EfficientFormer achieves a better trade-off between accuracy and latency. On iPhone Neural Engine, EfficientFormer-L1 runs at MobileNetV2×1.4\times 1.4 speed while achieving 4.5%4.5\% higher top-1 accuracy. In addition, EfficientFormer-L3 runs at a similar speed to EfficientNet-B0 while achieving relative 5.3%5.3\% higher top-1 accuracy. For the models with high performance (>83%>83\% top-1), EfficientFormer-L7 runs more than 3×3\times faster than EfficientNet-B5, demonstrating the advantageous performance of our models. Moreover on desktop GPU (A100), EfficientFormer-L1 runs 38% faster than EfficientNet-B0 while achieving 2.1% higher top-1 accuracy. EfficientFormer-L7 runs 4.6×\times faster than EfficientNet-B5. These results allow us to answer the central question raised earlier; ViTs do not need to sacrifice latency to achieve good performance, and an accurate ViT can still have ultra-fast inference speed as lightweight CNNs do.

Comparison to ViTs. Conventional ViTs are still under-performing CNNs in terms of latency. For instance, DeiT-Tiny achieves similar accuracy to EfficientNet-B0 while it runs 3.4×3.4\times slower. However, EfficientFormer performs like other transformer models while running times faster. EfficientFormer-L3 achieves higher accuracy than DeiT-Small (82.4%82.4\% vs. 81.2%81.2\%) while being 4×4\times faster. It is notable that though the recent transformer variant, PoolFormer , naturally has a consistent 44D architecture and runs faster compared to typical ViTs, the absence of global MHSA greatly limits the performance upper-bound. EfficientFormer-L3 achieves 1% higher top-1 accuracy than PoolFormer-S36, while being 3×\times faster on Nvidia A100 GPU, 2.2×\times faster on iPhone NPU and 6.8×\times faster on iPhone CPU.

Comparison to Hybrid Designs. Existing hybrid designs, e.g., LeViT-256 and MobileViT, still struggle with the latency bottleneck of ViTs and can hardly outperform lightweight CNNs. For example, LeViT-256 runs slower than DeiT-Small while having 1%1\% lower top-1 accuracy. For MobileViT, which is a hybrid model with both MHSA and MobileNet blocks, we observe that it is significantly slower than CNN counterparts, e.g., MobileNetV2 and EfficientNet-B0, while the accuracy is not satisfactory either (2.3%2.3\% lower than EfficientNet-B0). Thus, simply trading-off MHSA with MobileNet blocks can hardly push forward the Pareto curve, as in Fig. 1. In contrast, EfficientFormer, as pure transformer-based model, can maintain high performance while achieving ultra-fast inference speed. EfficientFormer-L1 has 4.4% higher top-1 accuracy than MobileViT-XS and runs much faster across different hardware and compilers (1.9×\times faster on Nvidia A100 GPU Computing, 2.3×\times faster on iPhone CPU, and 4.5×\times faster on iPhone NPU). At a similar inference time, EfficientFormer-L7 outperforms MobileViT-XS by 8.5%8.5\% top-1 accuracy on ImageNet, demonstrating the superiority of our design.

2 EfficientFormer as Backbone

Object Detection and Instance Segmentation. We follow the implementation of Mask-RCNN to integrate EfficientFormer as the backbone and verify performance. We experiment over COCO-2017 which contains training and validations sets of 118118K and 55K images, respectively. The EfficientFormer backbone is initialized with ImageNet-1K pretrained weights. Similar to prior work , we use AdamW optimizer with initial learning rate of 2×10−42\times 10^{-4}, and train the model for 1212 epochs. We set the input size as 1333×8001333\times 800.

The results for detection and instance segmentation are shown in Tab. 2. EfficientFormers consistently outperform CNN (ResNet) and transformer (PoolFormer) backbones. With similar computation cost, EfficientFormer-L3 outperforms ResNet50 backbone by 3.43.4 box AP and 3.73.7 mask AP, and outperforms PoolFormer-S24 backbone with 1.31.3 box AP and 1.11.1 mask AP, proving that EfficientFormer generalizes well as a strong backbone in vision tasks.

We further validate the performance of EfficientFormer on the semantic segmentation task. We use the challenging scene parsing dataset, ADE20K , which contains 2020K training images and 22K validation ones covering 150150 class categories. Similar to existing work , we build EfficientFormer as backbone along with Semantic FPN as segmentation decoder for fair comparison. The backbone is initialized with pretrained weights on ImageNet-1K and the model is trained for 4040K iterations with a total batch size of 3232 over 88 GPUs. We follow the common practice in segmentation , use AdamW optimizer , and apply a poly learning rate schedule with power 0.90.9, starting from a initial learning rate 2×10−42\times 10^{-4}. We resize and crop input images to 512×512512\times 512 for training and shorter side as 512512 for testing (on validation set).

As shown in Tab. 2, EfficientFormer consistently outperforms CNN- and transformer-based backbones by a large margin under a similar computation budget. For example, EfficientFormer-L3 outperforms PoolFormer-S24 by 3.23.2 mIoU. We show that with global attention, EfficientFormer learns better long-term dependencies, which is beneficial in high-resolution dense prediction tasks.

3 Discussion

Relations to MetaFormer. The design of EfficientFormer is partly inspired by the MetaFormer concept . Compared to PoolFormer, EfficientFormer addresses the dimension mismatch problem, which is a root cause of inefficient edge inference, thus being capable of utilizing global MHSA without sacrificing speed. Consequently, EfficientFormer exhibits advantageous accuracy performance over PoolFormer. In spite of its fully 44D design, PoolFormer employs inefficient patch embedding and group normalization (Fig. 2), leading to increased latency. Instead, our redesigned 44D partition of EfficientFormer (Fig. 3) is more hardware friendly and exhibits better performance across several tasks.

Limitations. (i) Though most designs in EfficientFormer are general-purposed, e.g., dimension-consistent design and 44D block with CONV-BN fusion, the actual speed of EfficientFormer may vary on other platforms. For instance, if GeLU is not well supported while HardSwish is efficiently implemented on specific hardware and compiler, the operator may need to be modified accordingly. (ii) The proposed latency-driven slimming is simple and fast. However, better results may be achieved if search cost is not a concern and an enumeration-based brute search is performed.

Conclusion

In this work, we show that Vision Transformer can operate at MobileNet speed on mobile devices. Starting from a comprehensive latency analysis, we identify inefficient operators in a series of ViT-based architectures, whereby we draw important observations that guide our new design paradigm. The proposed EfficientFormer complies with a dimension consistent design that smoothly leverages hardware-friendly 44D MetaBlocks and powerful 33D MHSA blocks. We further propose a fast latency-driven slimming method to derive optimized configurations based on our design space. Extensive experiments on image classification, object detection, and segmentation tasks show that EfficientFormer models outperform existing transformer models while being faster than most competitive CNNs. The latency-driven analysis of ViT architecture and the experimental results validate our claim: powerful vision transformers can achieve ultra-fast inference speed on the edge. Future research will further explore the potential of EfficientFormer on several resource-constrained devices.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [Yes]

Did you discuss any potential negative societal impacts of your work? [Yes]

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [N/A]

Did you include complete proofs of all theoretical results? [N/A]

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes]

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes]

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [Yes]

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes]

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes]

Did you mention the license of the assets? [Yes]

Did you include any new assets either in the supplemental material or as a URL? [N/A]

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A]

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A]

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]

Appendix

Appendix A Latency-Driven Slimming Algorithm

We provide the details of the proposed latency-driven fast slimming in Alg. 1. Formulations of the algorithm can be found in Sec. 4.2. The proposed latency-driven slimming is speed-oriented, which does not require retraining for each sub-network. The importance score for each design choice is estimated based on the trainable architecture parameter α\alpha.

Appendix B Ablation Analysis

Our major conclusions and speed analysis can be found in Sec. 3 and Fig. 2. Here we include more ablation studies for different design choices, provided in Tab. 3, taking the EfficientFormer-L3 as an example. The latency is measured on iPhone 12 with CoreML, and the top-1 accuracy is obtained from the ImageNet-1K dataset.

Patch Embedding. Compared to non-overlap large-kernel patch embedding (V5 in Tab. 3), the proposed convolution stem in EfficientFormer (V1 in Tab. 3) greatly reduces inference latency by 48%48\%, while provides 0.7%0.7\% higher accuracy. We demonstrate that convolution stem is not only beneficial to model convergence and accuracy but also boosts inference speed on the mobile device by a large margin, thus can serve as a good alternative to non-overlapping patch embedding implementations.

MHSA and Latency-Driven Search. Without the proposed 33D MHSA and latency-driven search, EfficientFormer downgrades to a pure 44D design with pool mixer, which is similar to PoolFormer (the patch embeddings and normalizations are different). By comparing EfficientFormer with V1 in Tab. 3, we can observe that the integration of 33D MHSA and latency-driven search greatly boost top-1 accuracy by 2.1%2.1\% with minimal impact on the inference speed (0.50.5 ms). The results prove that MHSA with the global receptive field is an essential contribution to model performance. As a result, though enjoying faster inference speed, simply removing MHSA greatly limits the performance upper bound. In EfficientFormer, we smoothly integrate MHSA in a dimension consistent manner, obtaining better performance while simultaneously achieving ultra fast inference speed.

Normalization. Apart from the CONV-BN structure in the 44D partition of EfficientFormer, we explore Group Normalization (GN) and channel-wise Layer Normalization (LN) in the 44D partition as employed in the prior work . By comparing V1 and V2 in Tab. 3, we can observe that the GN (V2-GN) can only slightly improve accuracy (0.3%0.3\% top-1) but incurs latency overhead as it can not be folded at the inference stage. Similarly, applying LN (V2-LN) gets higher latency than BN while the performance improvement is negligible. As a result, we apply the CONV-BN structure in the entire 44D partition in EfficientFormer.

Activation Functions. We explore ReLU and HardSwish (V3 and V4 in Tab. 3) in addition to GeLU employed in this work (V1 in Tab. 3). It is widely agreed that ReLU is the simplest and fastest activation function, while GeLU and HardSwish wield better performance. We observe that ReLU can hardly provide any speedup over GeLU on iPhone 12 with CoreMLTools, while HardSwish is significantly slower than ReLU and GeLU. We draw a conclusion that the activation function can be selected on a case-by-case basis depending on the specific hardware and compiler. In this work, we use GeLU to provide better performance than ReLU while executing faster. For a fair comparison, we modify inefficient operators in other works according to the supports from iPhone 12 and CoreMLTools, e.g., report LeViT latency by changing HardSwish to GeLU.

Dimension Consistent Design. We perform the latency analysis on the Nvidia A100 GPU to show that the proposed dimension-consistent (D-C) design is beneficial besides the iPhone. We deploy the PyTorch models with batch size 64 as ONNX format and use TensorRT to compile and benchmark the latency. We report the latency results averaged over 1,0001,000 runs in Tab. 4. For the non-D-C design, we revert the proposed Meta3D block into 4D implementation, where linear projections and MLPs are all implemented with CONV1×\times1-BN instead of 3D-Linear layers, and reshaping operations become necessary in order to perform multi-head self-attention. With this configuration, attention blocks can be arbitrarily placed along with Meta4D blocks without following dimension-consistent design, while frequent reshaping is introduced. We conduct the comparison on the following two models, both with a D-C version and a non-D-C one with the exact same computation complexity:

EfficientFormer-L7, which has 8 attention blocks.

DummyNet, a handcrafted dummy model with a total of 16 attention blocks.

As can be seen from Tab. 4, the proposed dimension-consistent design achieves faster inference speed than the non-dimension-consistent design for both EfficientFormer-L7 and the DummyNet.

Latency Driven Slimming. Besides the hardware-efficient architecture design, it is still crucial to find appropriate depth and width configurations for the model to achieve satisfactory performance. To understand the benefits of our latency driven slimming, we randomly sample networks from our search space that have the same computation, i.e., 1.3 GMACs, as our searched model EfficientFormer-L1. The sampled networks are denoted as Random 1 to Random 5, which are either deeper and narrower, or shallower and wider than EfficientFormer-L1. We train the sampled models on ImageNet-1K with the same training recipe as EfficientFormer-L1. The comparison between these models is shown in Tab. 5. As can be seen, our searched EfficientFormer-L1 has better latency or higher top-1 accuracy on ImageNet-1K than the randomly sampled networks, proving the advantages of our proposed latency driven slimming.

Appendix C Analysis of Hardware Utilization

EfficientFormer improves the latency vs. accuracy trade-off through better architecture design so that higher hardware utilization is achieved. To understand hardware utilization, we employ throughput in TFLOPS (Tera FLOPs per Second) as the evaluation metric, which is calculated by model computation cost (FLOPs) divided by execution time. Models with higher throughput (TFLOPS) better exploit the computation power of the hardware.

To fairly compare with baseline models under different computation complexity, we Linearly Scale the depth and width of EfficientFormer-L1 to obtain a series of models (EfficientFormer-LS-1 to EfficientFormer-LS-14), with the number of parameters ranging from 1.1M to 31.3M and MACs from 0.09G to 3.9G, and benchmark the latency and utilization on iPhone 12 NPU.

As in Fig. 4, super-tiny models still run at about 1ms, such as EfficientFormer-LS-{1, 2, 3}, where the throughput is low and the hardware is not fully exploited. Data processing and transferring become the bottleneck. As a result, making the model super small with sacrificed accuracy is less valuable. In contrast, our 1.3GMACs model, EfficientFormer-L1 lies at a sweet point, enjoying fast inference speed (1.6ms) while maintaining high accuracy.

Furthermore, we can observe that EfficientFormer variants outperform both CNNs and ViTs in hardware utilization across different computation complexity levels. For instance, at 4-GMACs level, EfficientFormer-LS-14 outperforms DeiT-S by 3.3×\times higher TFLOPS and outperforms PoolFormer by 2.2×\times, achieving comparable throughput to ResNet50. In the lightweight domain, EfficientFormer-LS-4 has 2.2×\times higher TFLOPS than EfficientNet-B0. We demonstrate that with the proposed hardware-friendly design, EfficientFormer naturally has better hardware utilization.

Appendix D Architecture of EfficientFormers

The detailed network archiecture for EfficientFormer-L1, EfficientFormer-L3, and EfficientFormer-L7 is provided in Tab. 6. We report the resolution and number of blocks for each stage. In addition, the width of EfficientFormer is specified as the embedding dimension (Embed. Dim.). As for the MHSA block, the dimension of Query and Key is provided, and we employ eight heads for all EfficientFormer variants. MLP expansion ratio is set as default (4), as in most ViT arts .