Dynamic Perceiver for Efficient Visual Recognition

Yizeng Han, Dongchen Han, Zeyu Liu, Yulin Wang, Xuran Pan, Yifan Pu, Chao Deng, Junlan Feng, Shiji Song, Gao Huang

Introduction

Convolutional neural networks (CNNs) and vision Transformers have precipitated substantial advancements in visual recognition. Despite concerted efforts towards scaling up vision models for superior accuracy , the high computational demands have acted as a deterrent to their deployment in resource-constrained scenarios. Research endeavours towards improving the inference efficiency of deep networks span a multitude of directions, including lightweight architecture design , pruning , quantization , etc. In contrast to traditional models, which adhere to a static computational graph during testing, dynamic networks can adapt their computation with varying input complexities, leading to promising results in efficient visual recognition.

In the field of dynamic networks, dynamic early-exiting networks construct multiple classifiers along the depth dimension, allowing samples that yield high classification confidence at early classifiers, referred to as “easy” samples, to be rapidly predicted without activating deeper layers. Existing implementations mostly build early classifiers on intermediate features (Fig. 1 (a)). However, it has been observed that classifiers will interfere with each other and significantly degrade the performance of the final exit. A widely held belief is that deep models generally extract features from a low level to a high level, and it is more appropriate to feed the high-level features at the end of a network to a linear classifier. Early classifiers in previous literature force intermediate low-level features to encapsulate high-level semantics and be linearly separable. This essentially means that feature extraction and early classification are intricately intertwined. This sub-optimal design invariably undermines the performance of dynamic early-exiting networks.

Ideally, it is expected that 1) there is a latent code which consistently embeds semantic information for direct use in classification tasks; 2) early classification and feature extraction should be decoupled, i.e., the acquisition of semantic information should be managed by a separate branch, thereby avoiding the necessity of sharing shallow layers in a feature extractor. Under these circumstances, the latent code needs to achieve linear separability, not the low-level image features, preserving the performance of late exits. The concept of incorporating a latent code is inspired by the general-purpose architecture, Perceiver . This model leverages asymmetric attention to iteratively distill inputs into a latent code, which is then employed for specific tasks. Despite its impressive ability to process various modalities, Perceiver’s application in visual recognition encounters a significant challenge in terms of computational cost, particularly when the pixel count in images is substantial.

In this paper, we propose a novel two-branch structure (Fig. 1 (b)), named Dynamic Perceiver (Dyn-Perceiver), for efficient visual recognition. Specifically, a feature branch extracts image features from a low level to a high level. Concurrently, a trainable latent code, engineered to encapsulate the semantics pertinent to classification, is processed by a classification branch. These two branches progressively exchange information via symmetric cross-attention layers, and the token number of image features is significantly reduced compared to the original Perceiver. Critically, multiple classifiers are situated solely in the classification branch, enabling early predictions without hindering feature extraction. The outputs from both branches are ultimately fused before being supplied to the final classifier.

Our design boasts three key advantages: 1) feature extraction and early classification are explicitly decoupled, and the experiment results in Sec. 4.2 demonstrate that the early classifier in our method even improves the performance of the last exit; 2) the Dyn-Perceiver framework is simple and versatile. It does away with the need for meticulously handcrafted structures as seen in previous approaches . In essence, we can construct the classification branch on any advanced vision backbones to attain top-tier performance. Such universality also allows Dyn-Perceiver to seamlessly serve as a backbone for downstream tasks such as object detection; 3) the theoretical efficiency of early exiting in Dyn-Perceiver can effectively translate into practical speedup on different hardware devices.

We evaluate the performance of Dyn-Perceiver with multiple visual backbones including ResNet , RegNet-Y , and MobileNet-v3 . Experiments show that Dyn-Perceiver significantly outperforms various competing models in terms of the accuracy-efficiency trade-off in ImageNet classification. Notably, the inference efficiency of RegNet-Y experiences a remarkable increase of 1.9-4.8×\times without any compromise in accuracy. The practical latency of Dyn-Perceiver is also validated on CPU and GPU platforms. Additionally, our method effectively enhances the performance-efficiency trade-off in action recognition (Something-Something V1 ) and object detection tasks. For instance, Dyn-Perceiver boosts the mean average precision (mAP) of RegNet-Y by 0.9% while diminishing its computation by 43% on the COCO dataset.

Related Works

Efficient visual recognition. Extensive efforts have been dedicated to improving the inference efficiency of deep networks. Popular approaches include network pruning , weight quantization , and lightweight architecture design . However, an inherent limitation of these static models is that they all treat different samples with equal computation, leading to inevitable inefficiency. In contrast, our Dyn-Perceiver can adapt its architecture (depth) to different inputs, effectively mitigating superfluous computation for “easy” samples.

Perceiver-style architectures. Our work draws inspiration from the general-purpose model Perceiver . Perceiver’s latent code directly queries information from raw inputs. Such generality comes at the cost of expensive computation. To this end, we adopt a visual backbone as feature extractor, thereby allowing the latent code to efficiently collect information from features, which contain considerably fewer tokens. Moreover, the performance of Dyn-Perceiver profits from our symmetric attention mechanism. Finally, Perceiver is a static model, which recursively executes attention layers for a fixed number of times. Dyn-Perceiver dynamically skips the computation of deep layers.

Mobile-Former also explores convolution-attention interactions. Dyn-Perceiver differs from Mobile-Former in two key aspects: 1) while Mobile-Former strives to construct an efficient static network, our model is a universal framework especially designed for dynamic early exiting; 2) the convolution’s input in a Mobile-Former block is the output from attention, rendering the inference pipeline a sequential process. Nevertheless, our computation in two branches is independent and hence more parallel-friendly.

Dynamic early exiting facilitates swift output predictions at shallower layers, reducing redundant computation in deep layers. Past observations have noted that the direct insertion of early classifiers degrades the performance of the final exit. Multi-scale dense network (MSDNet) and resolution adaptive network (RANet) partially address this via multi-scale structures and dense connections. However, their early classifiers are still appended on intermediate features. As a countermeasure, our model explicitly decouples feature extraction and early classification via a dual-branch architecture, which effectively improves the performance of the final exit. Furthermore, Dyn-Perceiver is a general and simple framework. It can be effortlessly constructed atop various backbones without requiring the intricately-designed architectures such as MSDNet. This adaptability allows Dyn-Perceiver to seamlessly function as a backbone for downstream tasks.

Method

In this section, we first provide an overview of the proposed Dyn-Perceiver (Sec. 3.1). Then the main components are explained (Sec. 3.2). We finally introduce the adaptive inference paradigm and the training strategy (Sec. 3.3).

Overall architecture. To explicitly decouple the feature extraction process and the early classification task, we propose a novel two-branch architecture consisting of 4 stages (Fig. 2). The first branch, refered to as the feature branch, can be designed as any visual backbone. In this paper, we implement it as a CNN for efficiency. The feature branch takes an image as input and generates feature maps (X0\mathbf{X}_{0} to X4\mathbf{X}_{4}) from a low level to a high level. The second branch, denoted as the classification branch, receives a trainable latent code Z0\mathbf{Z}_{0} as input. This latent code is randomly initialized and then processed by a series of self-attention operations. Following the common practice in popular vision models , we construct a token mixer (blue arrows in Fig. 2) between every two stages to reduce the token length and expand the hidden dimension of the latent code.

Symmetric cross attention. To incorporate the semantic information into the latent code, we adopt feature-to-latent (X2Z\mathbf{X2Z}) cross attention (green arrows in Fig. 2) at the start of each stage. Subsequently, the two branches conduct convolution and self-attention operations independently. The semantic information in the latent code is then integrated into the feature branch via latent-to-feature (Z2X\mathbf{Z2X}) cross attention (red arrows in Fig. 2) at the end of each stage.

Dynamic early exiting. The output from two branches are ultimately merged before being input to a linear classifier. Importantly, intermediate classifiers are added at the end of the last two stages of the classification branch to facilitate dynamic early exiting without disrupting feature extraction.

2 Main components

In this subsection, we present the main components in Dyn-Perceiver, generally in their order of execution.

Classifiers. To conduct dynamic early exiting without disrupting feature extraction, we build classifier heads after the last two stages only in the classification branch. Concretely, we pool the latent code along the token dimension and feed the result to a classification head. The classifier at the end of the feature branch is kept, as we find it slightly improves the dynamic inference performance. We also find that early classifiers at the first two stages bring limited improvements in dynamic early exiting. Finally, we concatenate the outputs from two branches and establish the last classifier based on the merged features.

Forward Knowledge Transfer (FKT). Inspired by the previous work on training multi-exit models , we propose to transfer the knowledge of early classifiers to deep ones. Specifically, a linear layer is attached to the output of an early classifier. The pooled latent code in the next stage is concatenated with the output of this linear layer before being fed to the classifier (Fig. 6). It is worth noting that instead of using a pretrain-finetune strategy as in , our FKT modules can directly improve the performance of both early and deep classifiers in end-to-end training (see the empirical analysis in Sec. 4.2). We believe that FKT could be viewed as a shortcut between classifiers, which also facilitates the optimization of early classifiers.

3 Inference and training

Dynamic early exiting. To reduce the redundant computation on “easy” samples, we conduct dynamic early exiting based on the classification confidence of early classifiers. The inference procedure for processing “hard” and “easy” samples is illustrated in Fig. 7. We can observe from Eq. 1 that the output of stage ii in the classification branch Zi\mathbf{Z}_{i} does not rely on the output of the same stage in the feature branch Xi\mathbf{X}_{i}. Therefore, the early prediction can be obtained by first activating a stage in the classification branch and its followed classifier. If the confidence (the max value of the Softmax probability) exceeds a threshold, the forward propagation terminates without activating deeper layers.

Training with self-distillation. We propose to use our last classifier to guide the training of early exits. Specifically, the loss function for the kk-th classifier can be written as

The overall training loss can be constructed by accumulating the loss from all exits: L ⁣= ⁣∑k=1KLk\mathcal{L}\!=\!\sum_{k=1}^{K}\mathcal{L}_{k}, and α\alpha in Sec. 3.3 is simply set as 0.5 in all our experiments.

Experiments

In this section, we first evaluate Dyn-Perceiver with different visual backbones on ImageNet , and then validate the proposed method in action recognition on Something-Something V1 (Sec. 4.1). Ablation studies (Sec. 4.2) and visualization (Sec. 4.3) are further presented to give a deeper understanding of our approach. Finally, we demonstrate the versatility of Dyn-Perceiver by using it as a backbone for COCO object detection (Sec. 4.4).

Datasets. ImageNet comprises 1000 classes, with 1.2 million and 50,000 images for training and validation. The images in ImageNet are of size 224×\times224. Something-Something V1 is a large-scale human action dataset that includes 98k videos, and we use the official training-validation split. The COCO dataset contains 80 categories with 118k training images and 5k validation images. We use the average FLOPs (floating-point operations) on the validation set of each dataset to measure the computational cost. The FLOPs are calculated with 8 224×\times224 frames per video on Something-Something V1 and are computed based on an input size of 1280×\times800 on COCO.

Models. We implement the feature branch with ResNet-50 , RegNet-Y and MobileNet-v3 . For the classification branch, we choose the initial token number LL of the latent code from {128,192,256} to construct different-sized models. The head number of self attention in stage ii is fixed as 2i−1,i ⁣= ⁣1, ⁣2, ⁣3, ⁣42^{i-1},i\!=\!1,\!2,\!3,\!4. The cross-attention layers all have 1 head. Other details are listed in Appendix A.

Inference and training. To perform dynamic early exiting on ImageNet, we randomly split 50,000 images from the training set. Then we vary the computation budget, solve the confidence thresholds on the split data as in , and evaluate the validation accuracy. The training setup for ImageNet classification is provided in Appendix B. On Something-Something V1, we replace the CNN backbone in TSM with ours and follow all the data-processing setups in . In COCO object detection, ImageNet-pretrained models are finetuned for 12 epochs with the default configuration of RetinaNet in MMDetection .

Results on ResNet are shown in Fig. 8 (a). The early-exiting performance of our Dyn-Perceiver is represented in gray curves, with the highest accuracy under each budget depicted by black curves. The multiple curves correspond to different models, whose detailed configurations are provided in Appendix A. We control the model complexity by manipulating the width (0.375-0.75×\times) of ResNet-50 and the number of initial tokens LL in the latent code. Our models are compared with various ResNet-based adaptive inference competitors, including layer skipping (Conv-AIG and SkipNet ), channel skipping (BAS-ResNet and Channel Selection ), and spatial-wise dynamic networks (DynConv and LASNet ). It can be observed that Dyn-Perceiver significantly outperforms other types of dynamic networks. Notably, apart from the performance, a key advantage of Dyn-Perceiver is its flexibility to adjust the computational cost with a single model. When the resource budget varies, we can simply set appropriate early-exiting thresholds to meet the constraint instead of training another model with different sparsity like other methods.

Results on RegNets. We further implement Dyn-Perceiver on RegNet-Y from 400M to 3.2G FLOPs and compare our method with multiple static backbones. The results in Fig. 8 (b) show the consistent improvement of our method across a wide range of computational budgets. For instance, Dyn-Perceiver reduces the computation by 4.8×4.8\times to achieve the same accuracy as a RegNet-Y-4GF. Compared with the recent Swin-Transformer and Vision Transformer with Deformable Attention (DAT) , Dyn-Perceiver reduces the computation by 1.8×1.8\times and 1.4×1.4\times respectively.

Comparison with early-exiting networks. Our RegNet-based Dyn-Perceiver is also compared with state-of-the-art dynamic early-exiting networks, including MSDNet , RANet , MSDNet trained with improved training techniques , glance-and-focus network (GFNet) , dynamic vision Transformer (DVT) , and the recent CF-ViT . The results in Fig. 8 (c) demonstrate that Dyn-Perceiver consistently outperforms these competitors.

Results on MobileNet-v3. We further validate Dyn-Perceiver on MobileNet-v3 with different width factors (0.75-1.5×\times). Our method is compared with various competitive baselines, including CNN (MobileNet-v3 ), vision Transformers (DeiT , T2T-ViT , PVT ), and hybrid models (MobileViT , MobleViT-v2 , EfficientFormer , EdgeViT , Lite Vision Transformer (LVT) , EdgeNeXt and MixFormer ). As can be seen from Fig. 10 that Dyn-Perceiver consistently outperforms the competitors in a wide range of computational budgets. For example, when the budget ranges in 0.2-0.8 GFLOPs, Dyn-Perceiver has ∼ ⁣ ⁣\sim\!\! 1.2-1.4×\times less computation than MobileNet-v3 when achieving the same performance.

Action recognition. We implement ResNet-based (for a fair comparison with baselines) Dyn-Perceiver in the TSM framework and compare our method with competitors including TSM , TRN , ECO and AdaFuse . As shown in Fig. 10, Dyn-Perceiver can seamlessly be applied in video classification and achieves a favorable trade-off between accuracy and efficiency.

The practical efficiency. We evaluate the practical efficiency of Dyn-Perceiver across multiple hardware platforms, including a mobile device (Nvidia Jetson TX2), a desktop CPU (Intel i5-8265U), and a server GPU (Nvidia A100). We set the batch size to 1 on CPUs and 128 on GPU. The accuracy-latency curves in Fig. 11 demonstrate that the exceptional theoretical efficiency of dynamic early exiting can effectively translate into the realistic speedup across different hardware platforms. For instance, while the recent Mobile-Former exhibits remarkable performance in theoretical efficiency, the MobileNet-v3-based Dyn-Perceiver models consistently run faster on hardware while achieving comparable accuracy. We conjecture this is because our two-branch structure can be executed in parallel, and the regular activation functions are more hardware-friendly compared to dynamic ReLU adopted in .

Comparison with Perceiver and Perceiver IO is presented in Tab. 1. By introducing the feature branch, the symmetric cross-attention mechanism, and the dynamic early-exiting paradigm, our method significantly reduces the computation without sacrificing accuracy.

2 Ablation studies

We conduct ablation studies with our RegNetY-400MF-based model to validate the effectiveness of our two-branch framework and different design choices.

The effectiveness of our two-branch framework. We first demonstrate that incorporating early exits (EE) into a standard model degrades its final performance. We experiment with a CNN (RegNet-Y-800M ) and a vision Transformer with a CLS token (T2T-ViT-7 ). An early exit is built on the feature map at stage 3 of the CNN and the CLS token at the 4th block of T2T-ViT-7, respectively. The accuracy of different exits is reported in Tab. 2. It is observed that the performance of both the CNN and the vision Transformer is significantly affected by the early exit. We can conclude that the CLS token cannot serve as our latent code perfectly, as it frequently participates in the attention operation with other patches in each block, and the network weights for processing the CLS token are shared with those for processing image features. In other words, feature extraction and early classification are still closely coupled.

Next, we conduct experiments with our two-branch structure, which is compared with two variants: the first has the same architecture but without early exits (EE), and the second incorporates an EE in the feature branch. The accuracy of different exits is listed in Tab. 2. The results suggest that: 1) the classification branch mitigates the accuracy drop brought by the EE to some extent, even if it is placed in the feature branch; 2) the EE in the feature branch still downgrades the last classifier’s performance by interfering with feature extraction; 3) building EE in the classification branch successfully decouples feature extraction and early classification and is therefore superior to the former choice. Moreover, the last exit even outperforms the variant without any EE. We conjecture that the EE provides a “deep supervision” for the classification branch. The above analysis indicates that our two-branch architecture is the key to avoiding the negative effect brought by early exits.

The effect of different components. We start from a “vanilla” two-branch structure that lacks cross-attention layers and FKT modules, and train it without self-distillation. In the vanilla model, we retain the first X2Z\mathbf{X2Z} cross attention, as we find the training divergent without it. Then we progressively add the components introduced in Sec. 3.2. The first line in Tab. 3 shows that the classification branch performs subpar without aggregating sufficient information from the feature branch. Next, the X2Z\mathbf{X2Z} cross attention significantly improves the performance of the classification branch. The 4th line of Tab. 3 demonstrates that the semantic information in the latent code also bolsters the performance of the feature branch. FKT and self-distillation further improve the accuracy of early classifiers. Finally, we remove the token mixers, which means that the token length and the channel number of the latent code are kept the same across different stages. We can witness that the token mixers are also important to the final performance.

3 Visualization

Fig. 12 show the images that are output by the first (“easy”) and the last (“hard”) exit of Dyn-Perceiver during dynamic early exiting. We can easily tell that “easy” samples generally contain simpler backgrounds, and the foreground objects usually have clearer appearances and standard poses. In the “hard” images, the foreground objects may have incomplete appearances (e.g. the balloon) or are very small in the scene (e.g. the Cacatua galerita). Interestingly, the harmonicas don’t even appear in the hard images, yet they are still correctly classified by the last exit. This indicates that our latent code captures rich semantic-level information to understand the “playing harmonica” action.

4 Object detection results

The recent early-exiting networks usually have specially designed architectures, and may not be suitable to apply on downstream tasks, e.g. object detection. In contrast, Dyn-Perceiver can be built on top of standard vision models, and therefore can seamlessly serve as a backbone for object detection. We implement a RegNet-based model in RetinaNet . Mean average precision (mAP) on the COCO validation set is used to measure the detection performance. Note that early exiting is not used in this task, and the experiment is mainly to demonstrate the generality of our model. The results in Tab. 4 suggest that Dyn-Perceiver outperforms the baselines even with less computation. The performance on object detection further validates that early classifiers in the classification branch will not downgrade the quality of the feature pyramid extracted by the feature branch. To the best of our knowledge, Dyn-Perceiver is the first dynamic early-exiting network that is empirically evaluated in the object detection task.

Conclusion

We introduced Dynamic Perceiver (Dyn-Perceiver), to explicitly decouple feature extraction and early exiting with a two-branch structure for efficient visual recognition. The inspiration came from the general-purpose architecture Perceiver using a latent code to directly query inputs. We adapted Perceiver to efficient visual recognition by introducing a feature branch. The latent code in Dyn-Perceiver is processed by a classification branch, and early exits are only inserted in this classification branch, thus not affecting the coarse-to-fine feature extraction process. The two branches interact with each other via symmetric cross-attention layers. Experiments on ImageNet demonstrated that our design effectively mitigates the negative effects brought by early exits. The framework consistently reduced the computational cost of different visual backbones on image classification, action recognition, and object detection tasks. Dyn-Perceiver significantly outperformed various competitive models in balancing accuracy and efficiency. Furthermore, evaluations on multiple hardware devices showcased the preferable inference latency of our method.

Acknowledgement. This work is supported in part by National Key R&D Program of China (2021ZD0140407), the National Natural Science Foundation of China (62022048, 62276150) and the Tsinghua University-China Mobile Communications Group Co.,Ltd. Joint Institute. We also appreciate the generous donation of computing resources by High-Flyer AI.

References

Appendix

We provide more details of our experiments, including the configuration of all models reported in the main text (Appendix A) and the training setup (Appendix B). We also provide the speed test results of Dyn-Perceiver implemented on ResNet (Appendix C).

Appendix A Model Configuration

In Fig. 8 and Fig. 10, we reported the results of different-sized Dyn-Perceiver. Here we list the detailed configuration of Dyn-Perceiver implemented on ResNet (Tab. 5), RegNet-Y (Tab. 6) and MobileNet-v3 (Tab. 7). SA is the abbreviation of self-attention.

Appendix B Training Settings

Image classification. The training settings of Dyn-Perceiver are demonstrated in Tab. 8. For simplicity, we use the same settings to train MobileNet-based models and the ResNet/RegNet-based models, except for the training epochs. For each experiment, we select batch size from {1024, 2048} based on model sizes and the GPU memory.

Action recognition. We use the TSM code basehttps://github.com/mit-han-lab/temporal-shift-module. and add temporal shift to our latent code Z\mathbf{Z} in self-attention blocks. We follow the settings in the official implementation and sum up the classification loss from different exits with the same weights as in ImageNet training.

Object detection. We finetune the ImageNet-pretrained checkpoints on COCO for 12 epochs following the official settings of RetinaNet in the MMDetection code basehttps://www.github.com/open-mmlab/mmdetection. Since the feature maps from different stages are required by the detection head, we do not perform early exiting here, and the experiment is only to demonstrate the capability of Dyn-Perceiver to serve as a detection backbone.

Appendix C Speed Test for ResNet-based Model

We also test the practical efficiency of our smallest ResNet-based Dyn-Perceiver (model 1 in Tab. 5) on the desktop i5 CPU and the A100 GPU. The latency-accuracy curves on CPU and GPU are presented in Fig. 13(a) and Fig. 13(b), respectively. It could be observed that Dyn-Perceiver achieves satisfying speedup on the two devices when achieving the same accuracy with the baseline.