AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition

Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang, Ping Luo

Introduction

There is a growing interest in adopting a general neural model to tackle a large variety of different tasks since it benefits in reducing the need for task-specific model design and training. Recently, Transformer demonstrates great potential in this goal considering its success in various fields, e.g., natural language processing (NLP) , visual recognition , dense prediction , Generative Adversarial Network (GAN) , reinforcement learning (RL) , robotics , and etc. However, existing literature in computer vision tend to focus on the same network with task-specific weights scenario, where a single network is used to train from scratch or fully fine-tune on a specific dataset, making it infeasible to maintain a separate model weight for every dataset when the number of task grows, especially for the increasing model capacity of state-of-the-art models (e.g., ViT-G/14 with over 1.8 billion parameters).

Different from prior arts, we step into the direction of developing same network with almost same weights and achieve superior performance than the full-tuning approach by only tuning less than 2% parameters, with the remaining over 98% parameters shared across different tasks. There are two challenges to learning universal representations using a single model. The first one lies in the pre-training stage, which requires algorithms that can learn well-generalized representations that are easy to be applied to many tasks. Recent arts in self-supervised learning can serve as a solution to this challenge. The second one, which is our main concern in this work, is to build an effective pipeline that can adapt the model obtained at the pre-training stage to various downstream tasks by tuning parameters as less as possible and keeping the left parameters frozen.

While fine-tuning pre-trained models has been widely studied in NLP , this topic is seldomly explored in the vision, where full tine-tuning of model parameters is still the dominant strategy for adapting vision transformers. However, the full fine-tuning cannot satisfy the goal of universal representation as it assigns an independent set of weights for every task. Linear probing is a straightforward approach to maintaining the pre-trained model fixed by only tuning a specific lightweight classification head for every task. However, linear probing tends to have an unsatisfactory performance and misses the opportunity of pursuing strong but non-linear features , which indeed benefit deep learning.

More recently, Bahng et.al., aimed to adapt pre-trained models by modifying raw input pixel space. Jia et.al., proposed Visual Prompt Tuning (VPT) to adapt transformer models for downstream vision tasks, which prepends several learnable parameters (prompts) to the patch embeddings and freezes the whole pre-trained backbone.

In this work, we propose a lightweight module, namely AdaptFormer, to adapt vision transformers by updating the weights of AdaptFormer. We introduce learnable parameters from the model perspective, which is different from VPT, which inserts learnable parameters into the token space. Our AdaptFormer is conceptually simple yet effective. It consists of two fully connected layers, a non-linear activation function, and a scaling factor. This module is set in parallel to the feed-forward network (FFN) of the original ViT model, as shown in Figure 2(b). This design is turned out to be effective for model transfer when processing scalable visual tokens for both image and video data (i.e., image data consists of a small scale of visual tokens while video data consists of a large scale). As shown in Figure 1, compared with the full-tuning strategy, AdaptFormer achieves comparable performance on video recognition with only about 0.1% tunable parameters. Meanwhile, with less than 2% tunable parameters, AdaptFormer surpasses the full-tuning solution by about 10% on top-1 accuracy. Similar approaches are also proposed in fine-tuning pre-trained language models (PLMs) .

The key contributions of this paper are summarized as follows: (1) We propose a simple yet effective framework, namely AdaptFormer, for adapting vision transformers to a large variety of downstream visual recognition tasks and avoiding catastrophic interference with each other. To the best of our knowledge, this is the first work that explores efficient fine-tuning in video action recognition. (2) We ablate many design choices and demonstrate the superior robustness of AdaptFormer when parameters scale up. (3) Extensive experiments on various downstream tasks demonstrate that AdaptFormer outperforms existing fine-tuning approaches significantly. By demonstrating the effectiveness of AdaptFormer on multiple visual benchmarks, we hope our work could inspire the research communities to rethink the fine-tuning mechanism in computer vision and make progress toward a flexible yet universal Transformer model for visual recognition.

Related Works

In the proposed AdaptFormer, we mainly introduce a plug-and-play module for efficiently fine-tuning the current vision Transformer models. In this section, we perform a literature review on related works from two perspectives, i.e., the vision Transformers, and efficient transfer learning for vision Transformers.

The Transformer architecture is first introduced in and has re-energized the natural language processing (NLP) field from then on . Inspired by its huge success, researches in the computer vision filed have also evolved into Transformer era since ViTs . The strong capability of modeling long-range relation has facilitated Transformer in various vision tasks, including image classification , object detection , semantic/instance segmentation , video understanding , point cloud modeling , 3D Object Recognition and even low-level processing . Furthermore, transformers have advanced the vision recognition performance by a large-scale pretraining . In such a situation, given the pre-trained Transformer models, which are more larger than the previously prevalent CNN backbones, one open question is how to fine-tune the big vision models so that they can be adapted into more specific down-stream tasks. To solve the open question, we propose AdaptFormer to transfer ViTs from the pre-trained pre-texts into the target tasks in a more effective and efficient way.

2 Efficient Transfer learning for Transformers

Transfer learning targets re-adopting a pre-trained model (either via the supervised or the unsupervised manner) as the starting point and further fine-tuning the specific model on a new task. In the NLP field, transferring the large pre-trained language models (PLMs) into downstream tasks has been the popular paradigm for a long time. Conventional arts set all the network parameters as learnable ones and adapt them to the target tasks. However, with the growth of model sizes and the complexity of the specific tasks, the conventional paradigm is inevitably limited by the huge computational burden. The NLP community has explored several ways for parameter-efficient transfer learning that only set a few parameters learnable and fine-tune them for efficiency. The pioneer works could be mainly categorized from the token and network perspectives . Basically speaking, the token-related methods typically prepend several learnable prefix vectors/tokens to the projected tokens within the multi-head self-attention layers (MHSA ). The philosophy behind it is to assist the pre-trained models in understanding downstream tasks with the guidance of extra token information. On the other hand, network-related methods integrate shallow modules to improve the model transferability. The introduced modules adapt the produced representations into the downstream tasks via features fusion.

Recently, with the emergence of a much more large-scale dataset , increasing researchers in computer vision have adopted the homologous paradigm, i.e., first pre-training and then fine-tuning, to advance the vision tasks. As for the second stage, traditional methods typically adopt the full-tuning arts in the downstream tasks. Rare attention has been drawn to the field of efficient adaptation, especially in the field of vision Transformers. Inspired by Prompting in NLP, introduced the learnable tokens in exploring the efficient adaptation for ViTs. We empirically found that the performance of prompting is hindered by the scale of tokens. That is to say, for the tasks where the number of tokens is on a small scale, e.g., image classification, Prompting is efficient for improving the model transferability. However, for larger scale tokens, e.g., video understanding, Prompting presents limited potential. This observation motivates us to introduce AdaptFormer, which is effective in the scenarios of scalable visual tokens.

Approach

We propose AdaptFormer for efficiently transferring large pre-trained vision transformer models to downstream tasks, in both image and video domains. AdaptFormer attains strong transfer learning abilities by only fine-tuning a small number of extra parameters, circumventing catastrophic interference among tasks. We illustrate the overall framework of AdaptFormer in Figure 2(b).

Each Transformer encoder mainly consists of two types of sub-layers, i.e., a multi-head self-attention layer (MHSA) and a MLP layer. In MHSA, the tokens are linearly projected and further re-formulated into three vectors, namely Q,K\bm{Q},\bm{K} and V\bm{V}. The self-attention calculation is performed on Q,K\bm{Q},\bm{K} and V\bm{V} by:

2 AdaptFormer

We propose a plug-and-play bottleneck module, namely AdaptMLPIn this paper, we use the term ‘AdaptMLP’ to denote the designed module and the term ‘AdaptFormer’ to represent the fine-tuning framework for Vision Transformers. Unless otherwise specified, we apply AdaptFormer to fine-tune the vanilla ViT backbone in this paper.. We denote the vision Transformer equipped with AdaptMLP as AdaptFormer.

Fine-tuning. During the fine-tuning phase, we only choose the newly added parameters to optimize and keep rest ones fixed. Specifically, the original model parts (blue blocks in Figure 2(b)) load weights from the pre-trained checkpoint and keeps parameters frozen. The newly added parameters (orange blocks) are updated on the specific data domain with the task-specific losses.

Inference. After fine-tuning, we still keep the shared parameters frozen as in the previous fine-tuning state, and additionally load the weights of the extra parameters that were fine-tuned in the previous stage. The single overall model is able to be adapted to multiple tasks with the assistance of lightweight introduced modules.

3 Discussion

Tunable parameters analysis. Our AdaptMLP module is lightweight. The total number of parameters introduced to per layer is 2×d×d^+d^+d2\times d\times\hat{d}+\hat{d}+d, which includes biases parameters. The middle dimension d^\hat{d} is a small value compared with dd (AdaptFormer still obtains a decent performance even when d^=1\hat{d}=1, as discussed in Sec. 4.5). Since most of the shared parameters are fixed and the number of newly introduced parameters is small (<2%<2\% of the pre-trained model parameters), the total model size grows slowly when more downstream tasks are added.

Applicability. We note that AdaptMLP is a plug-and-play module that can be adaptively inserted into existing popular vision transformer architectures since all of the backbones share the same MLP layers even though they differ in the MHSA architectures (as shown in Figure 2(b)). Compared to our methods, we notice that recent prompt-related approaches insert trainable parameters into the token space, as illustrated in Figure 3. They prepend learnable parameters either into the embedded tokens before linear projection or the key and value tokens after linear projection . Therefore, the prompt-related method can not be straightforwardly adapted to special MHSA variants, especially for the one that takes the pyramid spatial information into account . Besides, we empirically observe that prompt-related methods perform not well when the number of patch tokens grows up from image to video scale, as shown in Figure 1.

In summary, we present a strategy for tuning a pre-trained vision Transformer on a set of scalable vision recognition tasks (e.g.image domain and video domain). It adds limited learnable parameters for tuning while achieving comparable or even better performance than the full-tuning strategy. Moreover, AdaptFormer could serve as a generic module for a large variety of recognition tasks.

Insights of architecture design. The MLP module is important for ViTs. As illustrated in , MLPs prevent ViTs from producing a rank-1 matrix. Also, MLPs stop the ViT output from degenerations. Inspired by the above analysis, we believe an effective ViT adaptation shall focus on its MLPs rather than multi-head self attentions. Meanwhile, we learn from the inception framework that parallel design is an effective way for feature ensemble. With the parallel design, the domain-specific features produced by the adapter module can supplement the domain-agnostic features from the fixed branch for a better feature ensemble. Our following experiments will verify that the parallel performs better than the sequential design.

Besides, though many advanced Transformer-based models which have emerged since the success of ViT having different attention mechanisms within the Transformer block, they all share the similar MLPs (feed-forward network) structures. Therefore, our AdaptMLP can be easily plugged into these ViT variants. Moreover, AdaptMLP can also be applied to more recent attention-free models .

Experiments

We evaluate the effectiveness of AdaptFormer by conducting extensive visual recognition experiments in both the image and video domains. We first describe our experimental settings in Sec. 4.1, covering the pre-trained backbones, baseline methods, downstream tasks and training details. We then compare AdaptFormer with baseline methods and provide a thorough analysis in Sec. 4.2. In addition, we also conduct ablation studies to explore different experimental configurations and explain what makes for the superiority of AdaptFormer in Sec 4.5.

Pre-trained backbone. We adopt the plain Vision Transformer (ViT) , i.e., ViT-Base (ViT-B/16) as our backbone model and pre-train the model with both supervised and self-supervised approaches. Specifically, for image, we directly use the ImageNet-21k supervised pre-trained modelhttps://github.com/rwightman/pytorch-image-models/releases/download/v0.1-vitjx/jx_vit_base_patch16_224_in21k-e5005f0a.pth and MAE self-supervised modelhttps://dl.fbaipublicfiles.com/mae/pretrain/mae_pretrain_vit_base.pth. For video, we take both supervised and self-supervised pre-trained models from VideoMAE . More details about pre-training approaches and datasets can be found in Appendix.

Initialization of AdaptFormer. For the original networks, we directly load the weights pre-trained on the upstream tasks and keep them frozen/untouched during the fine-tuning process. For the newly added modules, the weights of down-projection layers are initialized with Kaiming Normal , while the biases of the additional networks and the weights of the up-projection layers are configured with zero initialization. The reason for the zero initialization of other layers is that in this way, the initial newly added parameters are initialized such that the new function resembles the original one at the start of the fine-tuning stage. We empirically found that if the initialization deviates too far from the identity function, the model is not stable to train.

Baseline methods. We compare AdaptFormer with three commonly used fine-tuning approaches, including (1)Linear probing: adding an extra linear layer on top of the backbone and tuning the added parameters for evaluation. (2) Full Fine-tuning: setting all the parameters learnable and tuning them together. (3) Visual Prompt Tuning (VPT): fine-tuning the extra token parameters as shown in Figure 3.

Downstream tasks. We evaluate our AdaptFormer on both image and video recognition tasks to verify its effectiveness. The specific datasets leveraged in this work are presented in the following.

∙\bullet Image domain : CIFAR-100 contains 50,000 training images and 10,000 validation images of resolution 32×32 with 100 labels. Street View House Numbers (SVHN) is a digit classification benchmark dataset. In total, the dataset comprises over 600,000 labeled images, containing 73,257 training samples, 26,032 testing samples and 531,131 extra training data. The Food-101 dataset consists of 101 food categories with a total of 101k images, including 750 training and 250 testing samples per category.

∙\bullet Video domain : Something-Something V2 (SSv2) is a large collection of video clips showing the people perform several normal actions in the daily life (e.g., moving stuff and opening the door). It consists of 168,913 training samples, 24,777 validation samples and 27,157 testing samples, making a total of 220,847 videos with 174 labels. HMDB51 is composed of 6,849 videos with 51 categories, making a split of 3.5k/1.5k train/val videos.

Implementation details. In this work, we use PyTorch toolkit to conduct all experiments on NVIDIA V100 GPUs. Unless otherwise stated, we use 8×\times8 GPUs for video experiments and 1×\times8 GPUs for image experiments. Our default configurations follow the linear probing settings in , which do not utilize many common regularization strategies, such as mixup , cutmix , color jittering and so on. More details can be found in Appendix.

2 Main Properties and Analysis

We compare the performance of different fine-tuning approaches in Table 1 with the backbones pre-trained via the self-supervised paradigms. The results show that AdaptFormer consistently surpasses linear probing and Visual Prompt tuning (VPT) methods. Specifically, AdaptFormer-64 outperforms VPT on image benchmark CIFAR-100, SVHN, and Food-101, by 3.46%, 2.87%, and 4.63% respectively. On the more challenging video action recognition dataset Something-Something V2, the superiority becomes even more significant, i.e., about 15%. Note that even compared with the full fine-tuning strategy, our AdaptFormer still outperforms by about 5% Top-1 accuracy on SSv2 dataset. To summarize, our AdaptFormer is highly parameter-efficient, as well as yielding good performance with parameter size at most 2% times than the full fine-tuning manner.

3 Scaling Tunable Parameters Up

Even though there are only limited parameters introduced, one might also argue that more tunable parameters of AdaptFormer contribute to its higher accuracy compared with VPT . We conduct experiments to make a comprehensive discussion on this aspect.

As described in Sec. 3.3, the number of tunable parameters can be adjusted by changing the number of introduced tokens for VPT, or the hidden feature dimension for AdaptFormer. As shown in Figure 5, we conduct experiments with a wide range of tunable parameters on both SSv2 and HMDB-51 datasets. Since AdaptFormer and VPT share the same number of parameters of classification head on a specific dataset, we only report the tunable parameters on the x-axis, which comes from the visual prompts (VPT) or weight/bias of the down-up fully-connected layers (AdaptFormer), without calculating the parameters of classification head. For VPT, the number of introduced tokens is chosen from {1, 2, 4, 8, 16, 32, 48, 64}. Similarly, the number of hidden dimensions in AdaptFormer is in {1, 2, 4, 8, 16, 32}. AdaptFormer has a slight performance gain or maintains the accuracy stably when the parameters scale up. On the contrary, the performance of VPT decreases dramatically when the parameters exceed the task-specific value. Moreover, choosing the most suitable number of token number becomes laborious since it might be task-specific (i.e.varying from one dataset to the other one). For example, the accuracy of VPT keeps going up when the number of tunable parameters increases up to 300K on SSv2, whereas it begins to drop when the number of tunable parameters exceeds 50K on HMDB-51.

We further study the optimization procedures of VPT by monitoring the test accuracy of the training stage. As shown in Figure 5, we gradually increase the number of tokens in VPT and plot the Top-1 accuracy of each epoch. The training stages are stable when the number of tokens is less than or equal to 4, e.g., {1, 2, 4}. However, when the number becomes 8 or larger, e.g., {8, 16, 32}, the training procedure collapses at about the tenth epoch and achieves poor performance at the end of the training stage. On the contrary, the optimization procedures of AdaptFormer are stable when the number of parameters varies across a large range, as shown in Table LABEL:tab:hidden_dim. The top-1 accuracy fluctuates within 1.5% when the number of parameters increases from 0.44M (dim=16) to 4.87M (dim=256).

4 Multi-Label Classification

We further conduct experiments on dataset with larger scale and diversity. Specifically, we evaluate AdaptFormer on NUS-WIDE for multi-label classification. NUS-WIDE contains 269,648 images collected from Flicker, which are annotated with 81 visual concepts. Since some images are not available on Flicker, we only use 220,000 images following . We utilize mean average precision (mAP) as performance metric.

Settings and results. Our training settings mainly follow ASL . Specifically, We trained all models for 40 epochs using Adam optimize and 1-cycle learning rate policy . The maximal learning rate is 0.001. As shown in Table 2, though AdaptFormer-64 achieves a slightly lower mAP than fine-tuning, it significantly reduces the amount parameters that need to be updated (from 85.86 to 1.25M). Moreover, AdaptFormer has an clear advantage over other fine-tuning approaches including linear probing and VPT.

5 Ablation Studies

We ablate our AdaptFormer to study what properties make for a good AdaptFormer and observe several intriguing properties. The ablation studies conducted in this work are all performed on the SSv2 validation set .

Middle dimension. The middle dimension controls the number of introduced parameters by AdaptFormer. Lower middle dimensions introduce fewer parameters with a possible performance cost. We ablate AdaptFormer on the middle feature dimension to study this effects. As shown in Table LABEL:tab:hidden_dim, the accuracy consistently improves when the middle dimension increases up to 64 and reaches the saturation point when the middle dimension is about 64 on SSv2 dataset. We note that our AdaptFormer can achieve a decent performance when the middle dimension reduces even to one, about 50.03% top-1 accuracy.

We conduct more extensive ablation studies on middle dimension in Appendix Table 10 and found that the optimal middle dimension varies per dataset. For example, the accuracy reaches saturation when the middle dimension equals 64 on SSv2, whereas for NUS-WIDE dataset, the mAP slightly improves when the middle dimension increases from 64 to 512. However, AdaptFormer with middle dimension as 512 has 0.75 mAP higher (59.82 vs. 59.07 mAP) than the one with 64 at the cost of about 8 times more parameters. Therefore, we choose the middle dimension=64 for both SSv2 and NUS-WIDE for a better trade-off.

Scaling factor. The scaling factor ss is introduced to balance the task-agnostic features (generated by the original frozen branch) and the task-specific features (generated by the tunable bottleneck branch). We evaluate AdaptFormer with multiple ss values and the results are summarized in Table LABEL:tab:s_factor. Different from the scaling factor in NLP field which prefer ss larger than 1 (e.g., s=4s=4 in ), we empirically found that the ss should be <1<1 for vision tasks, otherwise the fine-tuning would become unstable. Besides, we found that AdaptFormer achieves optimal performance with s=0.1s=0.1. A larger or smaller ss would bring slight performance drop. Thus, we choose s=0.10s=0.10 as a default setting.

AdaptFormer position. As shown in Table LABEL:tab:insert_form, we further ablate on the specific position to introduce the AdaptMLP block. We gradually increase the number of AdaptMLP layers with a step of three (start →\rightarrow end, both included). We observe that the performance of AdaptFormer has a positive correlation with the number of added layers. In addition, AdaptFormer prefers the top part (the one far away from the input image) of the network to the bottom part when introducing the same number of layers, e.g., AdaptFormer with 7 →\rightarrow 12 obtains over 14.5% higher accuracy than 1 →\rightarrow 6, though both equipped with six AdaptMLP layers.

Insertion form. We study the insertion formulation by comparing the parallel and sequential instances which are illustrated in Figure 7. As shown in Table LABEL:tab:insert_form, the parallel AdaptFormer is able to outperform the sequential one by 0.85% top-1 accuracy. The reason might be: (1) the parallel design maintains the original feature using an independent branch and aggregating updated context by element-wise scaled sum; (2) the sequential design is equivalent to adding more layers, which might cause optimization difficulty. Therefore, we adopt the parallel design as our default setting due to its superiority.

Number of frames. The number of embedded patch tokens increases linearly with the number of video frames for the plain ViT . We conduct experiments with the different number of frames, i.e., {2, 4, 8} and the results are shown in Figure 7. We observe that increasing the number of frames is beneficial for all these three fine-tuning methods. However, AdaptFormer consistently outperforms the linear manner (e.g., +30% top-1 accuracy on 8 input frames) and VPT method(e.g., +14% top-1 accuracy on 8 input frames).

6 Towards Visual Recognition Generalist Agent

In the above experiments, we typically utilize a modality-specific pre-trained checkpoint for the corresponding downstream tasks. For example, we use Kinetics-400 (video domain) pre-trained model for downstream video action recognition on Something-Something V2 and HMDB-51 benchmarks. Besides, we use ImageNet-21K (image domain) pre-rained model for downstream image classification on CIFAR-100, SVHN and Food-101 benchmarks. Our AdaptFormer achieves superior performances in this same network with modality-specific weights scenario.

Next, we take a further step to ask what would happen if using the same network with the modality-agnostic weights for multiple tasks in the multi-modalities downstream tasks?

We use the model pre-trained on ImagNet-21k to do action recognition on SSv2. As shown in Table 4, AdaptFormer is robust to domain shift caused by modality. The experimental results show that the linear probe approach obtains a very poor accuracy (i.e., 6.56% top-1 accuracy) when fine-tuning on SSv2. Meanwhile, VPT achieves a better performance than linear probe but it is not decent (i.e., 16.94% top-1 accuracy). Our AdaptFormer, compared to the above two methods, attains a promising 46.06% top-1 accuracy, which is even higher than the full-tuning schedule (+4.56%).

7 Visualization

To evaluate the quality of the produced features, we conduct t-SNE visualizations on AdaptFormer and other baseline methods. The features are extracted from the SSv2 validation set via the ViT-Base backbone. Figure 8 shows that the linear fine-tuning and the VPT methods tend to output mixed features as shown in Figure 8(a)-(b). Compared with the above two methods, the full fine-tuning strategy performs well in projecting features. However, it consumes huge computational sources to tune the whole network parameters. Figure 8(d) validates that our AdaptFormer facilitates ViT-Base in generating more separable representations with fewer learnable parameters.

Conclusion

We present a conceptually simple yet effective framework, AdaptFormer, for efficiently adapting a pre-trained Vision Transformer (ViT) backbone to scalable vision recognition tasks. By introducing AdaptMLP, our AdaptFormer is able to fine-tune the lightweight modules for producing features adapted to multiple downstream tasks. The extensive experiments on five datasets, covering both the image and the video domains, validate that our proposed methods are able to increase the ViT’s transferability with little computational cost. We hope our work will inspire future research in exploring more efficient fine-tuning methods for large vision models. One limitation is that AdaptFormer is only employed in recognition tasks in this work, it’s unclear whether it can work well in tasks beyond recognition, e.g., object detection and semantic segmentation. We leave it for the future exploration. Since our method is specially designed for efficient fine-tuning, we do not foresee obvious undesirable ethical/social impacts at this moment.

Acknowledgment. This work is supported by CCF-Tencent Open Fund. Ping Luo is supported by the General Research Fund of HK No.27208720, No.17212120, and No.17200622.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [Yes] Shown in Conclusion Section.

Did you discuss any potential negative societal impacts of your work? [Yes] Shown in Conclusion Section.

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [N/A]

Did you include complete proofs of all theoretical results? [N/A]

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] As a URL shown in the abstract.

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] Shown in supplementary materials.

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [N/A]

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] Please see Section 4.1

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes]

Did you mention the license of the assets? [Yes] Shown in supplementary materials.

Did you include any new assets either in the supplemental material or as a URL? [No]

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [Yes] We used publicly available datasets whose licenses allow research usage.

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [No] To the best of our knowledge, the data we used contains no personally identifiable information or offensive content.

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]

Appendix A Appendix

In this supplementary material, we will include the details about the pre-training and fine-tuning processes, the extensive experiments of AdaptFormer on hierarchical vision transformers (e.g., AdaptFormer-Swin), and the pseudo-code of AdaptMLP in a PyTorch-like style.

Image. We use MAE as our self-supervised pre-training method in the image domain, a simple yet effective method that first masks nearly 75% patches of the input image and then reconstructs the missing pixels. Specifically, we directly adopt the checkpointhttps://dl.fbaipublicfiles.com/mae/pretrain/mae_pretrain_vit_base.pthof ViT-B/16 for convenience, which is pre-trained on ImageNet-1K for 800 epochs.

Video. We use VideoMAE as our self-supervised pre-training method in the video domain, which is an direct extension of MAE to the video domain. VideoMAE utilizes the plain ViT architecture of joint space-time attention mechanism and an extremely high proportion of masking ratio (i.e., 90% to 95%) for pre-training. We also directly use the publicly available checkpointhttps://drive.google.com/file/d/1tEhLyskjb755TJ65ptsrafUG2llSwQE1/view?usp=sharing, which is pre-trained on Kinetics-400 .

A.1.2 Implementation Details of Fine-tuning

The implementation details are summarized in Table 5. The video experiments are conducted on 64 Tesla V100 GPUs, while the image experiments are performed on 8 Tesla V100 GPUs. For the optimizer, different from that adopts LARS, we leverage SGD for stable training on small-scale dataset (e.g., CIFAR10). The actual learning rate is calculated by: lr = base_lr×\timesbatchsize / 256 following the linear lr scaling rule . More detailed training configurations are presented in Table 5, including the batchsize, learning rate schedule and etc.

The experimental settings of image and video mainly follow the ones utilized in MAE and VideoMAE , respectively. We insert an extra BatchNorm layer without affine transformation (i.e. affine=False) before the final fully connected layer, following the common practice to normalize the pre-trained features . In addition, there is no flip augmentation during the fine-tuning stage for video data.

A.2 More Supplementary Results

In addition to the self-supervised pre-training presented in the main paper, we also evaluate AdaptFormer with the supervised pre-trained model. The results in Table 6 show that AdaptFormer still outperforms linear probe and VPT obviously. In addition, AdaptFormer surpasses full-tuning on four benchmarks (CIFAR100, SVHN, SSv2, HMDB51) with only 1.46% parameters. On the remaining benchmark (Food-101), AdaptFormer achieves an almost comparable performance to full-tuning (90.89% v.s. 90.96%).

A.2.2 AdaptFormer on Swin Transformer

Settings. We further demonstrate the effectiveness of AdaptFormer on hierarchical vision transformers, e.g., Swin . We name AdaptFormer applied to Swin as AdaptFormer-Swin, to distinguish plain AdaptFormer (without any suffix) which is applied to the vanilla ViT . It is noted that we can adopt AdaptMLP to Swin easily without any special modification as Swin and ViT share the same MLP architecture. However, VPT needs additional handcraft designs to be suitable for the shifted local windows in the prevalent hierarchical vision transformers, which hinders its general applications.

We utilize Swin-B and the video counterpart for image and video. Similarly, we also directly use the officially provided checkpointsImage: https://github.com/SwinTransformer/storage/releases/download/v1.0.4/swin_base_patch244_window877_kinetics600_22k.pth Video: https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window7_224_22k.pth, which are pre-trained on ImageNet-21K and Kinetics-600 .

Results. Since VPT is not applicable in Swin, we do not report its performance. Table 7 shows AdaptFormer-Swin performs well compared with other tuning strategies. For image benchmarks, our method can outperform full-tuning approach with only 1.43% parameters. Moreover, AdaptFormer-Swin surpasses linear probing by a significant margin, especially on the challenging dataset, SSv2. The results validate that AdaptFormer is able to generally boost the transferability of various vision Transformer variants.

A.3 Possible Architectures

We explore other possible architectures utilized in AdaptFormer. Specifically, we further replace the MLP architectures within the AdaptMLP module by the convolution layer, depthwise convolution layer, and LayerNorm layer. For fair comparisons, we carefully design the above modules to meet the comparable number of parameters (1̃.3M). The experimental results of different adapter modules are shown in Table 8, which validates that the simple MLP modules are simple yet effective compared with the other architectures. For example, our AdaptMLP module surpasses the AdaptConv module by 0.55% Top1 accuracy on SSv2 dataset.

A.4 Evaluation on ImageNet-1k datasets

We point out that in order to evaluate the adaptation performance across datasets, it’s an unreasonable setting to fine-tune the ImageNet-1k dataset with the ImageNet-21k pre-trained weights. This is because ImageNet-1K is a subset of the ImageNet-21K as introduced in . In contrast, in all the previous experiments, there is no overlap between the fine-tuning and pre-trained datasets. However, we document the experiments of fine-tuning with the ImageNet-1k dataset for the completeness.

We adopt exactly identical training configurations to conduct experiments in this subsection. We experiment with middle dimension = {1, 4, 16, 64} on ImageNet-1K, and the results are shown Table 9.

Results. Comparing the results of AdaptFormer with different middle dimension ({1, 4, 16, 64}) on ImageNet-1K, we find that AdaptFormer with the smallest number of parameters (AdaptFormer-1) achieves the best top-1 accuracy (82.33%). Furthermore, when the ‘middle dimension‘ increases from 1 to 4 or 16, AdaptFormer has a slight performance drop (AdaptFormer-4 (-0.07%) and AdaptFormer-16 (-0.09%)). Further increasing the middle dimension to 64 will cause a relatively clear performance drop (-0.47%).

Discussions. Although our AdaptFormer-64 does not have a clear advantage compared with VPT , our AdaptFormer-1 outperforms VPT by +0.65% top-1 accuracy with only 0.02M additional parameters. Besides, the trend of classification accuracy changing with middle dimension on ImageNet-1k is different from other datasets in our paper, e.g., AdaptFormer with middle dimension=64 achieves better top-1 accuracy than with middle dimension=1 on CIFAR-100. We empirically find introducing a small number of parameters (AdaptFormer-1) is sufficient for ImageNet-1K fine-tuning, while more introduced parameters will make a larger change to the original model and make it harder for optimization since ImageNet-1K is a subset of ImageNet-21K. However, for other datasets (e.g., CIFAR-100) with no overlap between the fine-tuning datasets and the pre-trained datasets, more introduced parameters are needed for learning better domain knowledge.

A.5 Extended experiments on middle dimension

We conduct the extended ablation studies on the middle dimension design in this sub-section. We aim to seek for a trade-off between model capacity (i.e., potential) and adaptation efficiency. In fact, the middle dimension has a main influence on the parameter size of adapter. The higher dimension brings more parameters while the efficiency and storage are limited. As shown in Table 10, we evaluate several numbers of middle dimension and found that using 64 is optimal to achieve accuracy, light-weight storage, and efficiency.

A.6 Analysis on the fine-tuning time and inference latency

To analysis the computational efficiency, we compare the fine-tuning time and inference time on a single NVIDIA A100-40G GPU. We utilize SSv2 video classification for this part. For fine-tuing, we experiment with batchsize of 32. For inference, we test the latency with multiple batch sizes to get a comprehensive comparison under various inference scenarios. All the time is measured in milliseconds averaged over 100 trials. The results are summarized in Table 11 and Table 12. As shown in Table 11, AdaptFormer only costs less than a half of the fine-tuning time compared with the full-tuning. Moreover, AdaptFormer significantly outperforms linear probing in terms of accuracy with a slight longer fine-tuning time. For inference, AdaptFormer introduce a negligible FLOPs and latency compared with the Linear/Full-tuning.

A.7 Discussion about ImageNet and Kinetics Pre-training

The type of spatiotemporal attention (divided vs. joint) determines whether the performance of the model pre-trained on ImageNet can outperform the model pre-trained on Kinetics.

A similar phenomenon has been discussed in recent work, Uniformer , independently. We borrow the experimental results from Table 4(c) in Uniformer paper . Specifically, the divided spatiotemporal attention prefers ImageNet to Kinetics-400 for the pre-training dataset. The performance of the divided attention model pre-trained on ImageNet outperforms the model pre-trained on Kinetics-400. On the contrary, the joint spatiotemporal attention prefers Kinetics-400 to ImgaeNet. The joint attention model attains higher top-1 accuracy with Kinetics-400 pretraining compared to ImageNet (53.8 vs. 52.0).

We adopt the joint spatiotemporal attention for all video-related experiments in this work (introduced in Appendix A.1.1). Therefore, our experimental phenomenon is consistent with the joint attention in , i.e., Kinetics pretraining is preferable.

A.8 Implementation

The core part of AdaptFormer is replacing the original MLP with AdaptMLP, which consists of the frozen original MLP and newly introduced Down →\rightarrow ReLU →\rightarrow Up layers, which are tunable at the fine-tuning stage. Algorithms 1 provides the implementation of AdaptMLP written in PyTorch .

For more implementation details, please refer to the provided source code.