ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Junting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao, Hongsheng Li

Introduction

In the NLP field, almost all the state-of-arts across a wide range of downstream tasks have been achieved by adapting from large pretrained models (a.k.a. foundation models ) such as BERT and GPT . The de facto standard approach to adapting a pretrained model to down-stream tasks is fine-tuning either fully or partially (e.g., linear probing by training the newly added multi-layer perceptron layers on the top alone), subject to the condition of adopting a similar network architecture as the pretrained model. Nonetheless, given increasingly larger whilst ever stronger foundation models (e.g., GPT-3 with 175B parameters), fully fine-tuning the whole model for every single downstream task would become prohibitively expensive and infeasible in terms of training cost and model storage. This could significantly restrict their deployment and usability in real-world applications. In this context, a series of NLP works has been introduced towards efficient transfer learning with better trade-offs between parameter and accuracy .

This trend has recently motivated the computer vision community. For example, the CLIP model , trained with 400 million web image-text pairs, achieves promising performances on a variety of image recognition and generation tasks. In the video domain, with significantly more computational cost and resources, Xu et al. trained a video variant of CLIP but excelled on a smaller number of downstream video tasks. This is partly attributed to two orders of magnitude more minor training data and limited availability of computing resources, as large video data is notoriously more difficult to collect, manage, and process than image data. Under these restrictions, large pre-trained image models are arguably still favorable in the selection of model initialization for video tasks.

In this work, we investigate a novel, critical problem of efficiently adapting large pre-trained image models for video downstream tasks, with a focus on the widely influential action recognition task. Considering that training video models is drastically more expensive in both computing resource and time than image models , this problem becomes particularly more useful and valuable in practice. On the other hand, it is also more challenging and non-trivial due to the extra necessity of overcoming the big gap between image and video in transfer learning. Especially, pre-trained image models lack the ability to infer temporal structured information, which however is critical in video understanding. In fact, the key design with state-of-the-art video models is usually about learning the temporal dimension based on contemporary image models. Although model initialization is still important, they largely go beyond the fine-tuning strategy, as architectural modification is often imposed in addition to full model training/fine-tuning per downstream task.

Given that this is a new problem, we first conduct a comprehensive benchmark using both various fine-tuning methods for image-to-video transfer learning and state-of-the-art video models . Regarding the pretrained image model, we select two Vision Transformer (ViT) models, with one from CLIP pre-training and the other pre-trained on ImageNet-21K . ViT is representative in terms of network architecture, pre-training algorithm, and training data scale. Crucially, we further propose an efficient yet effective Space-Time Adapter (ST-Adapter), capable of extracting and leveraging the pre-trained knowledge of a large image model to achieve superior video understanding at a small parameter cost. Specifically, ST-Adapter is formulated based on a novel parameter-efficient bottleneck with a sequence of operations including feature dimension reduction, spatial-temporal modeling, and feature dimension recovery. It is easy to implement and scalable for deployment since all the primitive steps are realized with standard operators (e.g., fully-connected layer, depth-wise 3D convolution). With such a lightweight design, our bottleneck can be cheaply integrated throughout the base network for enabling stronger layer-wise spatio-temporal learning. As a result, our model can be more rapidly optimized using fewer training epochs for significant convergence advantage.

We summarize the contributions as follows. (1) We investigate a new problem of parameter-efficient image-to-video transfer learning. Our motivation is to advocate the usability and deployment of increasingly larger whilst ever more powerful pre-trained image models in benefiting more challenging video understanding tasks. (2) We establish a benchmark for action recognition tasks by comprehensively experimenting with a variety of fine-tuning strategies and several state-of-the-art video understanding models. (3) We introduce a novel parameter-efficient Spatio-Temporal Adapter (ST-Adapter) for more effectively capitalizing a large pre-trained image model in video understanding. By grounding all the primitives on standard operators, ST-Adaptor is easy to implement and friendly to deployment. (4) Extensive experiments on action recognition datasets show that our ST-Adapter outperforms not only existing parameter-efficient alternatives and the full fine-tuning strategy, but also state-of-the-art video methods with the same network architecture and model initialization.

Related Work

Driven by the wider application of large pre-trained language models across a diversity of downstream tasks, the topic of efficient tuning has received increasing attention in NLP. Existing efficient tuning methods fall broadly into three categories. The first category is to introduce task-specific adapters . Specifically, an adapter consists of lightweight modules inserted between layers of a pre-trained model. To be parameter-efficient, only those newly added adapter modules need to be updated during task fine-tuning, whilst all the parameters of the large pre-trained model, which takes the majority proportion of the whole solution, are frozen. The second category is prompt tuning . Instead of manipulating the network architecture, these methods prepend a set of learnable tokens at the input point of the model or intermediate layers. Similarly, only these added tokens need to be optimized for each downstream task. The third category is learning weight approximation . In particular, only the low-rank matrices for approximating the weights need to be updated during training.

Early works for efficient transfer learning in vision focus on parameter sharing in the context of multitask learning . Recently, there are several works for extending the efficient tuning idea from NLP to vision tasks. CoOp and CoCoOp apply prefix tuning for adapting the CLIP model to various image recognition tasks. VL-Adapter achieves the performance comparable to full fine-tuning on challenging vision-language tasks. Commonly, their design focuses are all restricted to the text encoder of the CLIP model. More recently, introduce the idea of prompt learning to visual backbones. They obtained favorable results on various image recognition benchmarks. Moving a step further, in this work, we consider the more challenging adaptation problem from a pre-trained image model without temporal knowledge to video understanding tasks.

Action recognition in the unconstrained video has largely been dominated by deep learning methods, thanks to the availability of large video datasets, e.g., Kinetics and Something-Something . As a key component, the model architectures adopted by existing video methods has expanded from CNNs to Transformers . As temporal information is important for modeling the dynamics, a variety of motion learning techniques has been introduced . Further, different training methods have also been explored, e.g., unsupervised learning , and video-text contrastive learning . New opportunities for stronger video models are created following the introduction of large pretrained foundation models . For example, Wang et al. equipped the CLIP with temporal modules and good performance can be achieved after the model is fully fine-tuned on video datasets. Ju et al. adopted the CLIP model for video recognition tasks by learning video-specific prompts. In contrast, in this work, we explore the potential of the large pre-trained image models with the parameter-efficient adapter strategy. Importantly, despite the simplicity, we bring about more significant advantages in performance along with a new benchmark on parameter-efficient image-to-video transfer learning.

Methodology

To capitalize a large pre-trained image model for more challenging video understanding such as action recognition in a cross-modality manner, it is necessary to fill the intrinsic gap between image and video. For easier understanding, we start with an intuitive baseline based on temporal aggregation.

For more dedicated structural modeling in the time dimension with ViTs, a mainstream approach in the video domain is to develop various spatio-temporal attention mechanisms by further imposing temporal attention on top . We choose two representative video ViT models, TimeSformer and XViT , in our performance benchmark. However, state-of-the-art video ViT models often need to fully fine-tuned per task, which is parameter-inefficient, given that in this way we have to keep a separate copy of the whole fine-tuned model parameters for every single task.

We aim to propagate the success of Adapter from NLP to computer vision particularly the image-to-video transfer learning problem as discussed earlier. To that end, we introduce a novel Adapter tailored specially for spatio-temporal reasoning – a key capability for video understanding which, however, existing NLP Adapter variants lack.

2 Spatio-Temporal Adapter (ST-Adapter)

Typically, an image model only considers the ability of spatial modeling. The objective of our Spatio-Temporal Adapter (ST-Adapter) is to enable a pre-trained image model to reason about spatial and temporal information of video in a parameter efficient principle. In design, we consider a couple of practically-crucial criteria: (1) Smaller parameter size: The parameter cost for each downstream task should be small – the essential criterion for parameter efficiency. (2) Development friendliness: This is critical for real-world development and deployment. In practice, it is necessary that a model can be easily implemented using the standard highly optimized deep learning toolboxes (e.g., PyTorch, TensorFlow, TensorRT, and TorchScript), without tedious per-toolbox specialization. This also facilitates the realization of high inference efficiency across a diversity of running platforms due to the best usage of built-in software and hardware resources.

Under these considerations, we formulate the proposed ST-Adapter by sticking to commonly-adopted primitive operators alone. Starting with the above Adapter (Eq. (1)) originally developed for NLP tasks, we further introduce a spatio-temporal operator realized by a standard depth-wise 3D-convolution layer between the bottlenecks (Figure 1). In particular, our spatio-temporal operator enables layer-wise temporal inference efficiently, because it only operates in a compressed low-dimensional (e.g., 128D) feature space and the depth-wise convolution is highly efficient both in parameter and computation . As a result, this yields an introduction of tiny extra (∼\sim2%) parameters and (∼\sim0.3%) computation. Formally, our ST-Adapter can be expressed as:

3 ST-Adapter Integration

For proper adaptation, the adapter modules are often integrated between layers of a Transformer. In NLP, a variety of integrating designs have been investigated. For example, deploys two adapter modules per layer with one following the Multi-Head Self-Attention (MHSA) and the other following the Feed-Forward Networks (FFN) . On the other hand, suggest that adding only one adapter after the FNN suffices. Similarly, our ST-Adapter can be also integrated generally at distinctive positions. Empirically, we find that a decent performance can be achieved in case a single ST-Adapter is placed before the MHSA of each transformer block (Figure 1(a) and Table 5c).

Experiments

For the benchmark experiments, we use two popular video action recognition datasets.

Kinetics-400 (K400): The K400 dataset contains ∼\sim240k training videos and 20k validation videos labeled with 400 action categories. Most videos have a length of 10s or about 300 frames. While there is a great diversity in these videos, they are largely biased to spatial appearance .

Something-Something-v2 (SSv2): The SSv2 dataset consists of 220,487 videos covering 174 human actions. The video length ranges from 2 to 6 seconds. In contrast to K400, SSv2 presents richer temporal information with much higher significance .

Epic-Kitchens-100 (EK100): The EK100 dataset consists of 100 hours of video in egocentric perspective recording a person interacting with a variety of objects in the kitchen. Each video sample is labeled with a verb and a noun. We report top-1 verb and noun classification accuracy.

In all experiments, we use the standard ViT as our base backbone model. We conduct most of our experiments with the ViT-B/16 variant with 12 layers and 86M parameters, taking as input a sequence of patches at size 16×1616\times 16.

What was learned during pre-training directly decides the knowledge that can be transferred to downstream tasks, thus also the effectiveness upper bound of transfer learning methods. To this end, we benchmark the same backbone under two different pre-training strategies: pre-training with web-scale raw data that has been recently proposed by CLIP (400M image-text pair) and classical supervised pre-training on annotated data from ImageNet-21K (21k classes and 14M images).

All details, including training and testing settings and module instantiation details, are provided in the appendix.

We provide several transfer learning approaches in our benchmark for efficient image-to-video transfer learning. Note that the parameters of the linear classifier are always updated during training for all approaches.

Full Fine-tuning: Fully updating all the parameters when adapting for a specific target task.

Partial Fine-tuning: Only update the last ViT layer while keeping the rest of the parameter fixed.

Temporal Fine-tuning: We only tune the temporal attention modules (i.e., TA) in the SA+TA architecture.

Linear Probing: Freezing all the parameters except those in the linear classification layer.

Adapter : Adding small sub-networks between layers of a pre-trained model. During fine-tuning, we only update the newly added parameters introduced by the adapters.

Prompt Tuning : Prepending a sequence of learnable prompt tokens to the input visual patch tokens. During fine-tuning, only these newly added prompts are updated.

Attention Pooling Head: Replacing the original temporal average pooling with a temporal attention pooling layer (similar to the one used in ) before the classification head.

These approaches above do not incorporate temporal modeling to the image ViT. Hence, we further consider temporally augmented ViT architectures as introduced in state-of-the-art video methods:

Spatial Attention Only (SA): Space-Only TimeSformer .

Spatial Attention + Temporal Attention (SA+TA): The default TimeSformer with divided space-time attention (Fig. 1a).

Spatial Attention + Temporal Shift (SA+TS): XViT .

Note that not all fine-tuning protocols are compatible with each of these video ViT variants. Take SA+TS for example, the original model behavior is altered with channel shift, as a result, it is not compatible with Linear Probing that requires freezing all the parameters of the backbone.

2 Main Results and Analysis

Table 1 presents the results of fine-tuning a ViT-B/16 pre-trained with CLIP and ImageNet-21K. All baselines are built by combining existing efficient fine-tuning methods with three state-of-the-art ViT-based action recognition models. From the results we can see that:

(i) For CLIP pre-trained model, ST-Adapter performs on par with Full Fine-tuning (82.0 vs. 81.7 for K400 and 66.3 vs. 66.1 for SSv2) while updating far less parameters (7.2M vs. 121.57M). ST-Adapter significantly outperforms all other efficient fine-tuning methods. We see that baselines like Prompt Tuning and Partial Fine-tuning can provide non-trivial gain in performance compared to Linear Probe, but are still behind our ST-Adapter.

(ii) Our ST-Adapter can generalize across different pre-training datasets and methods. We can see that CLIP pre-train models dominate over ImageNet-21K pre-train ones. These results well match the shift of paradigm in current AI research , where pre-training no longer needs limiting to curated data and annotations to deliver good performance on downstream tasks, but can take advantage of broader scale web raw data.

Interestingly, we observe that SSv2, a motion-centric dataset in design, also benefits from stronger appearance (image) pre-training. We think this may attribute to that raw textual description can provide a much richer description (i.e., human-object relations) of the image than curated limited categorical labels. Full fine-tuning on SA+TS (XViT) performs slightly worse with CLIP pretrain than ImageNet-21k pretrain. We conjecture this is because the channel shift operation breaks the knowledge in the pre-training weights, and thus does not benefit much from stronger pre-training like CLIP.

Comparison to the state-of-the-art models.

We compare ViT with ST-Adapter to other state-of-the-arts methods on both K400 dataset in Table 2, SSv2 dataset in Table 3 and EK100 dataset in Table 4. We can observe that:

(i) With the proper adaptation method, we can simply turn a large image foundation model into a good video model by only tuning a few parameters. Our results are comparable to or better than previous methods tailored for such tasks. Our largest model with ViT-L backbone set a new state-of-the-art in K400 by achieving 86.7% top-1 accuracy.

(ii) It is noteworthy that, our method takes significantly fewer frames as input compared to other methods (8 vs. 16, 32, 64, 96). It is also reflected in terms of GFlops. Saying that the ViT was not designed for efficiency purposes like but the adapted CLIP ViT has achieved similar accuracy-efficiency trade-offs.

(iii) The paradigm of pre-training and fine-tuning has been widely adopted in most state-of-art methods to achieve good performance. Between them, most of the approaches start from image pre-trained models, and only a few can afford video pre-training. Note that for the Something-Something dataset, except MViT pre-trained on video data from scratch, the rest of methods are still initialized from image pre-trained weights. A good image pre-trained model with rich appearance information can facilitate temporal modeling in temporally challenging datasets like SSv2.

(iii) It is evident in Table 4 that our ST-Adapter consistently brings a big margin on egocentric videos. Also, we found that without our ST-Adapter, it is much more difficult to directly adapt CLIP pre-trained ViT on the domain of egocentric video with high sensitivity to the hyper-parameter setting. ST-Adapter eases the training process. It is worthy to note that, all current transformer based approaches need to be pre-trained first on image dataset and then fine-tuned on Kinetics dataset before fine-tuned with egocentric videos. In contrast, our ST-Adapter can be directly applied to an image model and trained with target egocentric video alone.

3 Ablations

Unless otherwise specified, we use ViT-B/16 backbone and 8 input frames in all ablation experiments, and we use one ST-Adapter with bottleneck width 384 before MHSA in each Transformer block.

By default, we insert a ST-Adapter to every Transformer block in the backbone, but we also show the performance impact of using fewer ST-Adapters. As shown in Table 5b, while more ST-Adapters tend to do better, ST-Adapters at deeper layers boost performance more than those at shallower layers. This observation is useful when we insert ST-Adapters into deeper models and having an Adapter for each block might be too expensive. We also show the performance when inserting ST-Adapters to different positions within a block. As shown in Table 5c, while the performance is relatively insensitive to the position of the Adapters, using multiple adapters in one block may substantially boost performance on some datasets, like SSv2 in our case.

We experiment with a different number of channels in the middle of the bottleneck design. As shown in Table 5a and Fig. 2a, our method is effective with a wide range of bottleneck width: even with a channel reduction to 64, our ST-Adapters still obtain relatively good performance, outperforming all baselines in Table 1 except for Full Fine-tuning (SA + TA). Even with a bottleneck width of 768, our ST-Adapters are still very parameter efficient, introducing only about 1/6 new parameters to a Transformer encoder block. In contrast to the inverted bottleneck design commonly used with depthwise convolutions , ST-Adapters work best with regular bottlenecks. The success of transfer learning with such low-rank projections again shows the rich knowledge and strong potential of modern foundation models.

In Fig. 2b we show an enlarged difference between full fine-tuned models and our ST-Adapters with low training budgets. When we reduce the number of training steps, the accuracy of full fine-tuned models drops significantly faster than models with ST-Adapters. This shows the advantage of our proposed modules when backbone models are large or computational resources are limited. We also report the total training GPU-hours and peak memory usage for three models: TimeSformer, ViT-B/16, ViT-B/16 with ST-Adapter (8 input frames, 16 samples per GPU on 8 V100 GPUs) in Table 6.

Fig. 2c showcases the impact of training data size on action recognition accuracy. Even with the same pre-trained weights, ST-Adapters tend to obtain higher performance than full fine-tuning especially on smaller datasets: the margin between the two models increases with the shrinkage of data. This shows that ST-Adapters are powerful tools to transfer to downstream tasks where only a small amount of labeled data is available.

We ablate the effect of kernel size in the depth-wise convolutions inside our proposed ST-Adapter. It is shown in Table 7 that the temporal span is most sensitive, suggesting the significance of temporal structural modeling as we focus on in this work.

Conclusions

In this work, we have presented a simple yet effective Spatio-Temporal Adapter (ST-Adapter) for enabling a less studied parameter-efficient image-to-video transfer learning. Fully using commonly adopt primitive operators, ST-Adapter is particularly designed to be both lightweight and easy to implement for friendly usability and deployment. This cross-modality adaptation is a practically critical capability considering that it is dramatically challenging and more costly to build a sufficiently strong large video model in reality. Encouragingly, extensive experiments on video action recognition show that our ST-Adapter can match or surpass both the full fine-tuning strategy as well as fully trained state-of-the-art video models, whilst having the benefit of (20 times less updated parameters) parameter-efficiency. Further, our method is also faster to train and consumes less computing resources with economic and environmental superiority. We believe this work is inspiring for the research of other video understanding tasks such as action localization and video summarization.

This work is supported in part by Centre for Perceptual and Interactive Intelligence Limited, in part by the General Research Fund through the Research Grants Council of Hong Kong under Grants (Nos. 14204021, 14207319).

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [No]

Did you discuss any potential negative societal impacts of your work? [No]

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [N/A]

Did you include complete proofs‘ of all theoretical results? [N/A]

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] Code will be provided on GitHub after blind review.

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] See A

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [No] The experiments are too expensive to repeat many times.

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] See A

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes] All are mentioned in 4

Did you mention the license of the assets? [No]

Did you include any new assets either in the supplemental material or as a URL? [Yes] Code will be provided on GitHub after blind review.

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [No]

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [No]

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]

Appendix A Appendix

All experiments are implemented in PyTorch . We use the configuration listed in Tab. 8 unless otherwise specified. In general, we use much simpler data augmentation techniques compared to end-to-end fine-tuning. Hyper-parameters were briefly tuned to ensure convergence on a 20% held-out validation set.

The training configuration used for all the baselines is summarized as follows:

Full Fine-tuning: we largely follow the training configuration provided in their original paper, except that we train all the CLIP initialized layers with 1/100 learning rate and weight decay. We found these changes are necessary to obtain reasonable results for CLIP pretrained models; Otherwise the accuracy on Kinetics-400 is less than 50%. We found 1/100 to be the best scaling among {1/10,1/100,1/1000}\{1/10,1/100,1/1000\} on Kinetics-400.

Partial Fine-tuning: we finetune only the last Transformer block and the classifier layer. For the SA+TA architecture, TA is only added to the last block since the previous blocks need to be frozen in a meaningful state. We use the identical training configuration as provided in the original paper (i.e., without reduction of learning rate or weight decay for any trainable weight) as we found it slightly improves accuracy for this baseline.

Other baselines use the same training configuration as our proposed method, as stated in the Implementation details section in the main manuscript.

It is observed that with the same model, CLIP pre-training is superior to ImageNet21K pre-training (not surprising due to the training data scale and richness difference). However, our main objective is to propose a parameter-efficient fine-tuning alternative to the standard full fine-tuning approach particularly for image-to-video adaptation. To that end, we have validated the effectiveness and efficiency of turning an image foundation model into strong video action recognition models by tuning only a small fraction of parameters, in comparison to previous state-of-the-art alternatives.

By reporting the results on two different pre-training datasets (i.e., ImageNet21K and CLIP datasets), we would like to demonstrate that our ST-Adapter can generalize across different pre-training datasets and methods. Moreover, it can shed light on the difference between a foundation model (pre-trained with noisy web-scale raw data) and an ImageNet pre-trained model (which has been standard pre-training over the last decade).

To further support our finding, we have also experimented with the latest SWAG foundation model. As seen in Table 9, our ST-Adapter with a SWAG model can achieve consistent results as with a CLIP model: Reaching similar accuracy in the same tendency whilst outperforming the strong full fine-tuning strategy on both action datasets.

Experiments on additional backbone architectures

we have additionally provided the results of ST-Adapters on Swin-B models in Table 10. The results of Swin space only and Swin joint attention are obtained with the training configure of but using (8 frames x 3 views) sampling setting. Although they are not directly comparable with the results reported in (32 frames ×\times 12 views for K400, 32 frames ×\times 3 views for SSv2), they are highly indicative within reasonable range. It is expected that on ImageNet-21k pretrained models our ST-Adapters underperforms full fine-tuning, especially when the locality inductive bias of Swin makes tuning on the downstream tasks easier. However, our ST-Adapter still exhibits strong temporal learning capability, matching the joint-attention Swin and outperforms space-only Swin by a large margin. Also, we observe higher data efficiency with our ST-Adapter: The Swin joint attention model on the SSv2 dataset relies on K400 pretraining (directly fine-tuning from ImageNet-21k results in slightly less than 60% accuracy). In contrast, Swin w/ ST-Adapter achieves 65.1% even when directly trained from ImageNet-21k weights. Note, we primarily aim at adapting foundation image models pretrained on larger datasets (e.g., CLIP) other than ImageNet-21k.

We provide an inference speed test in Table 11. We measure the latency at batch size = 1 and throughput at batch size = 32. It is shown that our model performs slightly lower than TimeSformer space only, indicating that just a small overhead is introduced in inference speed by ST-Adapter.

We verify our method on two additional smaller but also widely studied video recognition datasets, namely UCF-101 and HMDB-51 . For both cases, we finetune from a Kinetics-400 pretrained model, with all CLIP layers fixed and ST-Adapters set to 1/10 learning rate and weight decay, and train for 500 steps with a batch size of 128. Frames are sampled with a temporal stride of 8. All other training settings are identical to that used for Kinetics-400. For testing, we use 3 spatial views and 2 temporal views, and report the 3-split mean accuracy for both datasets. We compare with methods that take only RGB frames as input (without optical flow). The results are shown in Table 12. We observe similar top performance by our ST-Adapter in comparison to recent state-of-the-art competitors, including the latest CLIP based VideoPrompt by a large margin.

We provide qualitative results about the attention map change before and after adding the ST-Adapters in Fig. 3. Videos are sampled from Something-Something-v2 dataset and the attention map of the [CLS] token from the last Transformer block is shown. The visualization shows that with ST-Adapters, the model attends more to action related regions (e.g., hands, fore-ground objects or moving objects), while the CLIP model without adaptation tend to be distracted by the background.