Visual Prompt Tuning

Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, Ser-Nam Lim

Introduction

For a variety of recognition applications, the most accurate results are now obtained by adapting large foundation models pre-trained on massive curated or raw data, a finding that mirrors developments in natural language processing (NLP) .As pointed out in , all state-of-the-art models in contemporary NLP are now powered by a few Transformer-based models (e.g., BERT , T5 , BART , GPT-3 ) This also applies to vision-language field recently, i.e., CLIP . At first glance,this is a success story: one can make rapid progress on multiple recognition problems simply by leveraging the latest and greatest foundation model. In practice, however, adapting these large models to downstream tasks presents its own challenges. The most obvious (and often the most effective) adaptation strategy is full fine-tuning of the pre-trained model on the task at hand, end-to-end. However, this strategy requires one to store and deploy a separate copy of the backbone parameters for every single task. This is an expensive and often infeasible proposition, especially for modern Transformer-based architectures, which are significantly larger than their convolutional neural networks (ConvNet) counterparts, e.g., ViT-Huge (632M parameters) vs. ResNet-50 (25M parameters). We therefore ask, what is the best way to adapt large pre-trained Transformers to downstream tasks in terms of effectiveness and efficiency?

One straightforward approach is to turn to other strategies that we have perfected for adapting ConvNets to new tasks, as in Fig. 1(a). A popular approach is to fine-tune only a subset of the parameters, such as the classifier head or the bias terms . Prior research has also looked at adding additional residual blocks (or adapters) to the backbone . One could implement similar strategies for Transformers. However, in general these strategies under-perform full fine-tuning in accuracy.

We explore a different route in this paper. Instead of altering or fine-tuning the pre-trained Transformer itself, we modify the input to the Transformer. Drawing inspiration from the recent advances on Prompting in NLP , we propose a new simple and efficient method to adapt transformer models for downstream vision tasks (Fig. 1(b)), namely Visual-Prompt Tuning (VPT). Our method only introduces a small amount of task-specific learnable parameters into the input space while freezing the entire pre-trained Transformer backbone during downstream training. In practice, these additional parameters are simply prepended into the input sequence of each Transformer layer and learned together with a linear head during fine-tuning.

On 24 downstream recognition tasks spanning different domains using a pre-trained ViT backbone, VPT beats all other transfer learning baselines, even surpassing full fine-tuning in 20 cases, while maintaining the advantage of storing remarkably fewer parameters (less than 1% of backbone parameters) for each individual task (Fig. 1(c)). This result demonstrates the distinctive strength of visual prompting: whereas in NLP, prompt tuning is only able to match full fine-tuning performance under certain circumstances . VPT is especially effective in the low-data regime, and maintains its advantage across data scales. Finally, VPT is competitive for a range of Transformer scales and designs (ViT-Base/Large/Huge, Swin). Put together, our results suggest that VPT is one of the most effective ways of adapting ever-growing vision backbones.

Related Work

Transformer models have gained huge success in NLP . The triumph of the Transformer architecture also extends to various computer vision tasks, including image classification , object detection , semantic and panoptic segmentation , video understanding and few-shot learning , surpassing previous state-of-the-art approaches. Transformers are also being widely used in recent self-supervised pre-training methods . Given their superior performance and much larger scale compared to ConvNets, how to efficiently adapt Transformers to different vision tasks remains an important open problem. Our proposed VPT provides a promising path forward.

Transfer learning has been extensively studied for vision tasks in the context of ConvNets and many techniques have been introduced including side tuning , residual adapter , bias tuning , etc. Relatively little attention has been paid to vision Transformers adaptation and how well these aforementioned methods perform on this brand new type of architecture remains unknown. On the other hand, given the dominance of large-scale pre-trained Transformer-based Language Models (LM) , many approaches have been proposed to efficiently fine-tune LM for different downstream NLP tasks . Among them, we focus on the following two representative methods in our experiments for benchmarking purposes: Adapters and BitFit .

Adapters insert extra lightweight modules inside each Transformer layer. One adapter module generally consists of a linear down-projection, followed by a nonlinear activation function, and a linear up-projection, together with a residual connection . Instead of inserting new modules, proposed to update the bias term and freeze the rest of backbone parameters when fine-tuning ConvNets. BitFit applied this technique to Transformers and verified its effectiveness on LM tuning. Our study demonstrates that VPT, in general, provides improved performance in adapting Transformer models for vision tasks, relative to the aforementioned two well-established methods in NLP.

Prompting originally refers to prepending language instruction to the input text so that a pre-trained LM can “understand” the task. With manually chosen prompts, GPT-3 shows strong generalization to downstream transfer learning tasks even in the few-shot or zero-shot settings . In addition to the follow-up works on how to construct better prompting texts , recent works propose to treat the prompts as task-specific continuous vectors and directly optimize them via gradients during fine-tuning, namely Prompt Tuning . Compared to full fine-tuning, it achieves comparable performance but with 1000×\times less parameter storage. Although prompting has also been applied to vision-language models recently , prompting is still limited to the input of text encoders. Due to the disparity between vision and language modalities, in this paper we ask: can the same method can be applied successfully to image encoders? We are the first work (see related concurrent works ) to tackle this question and investigate the generality and feasibility of visual prompting via extensive experiments spanning multiple kinds of recognition tasks across multiple domains and backbone architectures.

Approach

We propose Visual-Prompt Tuning (VPT) for adapting large pre-trained vision Transformer models. VPT injects a small number of learnable parameters into Transformer’s input space and keeps the backbone frozen during the downstream training stage. The overall framework is presented in Fig. 2. We first define the notations in Sec. 3.1, then describe VPT formally in Sec. 3.2.

2 Visual-Prompt Tuning (VPT)

Given a pre-trained Transformer model, we introduce a set of pp continuous embeddings of dimension dd, i.e., prompts, in the input space after the Embed layer. Only the task-specific prompts are being updated during fine-tuning, while the Transformer backbone is kept frozen. Depending on the number of Transformer layers involved, our approach has two variants, VPT-shallow and VPT-deep, as shown in Fig. 2.

2.2 VPT-Deep.

2.3 Storing Visual Prompts.

VPT is beneficial in presence of multiple downstream tasks. We only need to store the learned prompts and classification head for each task and re-use the original copy of the pre-trained Transformer model, significantly reducing the storage cost. For instance, given a ViT-Base with 86 million (M) parameters and d=768d=768, 50 shallow prompts and deep prompts yield additional p×d=50×768=0.038p\times d=50\times 768=0.038M, and N×p×d=0.46N\times p\times d=0.46M parameters, amounting to only 0.04%\% and 0.53%\% of all ViT-Base parameters, respectively.

Experiments

We evaluate VPT for a wide range of downstream recognition tasks with pre-trained Transformer backbones across scales. We first describe our experimental setup in Sec. 4.1, including the pre-trained backbone and downstream tasks, and a brief introduction of alternative transfer learning methods. Then we demonstrate the effectiveness and practical utility of our method in Sec. 4.2. We also systematically study how different design choices would affect performance (Sec. 4.3), which leads to an improved understanding of our approach.

Pre-trained Backbones. We experiment with two Transformer architectures in vision, Vision Transformers (ViT) and Swin Transformers (Swin ). All backbones in this section are pre-trained on ImageNet-21k . We follow the original configurations, e.g., number of image patches divided, existence of [CLS], etc. More details are included in Appendix 0.A.

Baselines. We compare both variants of VPT with other commonly used fine-tuning protocols:

Full: fully update all backbone and classification head parameters.

Methods that focus on the classification head. They treat the pre-trained backbone as a feature extractor, whose weights are fixed during tuning:

Linear: only use a linear layer as the classification head.

Partial-kk: fine-tune the last kk layers of backbone while freezing the others, as adopted in . It redefines the boundary of backbone and classification head.

Mlp-kk: utilize a multilayer perceptron (MLP) with kk layers, instead of a linear layer, as classification head.

Methods that update a subset backbone parameters or add new trainable parameters to backbone during fine-tuning:

Sidetune : train a “side” network and linear interpolate between pre-trained features and side-tuned features before being fed into the head.

Bias : fine-tune only the bias terms of a pre-trained backbone.

Adapter : insert new MLP modules with residual connection inside Transformer layers.

Downstream Tasks. We experiment on the following two collections of datasets:

FGVC consists of 5 benchmarked Fine-Grained Visual Classification tasks including CUB-200-2011 , NABirds , Oxford Flowers , Stanford Dogs and Stanford Cars . If a certain dataset only has train and test sets publicly available, we randomly split the training set into train (90%) and val (10%), and rely on val to select hyperparameters.

VTAB-1k is a collection of 19 diverse visual classification tasks, which are organized into three groups: Natural - tasks that contain natural images captured using standard cameras; Specialized - tasks that contain images captured via specialized equipment, such as medical and satellite imagery; and Structured - tasks that require geometric comprehension like object counting. Each task of VTAB contains 1000 training examples. Following , we use the provided 800-200 split of the train set to determine hyperparameters and run the final evaluation using the full training data. We report the average accuracy score on test set within three runs.

We report the average accuracy on the FGVC datasets, and the average accuracy on each of the three groups in VTAB. The individual results on each task are in Appendix 0.D, as are image examples of these aforementioned tasks.

Tab. 1 presents the results of fine-tuning a pre-trained ViT-B/16 on averaged across 4 diverse downstream task groups, comparing VPT to the other 7 tuning protocols. We can see that:

VPT-Deep outperforms Full (Tab. 1(a)) on 3 out of the 4 problem classes (20 out of 24 tasks), while using significantly fewer total model parameters (1.18×\times vs. 24.02×\times). Thus, even if storage is not a concern, VPT is a promising approach for adapting larger Transformers in vision. Note that this result is in contrast to comparable studies in NLP, where prompt tuning matches, but does not exceed full fine-tuning .

VPT-Deep outperforms all the other parameter-efficient tuning protocols (Tab. 1(b,c)) across all task groups, indicating that VPT-deep is the best fine-tuning strategy in storage-constrained environments.

Although sub-optimal than VPT-deep, VPT-shallow still offers non-trivial performance gain than head-oriented tuning methods in Tab. 1(b), indicating that VPT-shallow is a worthwhile choice in deploying multi-task fine-tuned models if the storage constraint is severe.

We look at the impact of training data size on accuracy in the FGVC tasks (VTAB has only 1k training examples). We vary the training data between 10% and 80% and compare all methods. The same pre-trained ViT-B is used for downstream training. Task-averaged results for each method on different training data scales are presented in Fig. 3.

Fig. 3 shows that VPT-deep outperforms all the other baselines across data scales. Digging deeper, methods that use less trainable parameters, i.e., VPT, Linear, Adapter, Bias, dominate over Full in the low-data regimes. This trend, however, is reversed when more training data is available for Linear and Adapter. In contrast, VPT-deep still consistently outperforms Fullacross training data sizes. Although Bias offers similar advantages, it still marginally under-performs VPT-deep across the board (Fig. 3 right).

2.2 VPT on different backbone scales.

Fig. 4 shows VTAB-1k performance under 3 different backbone scales: ViT-Base/Large/Huge. VPT-deep is significantly better than Linear and VPT-shallow across all 3 backbone choices and 3 subgroups of VTAB-1k. More importantly, the advantages of VPT-deep over Full indeed still hold as the model scale increases, i.e., VPT-deep significantly outperforms Full on Natural and Structured groups, while offering nearly equivalent performance on Specialized.

2.3 VPT on hierarchical Transformers.

We extend VPT to Swin , which employs MSA within local shifted windows and merges patch embeddings at deeper layers. For simplicity and without loss of generality, we implement VPT in the most straightforward manner: the prompts are attended within the local windows, but are ignored during patch merging stages. The experiments are conducted on the ImageNet-21k supervised pre-trained Swin-Base. VPT continues to outperform other parameter-efficient fine-tuning methods (b, c) for all three subgroups of VTAB Tab. 2, though in this case Full yields the highest accuracy scores overall (at a heavy cost in total parameters).

It is surprising that the advantage of VPT-deep over VPT-shallow diminishes for Natural: VPT-shallow yields slightly better accuracy scores than full fine-tuning.

We ablate different model design choices on the supervised ImageNet-21k pre-trained ViT-Base and evaluate them on VTAB, with same setup in Tab. 1. See more in Appendix 0.B.

An important distinction between VPT and other methods is the extra learnable parameters introduced as inputs for the Transformer layers. Fig. 5 ablates different choices on how and where to insert prompts in the input space, and how they would affect the final performance.

Prepend or Add? Instead of prepending prompts to the sequence of the image patches embeddings Ei\mathbf{E}_{i} as described in Sec. 3.2, another option is to directly add prompts element-wise to those embeddings, keeping the Transformer’s input sequence length the same as before. Though this variant is competitive to Full in some cases (e.g., VTAB-Natural), its performance generally falls behind with the default Prepend in both deep and shallow settings. More discussion on this phenomenon is in Sec. 0.B.0.1.

Latent or pixel space? Instead of inserting the prompts as latent vectors for the first Transformer layer, one could introduce prompts in the pixel level before the Embed layer in Eq. 1, i.e., Prepend-pixel and Concat-channel. Fig. 5 shows that the adaption performance decreases for these two variants. For example, the accuracy score of prepending shallow prompts before the projection layer (Prepend-pixel) drops 6.9%\%, compared to the default prepending in the embedding space (Prepend) on VTAB-Natural. The performance further deteriorates (even as large as 30 accuracy scores drop on VTAB-Natural) if we instead concatenate a new channel to the input image (Concat-channel). These observations suggest that it’s easier for prompts to learn condensed task-dependent signals in the latent input space of Transformers.

3.2 Prompt Length.

This is the only additional hyper-parameter needed to tune for VPT compared to full fine-tuning. For easy reference, we also ablate two other baselines on their individual additional hyper-parameters, i.e., number of layers for Mlp and reduction rate for Adapter. As shown in Fig. 6, the optimal prompt length varies across tasks. Notably, even with as few as only one prompt, VPT-deep still significantly outperforms the other 2 baselines, and remains competitive or even better compared to full fine-tuning on VTAB-Structured and Natural.

3.3 Prompt Depth.

Fig. 7 ablates which and how many layers to insert prompts. Each variant reports the best prompt length selected with val set. VPT’s performance is positively correlated with the prompt depth in general. Yet the accuracy drops if we insert prompts from top to bottom, suggesting that prompts at earlier Transformer layers matter more than those at later layers.

3.4 Final Output.

Following the original configuration of ViT, we use the final embedding of [CLS], i.e., xN\mathbf{x}_{N}, as the classification head input, which is also the default setting in our ViT experiments. As shown in Fig. 8, if we use the average pooling on image patch output embeddings EN\mathbf{E}_{N} as final output (Image-pool), the results essentially remain the same (e.g., 82.4 vs. 82.3 for VTAB-Specialized). However, if the pooling involves final prompt outputs ZN\mathbf{Z}_{N} (Prompt-pool and Global-pool), the accuracy could drop as large as 8 points.

Fig. 9 shows t-SNE visualizations of xN\mathbf{x_{N}}, i.e., embeddings of [CLS] after the last Transformer layer and before the classification head, for 3 tasks in VTAB (SVNH , EuroSAT , Clevr/count ), one for each subgroup. All plots show that VPT-deep enables linearly separable representations while using less parameters than Full. We also observe that extra tunable parameters for every Transformer layer (VPT-deep) improve the performance, compared to VPT-shallow, which only inserts prompts for the first layer’s input. Interestingly on Clevr/count (Fig. 9(c)), VPT-deep and Full recover the underlying manifold structure of the task (counting objects in images vs. street number or landscape recognition), unlike VPT-shallow and Linear.

0.2 Apply VPT to more vision tasks.

We explore the feasibility of VPT beyond visual classification, by evaluating ADE20K semantic segmentation task with a Transformer model, SETR-PUP . It adds a standard ConvNet head to the ViT backbone to perform segmentation. The de-facto approach is still fully fine-tuning the pre-trained backbone together with the ConvNet head (Full). We include two more protocols for comparison: only update the head layers (Head Only), update head layers and bias vectors in the backbone (Bias). In Tab. 3, we report val mIoU results with and without multi-scale inference. Though parameter-efficient protocols could not compete with Full, VPT is still comparable with Bias. Notably, VPT offers competitive results to a fully fine-tuned state-of-the-art ConvNet model (DeepLab v3+ ), while tuning significantly less parameters (15M vs. 64M, respectively).

0.3 Apply VPT to more pre-training methods.

In addition to the backbones pre-trained with labeled data, we experiment with two self-supervised objectives: MAE and MoCo v3 . Tab. 4 reports the results on VTAB-1k with ViT-B. We observe that both variants of VPT surpass Linear, yet the comparisons among other techniques are less conclusive. For MAE, other parameter-efficient methods, e.g., Partial-1, outperform both VPT and Linear. In the case of MoCo v3, VPT no longer holds the best performance, though it is still competitive with the others. This suggests that these two self-supervised ViTs are fundamentally different from the supervised ones in previous sections. Exactly why and how these differences arise remain open questions.

0.4 Apply VPT to ConvNets.

We examine the idea of adding trainable parameters in the input space of ConvNets: padding both height and width by pp learnable prompt pixels for the input image. Though this operation seems unconventional, we implement VPT this way given there is no obvious solution to add location-invariant prompts similar to the Transformer counterparts. In fact this approach has been explored before in the adversarial attack literature . The value of pp in our experiment is 2 orders of magnitude smaller than previous work: e.g., 5 vs. 263. Most importantly, we cast this idea in the lens of transfer learning. See Sec. 0.C.0.1 for more discussion.

Tab. 5 presents the results for ConvNeXt-B (pre-trained on ImageNet-21k) and ResNet-50 (pre-trained on ImageNet-1k), respectively. VPT works well in a larger ConvNet backbone, ConvNeXt-B, offering accuracy gains over other sparse tuning protocols (b, c), and outperforming Full on 8 out of 19 cases. The advantages of VPT, however, diminish with smaller ConvNet (ResNet-50), as there is no clear winner for all 19 VTAB-1k tasks.

We present Visual Prompt Tuning, a new parameter-efficient approach to leverage large vision Transformer models for a wide range of downstream tasks. VPT introduces task-specific learnable prompts in the input space, keeping the pre-trained backbone fixed. We show that VPT can surpass other fine-tuning protocols (often including full fine-tuning) while dramatically reducing the storage cost. Our experiments also raise intriguing questions on fine-tuning dynamics of vision Transformers with different pre-training objectives, and how to transfer to broader vision recognition tasks in an efficient manner. We therefore hope our work will inspire future research on how best to tap the potential of large foundation models in vision.

Acknowledgement. Menglin is supported by a Meta AI research grant awarded to Cornell University, Luming and Bharath is supported by NSF IIS-2144117, Serge is supported in part by the Pioneer Centre for AI, DNRF grant number P1. We would like to thank Alexander Rush, Yin Cui for valuable suggestions and discussion.

We use PyTorch to implement all experiments on NVIDIA A100-40GB GPUs.

We use val set of each dataset to find best prompt length pp, see Sec. 3.2. The prompt length is the only VPT-specific hyper-parameter that we tune. For Transformer backbones, the range of pp is {1,5,10,50,100,200}\{1,5,10,50,100,200\} and {1,5,10,50}\{1,5,10,50\} for ViT and Swin, respectively. The maximum choice of pp is approximately close to the number of image patch tokens within each MSA for both architectures (ViT: 196, Swin: 49). We also apply a dropout of 0.10.1 for VPT-deep. For ConvNets, the range of pp is {1,3,5,7,9,11}\{1,3,5,7,9,11\}. Each prompt is randomly initialized with xavier uniform initialization scheme . We follow the original backbone’ design choices, such as the existence of the classification tokens [CLS], or whether or not to use the final [CLS] embeddings for the classification head input.

A.1.2 Adapter.

Adapters insert extra lightweight modules inside each Transformer layer. One adapter module generally consists of a linear down-projection (with a reduction rate rr), followed by a nonlinear activation function, and a linear up-projection, together with a residual connection. exhaustively searched all possible configurations and found that only inserting adapters after the FFN “Add & LayerNorm” sub-layer works the best. Therefore we also use this setup in our own implementation. We sweep the reduction rate rr in {8,64,256}\{8,64,256\}.

A.1.3 Augmentation and other hyper-parameters.

We adopt standard image augmentation strategy during training: normalize with ImageNet means and standard deviation, randomly resize crop to 224×\times224 and random horizontal flip for five FGVC datasets, and resize to 224×\times224 for the VTAB-1k suite.Following the default settings in VTAB, we don’t adopt other augmentations Tab. 6 summarizes the optimization configurations we used. Following , we conduct grid search to find the tuning-specific hyper-parameters, learning rate, and weight decay values using val set of each task. Following the linear scaling rule , the learning rate is set as base_lr×b/256\times b/256, where bb is the batch size used for the particular model, and base_lr is chosen from the range specified in Tab. 6. The optimal hyper-parameter values for each experiment can be found in Appendix 0.D.

A.1.4 Datasets and pre-trained backbones specifications.

Tabs. 7 and 8 summarize the statistics and details of the evaluated classification datasets and all the pre-trained backbones used in the paper. Fig. 10 includes image examples of all 24 classification tasks evaluated.

A.2 Semantic Segmentation Experiments

ADE20K is a challenging scene parsing benchmark with 150 fine-grained labels. The training and validation sets contain 20,210 and 2,000 images respectively. We utilize the public codebase MMSegmentation in our implementation.See the MMSegmentation GitHub page The ViT-L backbone is supervisely pre-trained on ImageNet-21k.ViT-L/16 checkpoint

SETR is a competitive segmentation framework using ViT as the encoder. PUP is a progressive upsampling strategy consisting of consecutive convolution layers and bilinear upsampling operations. Among multiple decoder choices, PUP works the best according to MMSegmentation’s reproduction therefore we also use it as in our implementation. MMSegmentation’s reproduction on SETR

When applying VPT to SETR-PUP, we only insert prompts into SETR’s ViT encoder backbone. For the decoder, only image patch embeddings are used as inputs and prompt embeddings are discarded. Same as recognition tasks, only the PUP decoder head and prompts are learned during training and the ViT backbone is frozen.

For full fine-tuning, we use the same hyper-parameters as in MMSegmentation. For HeadOnly, Bias, and VPT, we use the hyper-parameter sweep on learning rate {0.05, 0.005, 0.0005, 0.001}. The optimal learning rate is 0.005 for all methods. We sweep prompt length p∈p\in {1, 5, 10, 50, 100, 200}. For VPT, we also change the learning rate multiplier to 1.0 instead of the default 10.0, so the decoder head and prompts share the same learning rate. Other hyper-parameters remain the same as full fine-tuning.

Appendix 0.B Extended Analysis

As shown in Tab. 1, by expanding the input sequence with learnable prompts, VPT achieves better performance than Full on the 20 out of 24 tasks evaluated. To investigate whether the advantage of VPT is due to its enlarged input sequence length, we experiment on two more variants: (1) the prompts are kept frozen during fine-tuning stage (Prompt-Fixed). (2) only tuning the [CLS] token ([CLS]-Learned). From Fig. 11 we can see that, updating prompt embeddings (Prompt-Learned) offers significant gains, while Prompt-Fixed yields comparable results w.r.t. Linear. This suggests that the final performance of VPT is mainly contributed by the learned prompt embeddings instead of the enlarged sequence length. Updating the [CLS] token performs similarly as updating 1 prompt ([CLS] vs. Learnedp=1), but still lags behind the default setting where we manually select the best number of prompt tokens based on the val set.

B.0.2 Sharing prompts.

We examine the effect of sharing parameters of prompts in Fig. 12 by setting the same prompt embedding within Transformer layers (Shared-intra), among all layers (Shared-inter), and for all prompts inserted in the Transformer (Shared-all). We can observe that: (1) Sharing prompts within layer (Shared-intra) performs competitively or slightly outperforms the performance of using one prompt (Defaultp=1), further demonstrating the value of expanding input sequence. (2) Although Shared-intra under-performs Default in general, surprisingly, Shared-inter slightly outperforms our default VPT-deep while using similar number of trainable parameters (total number of parameters for all VTAB tasks: 1.14×\times vs. 1.13×\times for Shared-inter vs. Default, respectively). Closer examination reveals that the optimal prompt length pp for Shared-inter is in general larger than Default, i.e., average prompt length on all VTAB tasks: 64.58 vs. 60.94, for Shared-inter vs. Default, respectively. (3) Sharing the same prompt embedding both among and within layers (Shared-all) deteriorates performance, but still surpass the linear probing results across three VTAB subgroups.

B.0.3 Prompt initialization.

In NLP, prompt tuning could benefit from more sophisticated prompt initialization, as shown in . We investigate if this is the case for visual prompting as well. We utilize prototype representations for downstream target classes so that the prompts are initialized with embeddings that enumerate the output space. Since we want the model to produce an output embedding that is close to one of these prototype representations given a test example, initializing prompts in this manner might give the model some hints about the target categories thus help improve the optimization process.

We compare the fine-tuning performance using the above initialization strategy (CLS) against the default random initialization (Random) in Fig. 13. We also report results when we fix the prompts during the fine-tuning stage (⋅\cdot-fixed). As shown in Fig. 13, it’s quite surprising that our default random initialization (Random) works the best in general, consistently across different subgroups of VTAB without extra pre-processing steps described above (CLS). CLS works comparably in Natural and Specialized subgroups. Utilizing the per-class averaged [CLS] features, we also tried several other different implementation variants, including using per-layer [CLS] embeddings for VPT-deep instead of only the final output [CLS] vector. They perform either the same as or even much worse than the CLS strategy above, and none of them is able to out-perform the default Random.

B.0.4 Prompt depth vs. prompt length.

In Fig. 7, we ablate the number of layers we insert prompts in. For each prompt depth variant, Fig. 7 reports the results using the best prompt length for each task (“⋅→⋅\cdot\rightarrow\cdot (best)” in Fig. 14). Here we adopt another setting where the best prompt length from 1→121\rightarrow 12 are used for all other prompt depth variants. Comparing both “⋅→⋅\cdot\rightarrow\cdot (best)” and “⋅→⋅\cdot\rightarrow\cdot”, we observe that there are varied sensitivities to prompt length for different depths, especially if we insert prompts in nine layers only (3→3\rightarrow12, 12→312\rightarrow 3).

B.0.5 Combine VPT with Bias Tuning.

Our experiments in the main paper reveal that Bias is a competitive parameter-efficient tuning baseline (e.g., Tab. 1(c)). Based on this observation, we explore another protocol where we update both prompts and the bias terms of the pre-trained backbone, keeping everything else in the backbone frozen (VPT+Bias). As shown in Tab. 9, to our surprise, incorporating Bias with VPT does not yield superior results in general, even undermines VPT-deep for all 3 task subgroups. This suggests that these two methods are not necessarily complementary to each other.

B.0.6 Prompt ensembling.

demonstrated prompt’s efficiency in the context of model ensembling. For an ensemble of kk models, we only need to store the learnt prompt vectors instead of kk copies of the whole fine-tuned model parameters (e.g., k×2.5k\times 2.5GB for ViT-H). Furthermore, given one test example during inference time, only one forward pass is executed with a specially-designed batch with replicated original data but varied prompts.

Given such advantages, we also investigate VPT’s effectiveness on prompt ensembling. We train 5 different prompts for each VTAB task with different random seeds, using the same pre-trained ViT-B backbone and hyper-parameters as in Tab. 1. Fig. 15 shows that the ensembled VPT-deep outperforms the average or even the best single-prompt counterparts, as well as other ensembled fine-tuning methods including Full.

B.0.7 Test of statistical significance.

We conduct non-parametric paired one-tailed tt-test (the Wilcoxon signed-rank test ) on whether VPT-deep’s performance is greater than other fine-tuning methods across 19 VTAB tasks (the null hypothesis H0H_{0} states that the mean VTAB performance difference between VPT-deep and alternate baseline method is zero. The alternative hypothesis H1H_{1} states that VPT-deep outperforms the baseline method on VTAB). Tab. 10 presents the pp-values of each test, with the number of observations equal to 19 for each method compared (we use the averaged accuracy scores among 5 runs for 19 VTAB tasks and all fine-tuning methods). For all of the fine-tuning protocols compared, VPT-deep’s improvements are statistically significant (p<0.05p<0.05).

We also conduct un-paired one-tailed tt-test with unequal variances (Welch’s tt-test ), comparing the individual runs (the number of observations = 5) for each VTAB task (H0H_{0} states that VPT-deep and the other baseline perform the same for a specific VTAB task, while H1H_{1} states that VPT-deep outperforms the other baseline for a specific VTAB task). Fig. 16 presents the pp-values for each <<VPT-deep, baseline method>> pair on each task. We reject H0H_{0} on 127 out of 19×8=15219\times 8=152 cases (p<0.05p<0.05). Compared to Full, VPT-deep achieves statistically significant better performance on 11 out of 19 tasks.

B.0.8 Effect of different fine-tuning hyper-parameters.

In Fig. 17, we present different tuning protocol’s performance on different fine-tuning hyper-parameters including learning rate and weight decay. For our proposed VPT-deep, we also ablate different choices of prompt length pp, which is the only hyper-parameter that needs to be manually tuned. All experiments are evaluated on the val set of KITTI/Distance task (VTAB-Specialized). We observe different behaviors between Linear and VPT. Both methods freeze backbone parameters during fine-tuning stage. Linear probing is more sensitive to weight decay values in general, whereas VPT is influenced by both learning rate and weight decay values. VPT with larger prompt length is also less sensitive to the choice of learning rate.

B.0.9 Effect of image resolution.

The original ViT paper found that fine-tuning with higher image resolutions (384×\times384) is beneficial to downstream recognition tasks. All recognition experiments presented in the main paper are fine-tuned on 224×\times224 resolution. As shown in Tab. 11, we re-run the VTAB experiments with the same setup as in Tab. 1 but in the 384 resolution instead of the default 224. We can see that, VPT-deep still achieves the best performance among all parameter-efficient tuning protocols, and even outperforms full fine-tuning on 15 out of 19 tasks. Although the increase of image resolutions doesn’t lead to better full fine-tuning performance in general, it indeed slightly boosts VPT-deep’s performance.

Another interesting observation from Tab. 11 is that with 224 fine-tune resolution and a larger value of p=380p=380, VPT could achieve similar or better performance compared to Full with 384 resolution, while using the same input sequence length yet significantly less trainable parameters.

B.0.10 Empirical computational cost.

One possible limitation of VPT is the extra input sequence length for Transformers. In theory the complexity of MSA is quadratic w.r.t. the input sequence length, but this might not be the case for real-world speed due to hardware details like lane widths and cache sizes . In Tabs. 12 and 18, we study the empirical computational cost, i.e., latency, and peak GPU memory usage at both training and inference times, for all the fine-tuning protocols studied. All experiments use the same A100 GPU with a batch size 64 for both training and inference. We can see that the theoretical quadratic scaling w.r.t. sequence length barely happens to VPT. For instance, doubling the length (p=200p=200 vs. m=198m=198) basically only lead to 2×\times (instead of 4×\times) inference latency and peak GPU memory w.r.t. full fine-tuning. For training, the latency would be largely reduced with less number of prompts.

An equivalent implementation of VPT during test time is directly prepend the parameters to the key and value arrays inside the self-attention module of Transformer (VPT-prefix). While we found that such implementation does not lead to accuracy improvement on VTAB datasets, it reduces the computation cost during inference. Figure 19 shows the comparison with different values of pp. VPT-prefix reduces test-time latency and peak GPU memory with a large margin especially when pp becomes large.

Appendix 0.C Further Discussion

The differences are: (1) the number of learnt parameters injected in the input space in AR literature is nearly 20 times larger than ours (264k vs. 13k). VPT is significantly more parameter-efficient; (2) AR has shown its effectiveness in ConvNet, while VPT can be applied to broader architectures, including ViT, Swin. Furthermore, VPT is more general with the option of diving into deeper layers of pre-trained backbone (Fig. 2), whereas AR strictly applies to the first input layer of ConvNets. (3) another distinction is that our setting update both prompts and classification head, while AR directly use the pre-trained classification head. Our setup is more general and could be applied to models with a broader range of pre-training objectives (e.g., MAE , which does not include a pre-trained classification head) and broader vision tasks (e.g., segmentation).

C.0.2 Visual prompt vs. textual prompt.

Our paper also discover discrepancies between visual and textual prompts: we show that VPT could even outperform full-model fine-tuning on 20 out of 24 cases, which is in contract to the NLP’s related work . We also found that random initialized prompts works better in Fig. 13, and prompts at earlier layers matters more (Figs. 7 and 14), which are also different from observation on the NLP side . These discrepancies indicate that visual prompting might be fundamentally different from text prompts thus in need of further investigation.

Appendix 0.D Supplementary Results

Tabs. 13 and 14 present per-task results for 24 classification tasks evaluated in Tab. 1.

D.0.2 Per-task results on training data ablations.

Fig. 20 presents the per-task results for five FGVC datasets. We observe a similar trend in Fig. 3: while all parameter-efficient methods outperform full fine-tuning in small-to-medium data regime, VPT-deep consistently surpasses Full across data scales for five FGVC tasks.

D.0.3 More t-SNE visualizations.

In Fig. 21, We presents more t-SNE visualizations, similar to Fig. 9, for all VTAB datasets with less than or equal to 20 target classes.

References