Exploring Visual Prompts for Adapting Large-Scale Models

Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, Phillip Isola

Introduction

When we humans learn a new task, we tend to start from our current knowledge base and extrapolate thereof. A child who is starting to speak and comprehend sentences quickly develops the ability to parse the emotional context that accompanies a sentence. For example, the sentence “I missed the school bus” carries a particular emotion such that if followed by “I felt so [MASK]”, the child can provide an appropriate emotion word. This paradigm, aptly named prompting, has recently been popularized in NLP, where large pre-trained language models are adapted to new tasks by converting the downstream dataset into the format of the pre-training task. Without updating any of its parameters, the language model uses its existing knowledge base to fill in the mask in the provided prompt, hence becoming an expert in the new task. Currently, prompting methods are dominantly NLP-specific , despite the fact that the framework serves a general purpose: adapt a frozen pre-trained model by modifying the data space. Considering the generality, can we create prompts in the form of pixels? Broadly, can we steer frozen visual models to solve a new task by modifying pixel space?

Adversarial reprogramming is a class of adversarial attacks where input perturbations repurpose a model to perform a task chosen by the adversary. Despite having different termsFor convenience, we unify the term and use “visual prompt” to denote any pixel-space modification to the input image for model adaptation. and motivations, this input perturbation essentially acts as a visual prompt — it adapts a model to new tasks by modifying pixels. Existing methods , however, have focused on adversarial goals or demonstrated limited application to relatively small-scale datasets and models. Having originated from different communities, adversarial reprogramming and prompting share the general idea : perform data-space adaptation by transforming the input (i.e., prompt engineering) and/or output (i.e., answer engineering).

Inspired by the success of natural language prompting, we aim to investigate the efficacy of visual prompting for adapting large-scale models in vision. As pixel space is inherently continuous, we follow the recent approach that treats prompts as a continuous task-specific vector . We learn a single image perturbation (i.e., “soft prompt”) via backpropagation while having the model parameters frozen. We map the model outputs to downstream labels by using a discrete text prompt for CLIP and hard-coded mapping for vision models (Figure 1).

How is visual prompting different from existing adaptation methods? Currently in vision, standard adaptation methods are fine-tuning and linear probe. Both approaches require some level of access to the model: entire parameters in the case of fine-tuning and model outputs (usually activations at the penultimate layer) in the case of linear probe. In contrast, visual prompting adapts the input to a model. After acquiring the visual prompt, it does not require model access at test time. This opens up unique applications ; input-space adaptation puts control in the hands of the end-user of the system. For instance, a pedestrian could wear a visual prompt that improves their visibility to cars, without having access to the car itself, nor its vision system.

We conduct comprehensive experiments across four pre-trained models and 15 image classification datasets. We demonstrate that visual prompting is surprisingly effective for CLIP and robust to distribution shift, achieving performance competitive with, and sometimes beyond, standard linear probes. We further analyze what properties of the downstream dataset, prompt design, and output transformation affect performance. Note that our goal is not to achieve the state-of-the-art performance on specific tasks, but instead to broadly explore a new paradigm for visual adaptation. The surprising effectiveness of visual prompting provides a new perspective on how to adapt and use pre-trained models in vision.

Related Work

Our investigation is inspired by the recent success in natural language prompting. Prompting in NLP reformulates the downstream dataset into a (masked) language modeling problem, so that a frozen language model directly adapts to a new task without updating any parameters. A prompt consists of constructing a task-specific template (e.g., “I felt so [MASK]”) and label words (e.g., “happy/horrible”) to fill in the blank . However, hand-crafting the right prompt requires domain expertise and a significant amount of effort.

Prefix tuning or prompt tuning mitigates this problem by learning a “soft prompt” via backpropagation, while having the model parameters fixed. Prefix tuning learns a task-specific continuous vector (i.e., prefix) that allows language models to adapt to various generation tasks. While prefix tuning prepends the prefix to each encoder layer, prompt tuning further simplifies by only prepending tunable tokens to the input. When applied to large models with billions of parameters, a properly optimized prompt achieves competitive performance to fine-tuning the entire model, while significantly reducing memory usage and per-task storage. As prompts in pixel space are inherently continuous, we follow this line of work and optimize the pixels directly.

2 Prompting with Images

There have been initial approaches that attempt to prompt with images. Similar to prefix tuning, Frozen creates a image-conditional prompt by training a vision encoder using gradients from a frozen language model. The images are represented as a continuous embedding from the vision encoder and used as a visual prefix to allow frozen language models to perform multi-modal tasks. CPT converts visual grounding into a fill-in-the-blank problem by creating visual prompts with colored blocks and color-based textual prompts. However, both of these approaches focus on extending the capabilities of a language-based model. On the other hand, we focus on investigating the efficacy of prompting for visual representations and image classification datasets. In other words, we assume that the pre-trained model consists of a visual encoder and focus on reformulating image datasets. Visual prompt tuning is concurrent work that proposes visual prompts specific to Vision Transformers . It uses deep prompt tuning by prepending a set of tunable parameters to each Transformer encoder layer.

3 Adversarial Reprogramming and Unadversarial Examples

ß Adversarial reprogramming is a type of adversarial attack where a single, class-agnostic perturbation reprograms a model to perform a new task chosen by the attacker. Despite its adversarial goal, the framework essentially serves the same purpose as prompting: adapt a frozen model to new tasks by modifying the input and/or output of the downstream dataset. However, existing methods in vision are designed to achieve an adversarial goal or demonstrate limited application to small-scale vision models and simple datasets. Similarly, unadversarial examples aim to increase performance on the (pre-)trained task. It learns an image perturbation that improves performance on a specific class (i.e., class-conditional). In our work, we revisit adversarial reprogramming as a form of visual prompt and investigate its efficacy in adapting large-scale models in vision.

4 Adapting Pre-trained Models in Vision

Figure 2 provides a summary of different methods for adapting a pre-trained model. Fine-tuning and linear probing are highly flexible in their usage: they can be used to adapt the model to a new domain of inputs or to a new task with different output semantics. However, they also require some level of access to the model: parameters in the case of fine-tuning and model outputs (usually activations at the penultimate layer) in the case of linear probes. Domain adaptation is an interesting alternative to model adaptation in that it only modifies the inputs to the model using techniques such as image-to-image translation . Like domain adaptation, visual prompting also modifies the inputs to a model. Therefore, once the end user has found the visual prompt, it does not require having control over the model itself at test time. This opens up unique applications; for example, users can feed domain-adapted images to online APIs that can only be manipulated via their inputs. Domain adaptation focuses on adapting a source domain to look like a target domain, requiring both source and target datasets available at hand. On the other hand, we demonstfirate that visual prompting can steer model in more arbitrary ways; for example, a model that performs one classification task can be adapted to perform an entirely different classification task, with new output semantics, just by perturbing the input pixels. Also, whereas domain adaptation methods are typically input-conditional, the visual prompts we explore in this paper are fixed (i.e., input-agnostic) across an entire dataset, as in NLP where the same natural language prompt is added to all model queries.

Methods

Under different terms, prompting and adversarial reprogramming serve the same purpose: data-space adaptation. They generally consist of two stages : input transformation and output transformation. The goal of the input transformation (or prompt engineering) is to design a proper prompt that specifies the task which is applied to the input. The goal of the output transformation (or answer engineering) is to map the model’s output/answer to the target label. We introduce different design choices for vision and vision-language models according to their pre-trained task.

Prompts in pixel form can essentially be applied to any visual representation. Therefore, we select three vision models and one vision-language model: Instagram-pretrained ResNeXt (Instagram) , Big Transfer (BiT-M) , ResNet trained on ImageNet-1k (RN50) , and CLIP . Vision models are trained to predict a fixed set of predetermined classes and typically require learning a separate layer to predict unseen classes. In contrast, CLIP is a vision-language model that is able to perform flexible zero-shot transfer to unseen classes using text prompts. We summarize the pre-trained model details in the Appendix. We select models across varying input modalities, pre-trained dataset size, and model architecture to evaluate the practical utility of visual prompts. For Instagram-pretrained ResNeXt, we use the model additionally fine-tuned on ImageNet-1k.

2 Input Transformation

There can be several ways of designing a visual prompt. As pixel space is less discrete compared to natural language, it is difficult to handcraft prompts as in NLP (e.g., “a photo of a [LABEL]” for image classification). In fact, it is unclear what type of visual context is useful for each downstream task (e.g., what visual information would be useful for specifying satellite image classification?). Intuitively, a visual prompt does not necessarily need to be interpretable to humans; it’s a visual cue that aids the decision of a machine learning model. Thus, let the model optimize the visual context! We follow a simple gradient-based approach where we directly optimize the visual prompt via backpropagation.

Given a frozen pre-trained model FF and a downstream task dataset D={(x1,y1),…,(xm,ym)}\mathcal{D}=\{(x_{1},y_{1}),\dots,(x_{m},y_{m})\}, our objective is to learn a single, task-specific visual prompt vϕv_{\phi} parameterized by ϕ\phi. The prompt is added to the input image to form a prompted image x+vϕx+v_{\phi}. During training, the model maximizes the likelihood of the correct label yy,

while the gradient updates are applied only to the prompt parameters ϕ\phi and the model parameters θ\theta remain frozen. During evaluation, the optimized prompt is added to all test-time images,

which are then processed through the frozen model FF.

Note that our goal is to explore visual prompts as a practical adaptation method. Therefore, we do not necessitate any adversarial constraint of making the perturbations imperceptible. Also, adversarial reprogramming assumes the downstream dataset to be lower-resolution than the pre-trained dataset, such that the input perturbation is padded around the downstream dataset. In real-world applications, the downstream dataset can have varying resolutions. Thus, we resize every dataset to the input size of the pre-trained model and add the prompt directly to the input region.

2.2 Prompt Design

There are several ways to design a visual prompt in terms of template and size. We explore three visual templates: pixel patch at random location, pixel patch at fixed location, and padding. We explore various prompt sizes pp, where the actual number of parameters is Cp2Cp^{2} for patches and 2Cp(H+W−2p)2Cp(H+W-2p) for padding, where CC, HH, WW are the image channels, height and width respectively. Section 6.2 shows that padding with p=30p=30 achieves the best performance over other design choices. We use this as default for all our experiments.

3 Output Transformation

To map model outputs to the target label, we take a different approach for vision models and CLIP. Standard vision models treat image classes as a numeric id (e.g., “cat” is mapped to “index 1”). We use a hard-coded mapping and arbitrarily map downstream class indices to pre-trained class indices, discarding unassigned indices for loss computation. For CLIP, a vision-language model, we utilize text prompts as our output transformation function. Image classes are represented by text (e.g., “cat”) which are then prompted (e.g., “a photo of a [object]”) to specify context of the downstream task. Note that we use a single, fixed text prompt (see Appendix) and only optimize the visual prompt. We follow the protocol for CLIP zero-shot transfer and calculate cosine similarity of the embeddings for every class, which is normalized into a probability distribution via softmax. The class with the highest probability is chosen as the model output. The full overview for vision models and CLIP is illustrated in Figure 1.

4 Implementation Details

To learn the visual prompt, the objective function for CLIP is identical to its evaluation setting, i.e., we only compute cross entropy loss over images, where a set of prompted text strings is processed through the text encoder to produce weights of a linear classifier . For vision models, we compute cross entropy loss over new class indices. For all experiments, we use the padding template with prompt size of 30. All images are resized to 224×224224\times 224 to match the input size of pre-trained models, and preprocessed identical to the evaluation setting of each model. We find that closely following the pre-trained model’s evaluation setting is important for learning a good prompt. All visual prompts are trained for 1,000 epochs. We use SGD with a learning rate of 40, which is decayed using cosine schedule . We use a batch size of 256 for CLIP, 128 for BiT-M and RN50, and 32 for Instagram.

Experimental Setup

To evaluate how well visual prompts adapt a model to new tasks, we measure performance across 12 datasets: CIFAR100, CIFAR10 , Flowers102 , Food101 , EuroSAT , SUN397 , DTD , UCF101 , SVHN , OxfordPets , Resisc45 , and CLEVR . We also measure robustness to distribution shift, i.e., training distribution differs from the test distribution, by evaluating on three image classification datasets in WILDS : Camelyon17 , FMoW , and iWildCAM . For Camelyon17, training and test sets comprise tissue patches from different hospitals. For FMoW, training and test sets are from different regions and years. Finally, iWildCAM consists of photos from disjoint sets of camera traps. Note that we learn the visual prompt on the training set and evaluate its performance on the test set.

2 Baseline Methods

To measure how visual prompting performs compared to existing adaptation methods (Figure 2), we compare fine-tuning, linear probes, and text prompting (i.e., zero-shot transfer). Fine-tuning and linear probe are standard adaptation methods in vision. Fine-tuning update the entire model parameters during adaptation. Linear probe is a lightweight alternative which adapts the model outputs (usually activations at the penultimate layer) by learning a linear layer, while having the model parameters frozen. For text prompting, we use “This is a photo of a [LABEL]” as default. For CLEVR, we use “This is a photo of [LABEL] objects”, with class label “three” to “ten”. For Camelyon17, we use “a tissue region [LABEL] tumor”, with class label “containing” and “not containing”.

Results

We first compare prompting performance with linear probe, the current de facto approach to lightweight adaptation. Figure 4 shows average test accuracy across 12 datasets for each pre-trained model. Prompting with vision models, or adversarial reprogramming, shows significant performance gap (+40%) to standard linear probe. On the other hand, we find that prompting is surprisingly effective for CLIP, achieving competitive performance to linear probe. In particular, prompting outperforms linear probe on EuroSAT, SVHN, and CLEVR, by 1.1%, 23%, and 15.4% respectively (Table 1). On average, learning a visual prompt achieves 24% performance gain compared to using text prompt only (i.e., “zero-shot transfer”). Interestingly, we find that the performance of visual prompts varies across datasets (Figure 4). Regarding this phenomenon, we further analyze what properties of the dataset affect performance in Section 6.1. We report full results across 12 datasets for vision models in the Appendix.

2 Robustness to Distribution Shift

As model parameters remain frozen, prompting prevents modifying the general knowledge base of the pre-trained model. This reduces the possibility of overfitting to spurious correlations in the downstream dataset, thereby improving robustness to distribution shift. Using the WILDS benchmark , we learn visual prompts from training sets that contain images from a particular domain, and see how it transfers to test sets from different domains (e.g., images from different hospitals, regions, years, cameras). Table 2 show that average performance gap compared to linear probe and fine-tuning is further reduced to 4.5% and 3.5% respectively. On Camelyon17, visual prompting outperforms both linear probe and fine-tuning by 4.9% and 6.5% respectively. This suggests the practical utility of prompting in real-world deployments, where diverse range of domain shifts naturally arise. We report robustness results for vision models in the Appendix.

Understanding Visual Prompts

In this section, we investigate visual prompting performance in regard to properties of the downstream dataset, prompt design (i.e., input transformation), and output transformation.

We find that the performance of visual prompting varies across downstream datasets. As shown in Figure 4, the best-performing dataset achieves +83.2% accuracy gain, while the worst-performing dataset has -1% accuracy loss. To explain this phenomenon, we first hypothesize that visual prompts bridge the distribution gap by converting the unfamiliar downstream dataset to look more similar to the pre-trained dataset. Under this hypothesis, visual prompts would not help datasets already within the pre-trained distribution, yet could help datasets that are severely out-of-distribution. While CLIP’s pre-trained dataset is not available to the public, it is excessively tuned to achieve state-of-the-art zero-shot performance on ImageNet. Thus, we use ImageNet as a proxy. We validate our hypothesis by measuring the distributional similarity between ImageNet and downstream datasets using the FID score . We compare these scores to accuracy gain from visual prompts. Due to computation limitations, we randomly sample 100k images from the ImageNet-1k training set to compute the metrics. In Figure 6, we observe a general performance gain as the downstream dataset becomes more out-of-distribution to ImageNet (e.g., CLEVR, SVHN).

Another hypothesis regards to learning a single prompt per dataset. While this may be sufficient for datasets with low perceptual diversity, a single visual prompt may fail to capture the full distribution as the diversity increases. We measure perceptual diversity using LPIPS . For each dataset, we measure LPIPS between two randomly sampled image pairs and report the average score. Figure 6 shows that a learning single visual prompt achieves better performance gain for datasets with low perceptual diversity.

2 Prompt Design

Choosing the right prompt design (i.e., template and size) can highly affect performance. We perform an ablation study on three different templates: pixel patch at random location, pixel patch at fixed location, and padding, across prompt size p=1,…,224p=1,\dots,224. We measure accuracy on the EuroSAT dataset using a frozen CLIP. Figure 6.2 shows that using a fixed-location template (i.e., padding, fixed patch) yields better performance. For fixed-location templates, we find that performance improves as prompt size increases (i.e., more trainable parameters), then it starts to drop for +70k parameters. Surprisingly, we find that our simplest approach — adding a single-pixel prompt — can yield a 3% improvement over text-prompted CLIP (Figure 7). Overall, padding with p=30p=30 achieves the best performance in our experiments. We believe this is because our application scope is image classification, where the object of interest tends to be located in the center of the image. We believe other visual tasks may require significantly different design choices. Refer to Section 3.2.2 on how the actual number of parameters are calculated.

3 Output Transformation

We investigate prompting performance in regard to how we design the output transformation. For vision models, we follow and use hard-coded mapping; downstream class indices are arbitrarily assigned to pre-trained class indices. We analyze how this mapping affects downstream performance. Using a subset of OxfordPets, we construct a simple toy dataset for classifying dogs and cats. Using ResNet trained on ImageNet-1k (RN50), we compare two cases: (1) downstream classes are assigned to pre-trained classes with similar semantics (unseen “dog” assigned to pre-trained “chihuahua” index), (2) we swap the indices (cat assigned to dog index, vice versa). (1) achieves 100% and (2) achieves 62.5%; having similar semantics between class indices is critical for performance. This may explain the performance gap between vision models and CLIP.

For CLIP, a vision-language model, we use text prompts for output transformation. As we learn visual prompts via backpropagation, the learning signal is dependent on the text prompt we use. It has been reported that CLIP’s zero-shot accuracy can be significantly improved by using a better text prompt . Therefore, we hypothesize that the quality of text prompt affects the performance of visual prompts. On EuroSAT, we measure text prompt quality by the zero-shot performance of CLIP. Figure 8 shows that the performance gain from visual prompts is higher for text prompts with low zero-shot performance. In other words, visual prompting can compensate for low-quality text prompts. As manually searching for the best text prompt is extremely laborsome, this result highlights the usefulness of visual prompts.

Discussion

In this paper, we have investigated a method to perturb inputs to a pre-trained model in a manner which improves classification accuracy. A broader interpretation of visual prompting is to think of it as a way to steer a pre-trained model in any direction by modifying its input space. For instance, a visual prompt for an image-to-image model could be used to change the visual style of the input. Even though we have explored “universal” visual prompts in this work (i.e., a single prompt that apply to all input images), prompts could also be made input-conditional and hence less universal but perhaps more accurate. The specific design choices including (a) input-specific or input-agnostic, (b) improving or decreasing accuracy, and (c) type of the pretrained-model, can be modified to create future interesting applications of prompting.

One natural question that arises following our exposition is in what situations would one prefer visual prompting over fine-tuning or a linear probe? Fine-tuning assumes that the model can be modified which may not always be the case (e.g., if the model is exposed by a API owned by a third-party). While prompting does under-perform linear probe in some cases, we would like to stress that the goal of this work is to show the existence of a prompting mechanism in “pixel space”, which works across multiple datasets and pre-trained models, and reveals new avenues for how vision models can be effectively adapted. Our focus in this work is not to outperform state-of-the-art; we note that there are several approaches one could use to improve performance further including ensembling multiple prompts, using prompts in conjunction with linear probe or fine-tuning, or scaling the pre-trained model (e.g., ViT-L/14 of CLIP, which is unfortunately not available to the public). We leave these for future work.

Conclusion

While standard adaptation methods in vision focus on introducing a separate task-specific head and adapt the model parameters or activations, we investigate visual prompting as practical adaptation method. We use a gradient-based scheme to learn a single, input-agnostic perturbation that repurposes a frozen model to perform a downstream task. Through various experiments across pre-trained models and datasets, we have demonstrated that CLIP is particularly suitable for visual prompting, achieving competitive results to linear probe. We hope that our unique findings will spur further research into: (1) better understanding pixel-space adaptation — when and why they are effective at steering deep networks, and (2) developing better visual prompts that further add to our repertoire of mechanisms for creating flexible and adaptable vision systems.

Acknowledgements

We would like to thank Lucy Chai, Caroline Chan, Joanna Materzynska, Xavier Puig Fernandez, Minyoung Huh, Tongzhou Wang, and Yen-Chen Lin for proofreading the paper. We thank Judy Hoffman for helpful discussion and advice. This work was partially supported by funding from MIT STL and an MIT RSC award from the NEC fund.

References

Appendix A Appendix

In Figure 9, we compare performance using different model architectures on CIFAR100. We compare the original ImageNet-pretrained ResNets released by , namely ResNet-18, ResNet-50, ResNet-101, ResNet-152. For Instagram-pre-trained ResNeXt , we compare two models (32x8d, 32x16d). For Big Transfer , we use four BiT-M models (ResNet-50, ResNet-101, ResNet-50x3, ResNet-101x3). For ResNet-based CLIP models, we compare two models trained on 224×224224\times 224 images (ResNet-50, ResNet-101). For CLIP models that use the Vision Transformer , we compare the two released models (ViT-B/32, ViT-B/16). For vision models, performance does not necessarily increase for larger models. For CLIP, we observe superiority of ViT-based models over ResNet-based models.

A.2 Dataset Statistics

Table 6 illustrates description of the datasets and the corresponding text prompt used for adapting CLIP. For OxfordPets, Flowers102, Food101, SUN397, DTD, EuroSAT, and UCF101, we used the data splits provided by . For other datasets, we used the officially provided data splits.

A.3 Change Log

In ArXiv v1, we overlooked the adversarial reprogramming literature. In fact, visual prompting for vision models is essentially the same as adversarial reprogramming! In the current version, we have clarified this and removed claims of methodological novelty. We have reframed the paper as an exploration of the viability of visual prompts as a practical adaptation method for modern large-scale models. We thank Seong Joon Oh and users on twitter for pointing out the connection to adversarial reprogramming.