What does CLIP know about a red circle? Visual prompt engineering for VLMs
Aleksandar Shtedritski, Christian Rupprecht, Andrea Vedaldi
Introduction
Large Language Models (LLMs) such as GPT-2/3 and ChatGPT have demonstrated surprising emerging behaviours. For example, these models can perform language translation without being explicitly trained for it, in a zero-shot manner. This can be partially explained by the fact that occurrences of the desired behaviours, such as translating between two languages, naturally occur in their enormous training corpus, which is, essentially, the Internet.
Interesting emergent behaviours have been observed in large Vision-Language Models (VLMs) like CLIP too. For example, CLIP can be used for zero-shot classification by checking the compatibility of a given image with prompts such as “an image of a ”, where is one of a set of class hypotheses to be tested.
Emergent behaviours are elicited by supplying suitably crafted inputs to the VLMs, often called prompts. As in the example above, researchers have mostly focused on engineering textual prompts, manipulating the textual input of the model. This approach is inspired by LLMs, where manipulating the textual modality is the only available option. However, VLMs are inherently multimodal and offer the possibility of manipulating both modalities, textual and visual. While the textual modality is the natural choice for expressing semantics, the visual modality can be better for expressing geometric properties such as location.
In this paper, we thus explore visual prompt engineering Note the difference between visual prompt tuning, a setting previously explored, where the prompts are task-specific learnable tokens, and visual prompt engineering, where we apply a fixed augmentation in pixel space.. We do so with two goals. The first goal is to contribute one more practical tool for extracting useful information from VLMs in a zero-shot manner. We demonstrate this by obtaining state-of-the-art zero-shot results in referring expressions comprehension by engineering visual prompts. The second goal is to characterise interesting and unexpected properties of the VLMs and their training data, including identifying some behaviours that can raise ethical concerns.
Perhaps the most surprising of our findings is the effectiveness of a particular type of visual prompting: drawing a plain red circle on top of the image (Fig. 1). We show that this simple intervention steers the VLM to analyse/talk about the image region contained in the circle. This behaviour can then be used for tasks such as naming a specific object or object part or detecting particular image regions based on a description. The latter, for instance, is achieved by marking each object proposal with a red circle and using the VLM to find the best match with respect to the provided referring expression, achieving strong results on multiple benchmarks in the unsupervised regime. Furthermore, we show that prompting with a circle also works for finer-grained localization, marking specific object parts or keypoints instead of just whole objects.
We further contrast marking an image with the alternative of cropping it, which, from sliding window classifiers to region neural networks, is the canonical approach to steer the focus of an image-level predictor to a particular image region. We show that, for VLMs at least, marking is significantly more effective than cropping, possibly because it does not lose contextual information like the latter.
Apart from the practical applications, our findings reveal unexpected and intriguing properties of VLMs. We show empirically, that marking with a red circle is optimal among a selection of possible markers (variants of the circle, boxes, arrows, etc.). Presumably, the VLMs understand red circles out of the box because these appear sufficiently frequently in the training corpus, i.e., the Internet. While we do not have access to the full training data of CLIP, we corroborate this intuition by seeking examples of such images in YFCC15M, a dataset of CC-BY images.
Our analysis shows that red circles are indeed present even in a (comparatively small) dataset of images like YFCC15M, but they are rare. It is a testament to the extraordinary capacity of VLMs that such a behaviour can be learned from such rare events, without an explicit focus on doing so. We test models of different sizes/capacities and show that only the larger models exhibit this behaviour reliably, which corresponds to our intuition.
Finally, we note that the ability of VLMs to learn even from rare events such “red circles” can acquire both desirable and undesirable behaviours. Red circles, in particular, can have a negative connotation in the training data as they are often used by news outlets to mark missing people or criminals and, evidently, the model learns from such examples. As a result, we show that drawing a red circle in an image increases the probability that the model would characterise a person as a criminal or as a missing person.
To summarise, we make the following main contributions: (1) We propose marking as a new form of visual prompt engineering that is effective in extracting useful emergent behaviours in VLMs like CLIP; (2) We use the latter to achieve state-of-the-art zero-shot referring expressions comprehension using a VLM; (3) We provide an analysis of why marking is effective for these models, and link that to the training data and large model capacity; (4) We show that visual prompt engineering can also elicit unwanted behaviours, such as triggering problematic biases in the VLMs, revealing potential ethical issues.
Related work
Emergent Behaviour from Large Scale Pretraining has mainly been observed in Large Language Models (LLMs). Most notably, GPT-2 , GPT-3 , and ChatGPT have been shown to be capable of tasks such as zero-shot translation, question answering, arithmetic, as well as planning actions for embodied agents . Fine-tuning LLMs can also lead to models that can generate code from docstrings or solve math problems . Only a few emergent zero-shot behaviours have been reported for VLMs like CLIP, mainly for classification and OCR . Generative VLMs like FLAMINGO and BLIP excel in captioning and visual question-answering tasks, but also have no way of solving pixel-level computer vision tasks.
Prompting VLMs is most commonly performed by prepending a set of learnable tokens to the text input , vision input , or both text and vision inputs , in order to easily steer a frozen CLIP model to solve a desired task. learn augmentations in pixel space, such as padding around the image, or changing a patch of the image, which are optimized with gradient descent on a downstream task. cast image inpainting as a visual prompting task, using a generative model trained on figures from academic papers. Coloring regions of an image has been used for the VCR task , where a model is finetuned on annotated images . Colorful Prompt Tuning (CPT) color regions of an image and use a captioning model to predict which object in an image an expression refers to by predicting its color. Similarly to CPT, we augment the input image in pixel space and perform zero-shot inference. However, we annotate the image in a human-like manner and show that our method is more powerful and more flexible than CPT.
Referring Expression Comprehension (REC) aims to localize a target object in an image that corresponds to a textual description. Most approaches to REC start with object proposals, for example, generated with Faster-RCNN , and learn to score them . REC is sometimes considered together with referring expression generation — the task of generating a description of a given region. use a comprehension model to guide a generator, whereas jointly train a detector with a caption generator. Some works model the scene as a graph or use language parsers and grammar-based methods , leading to a more interpretable result. More recently, transformer architectures have been used . perform text-modulated object detection, where a transformer decoder takes the referring expression as an input and predicts a bounding box. train with a text-to-pixel contrastive loss, which allows for a text-driven segmentation or detection at test time.
Unsupervised Referring Expression Comprehension is a less explored area, only made possible with the introduction of large pre-trained models such as CLIP . ReCLIP crops object proposals and ranks them using CLIP before an ad-hoc postprocessing step to take into account relations such as left/right, smaller/bigger, etc. CPT colors object proposal boxes and use a pre-trained captioning model to auto-regressively predict which colored proposal corresponds to the query description. Pseudo-Q generates descriptions for multiple objects in an image, which is used to train a REC network. However, this model is not fully unsupervised as the pseudo descriptions it uses are generated using a captioning model trained on COCO.
Visual Reasoning Using Large Pretrained Models has been an area of significant interest in the last few years. In addition to referring expression detection , CLIP has been used for semantic segmentation . use CLIP to assign text labels to object parts after doing part co-segmentation in the latent space of a GAN. utilize CLIP for open-vocabulary segmentation by using a general-purpose mask proposal network and CLIP as a classifier. CLIP has also been used for unsupervised object proposal generation and open-set detection . Semantic segmentation also emerges from image only or image-text self-supervision.
Bias of VLMs is an increasingly popular area of research, as downstream applications come with the risk of perpetuating biases and stereotypes existing in the training data. However, methods for assessing the bias of a VLM are still not well established. measure the misclassification rate of CLIP of faces of people of different races with non-human and criminal categories, whereas measure fairness in retrieval results. Here, we show a different kind of bias, where the addition of a red circle over a person can trigger a negative connotation.
Method
One of the most striking capabilities of VLMs is their ability to solve a variety of classification tasks with little to no further training at all, in a zero-shot manner. This is done by reducing the task of interest to that of evaluating the VLM on suitably-engineered image and text pairs.
For example, given an image-caption pair , consider the problem of localizing a named object keypoint in the image. We can cast this as a question-answer problem, where the question is the name of the object keypoint (e.g., “right ear”, “front left leg”, …) and the answer is one of a discrete set of image locations.
Because the VLM computes a compatibility score between an image and the text , it cannot be used to map the question to the answer directly. However, via prompt engineering, we can use the VLM to construct a compatibility score between question and answer, conditioned on the input image-text pair . This score is in general given by the expression
where and are versions of the input image and text, obtained by transforming the latter to reflect the question-answer pair .
The specific way Eq. 1 should be applied to a problem depends on the specific nature of the latter. For example, in the problem of localizing the named keypoints, it is natural to encode the name of the keypoint via the textual modality and its 2D location via the visual modality. For instance, in order to answer the question for a given input image with caption , we can engineer the textual prompt to encode a description of the named entity. Likewise, we can engineer the visual prompt in such a way as to ‘select’ the location in the image, using one of the methods discussed in Section 3.2. With this, we can answer the question by finding that maximizes the score which specializes Eq. 1.
In the following sections, we provide further details and apply these ideas to a few concrete tasks.
2 Visual prompting via marking
The usual way of encoding location information in a visual prompt is to crop the image around the desired location, meaning that is the image cropped around . This idea has been used extensively with VLMs, including to interpret referring expressions, where maximizing a score of the form seeks for the image crop that best matches the referring expression .
In this paper, we explore an alternative approach for visual prompting that uses the concept of marking the desired region in the image. Marking quite literally means overlaying to the image a circle, a box, or an arrow, which visually indicates the desired location .
While the idea of marking may sound strange, it is interesting for two reasons. First, differently from cropping, a marked image preserves almost all the information contained in the input image , including contextual information that crops lack. Second, we show that marking works well with VLMs, outperforming cropping-based prompt engineering in some prediction tasks.
While the simplest marking consisting of a red circle is particularly effective, in Section 4 we explore several different ways of generating markings. We refer the reader to that section for further details and examples.
3 Tasks
We study the idea of mark-based prompt engineering by considering several zero-shot prediction tasks, from simple tasks such as matching keypoints to their names to more complex ones such as referring expression comprehension.
The first and simplest task that we consider is matching the name of the keypoints of an object to their 2D locations in an image. The input is an image , a set of keypoint names , and a set of corresponding keypoint locations . The number of names and locations is the same () and the goal is to match the two. We express the latter as predicting the square permutation matrix that associates each name to its corresponding location (i.e., ).
In order to predict , we use Eq. 1 to define the cost of associating name to location as where is obtained either via cropping or marking and is just the name of the keypoints prefixed by the string “an image of”. For this problem, the role of questions and answers is symmetric and we decode the cost matrix into a permutation matrix via optimal transport:
where is a temperature parameter. This optimization problem is solved efficiently via the Sinkhorn-Knopp algorithm , which renormalizes matrix .
Keypoint Localization.
The second task is a more useful and difficult variant of the first. The goal is still to localize a named keypoint in an image, but this time the locations are a subset of a regular grid. These are further restricted to a salient image region extracted by using the unsupervised saliency method of to avoid testing irrelevant locations in the background. The difference compared to naming keypoints is that this version of the problem does not assume prior knowledge of the possible locations of the keypoints. Given the name of a keypoint, its location is then obtained as where and are as defined previously.
Referring Expression Comprehension.
Comprehending a referring expression means detecting an object in an image that corresponds to a textual specification that explicitly refers to it (e.g., “fourth dog from the right”). Similarly to prior work , given an image , we approach this problem by extracting first a set of object proposals using the method from and interpret those as the set of possible answers . The set of questions is instead a collection of referring expressions extracted from a given benchmark dataset. For each referring expression, the best matching proposal is then given by
The engineered prompts and are defined as in Section 3.3. In this case, we found it useful to subtract from the score the average with respect to all possible referring expressions . This weighs down hypotheses such as faces that are visually very salient and tend to respond very strongly to all questions .
Experiments
We study the properties of visual marking in VLMs by considering first the three tasks of Section 3.3: naming keypoints, localizing keypoints, and referring expression comprehension.
Naming keypoints is a comparatively simple problem that has no direct application; however, it is simpler and faster to evaluate than the other tasks, so we use it to ablate various aspects of our method.
For this task, we consider the CUB-200-2011 (CUB) and SPair71k datasets. The first contains named keypoint annotations for each image, whereas the second only annotates matching keypoints in pairs of images, but does not name them. We thus augment the latter, manually naming each keypoint instance in each animal image. We further crop the images from SPair71k with the provided bounding boxes. For the VLM, we use the ViT-L/14@336px backbone. Please see the sup. matt. for details.
Results.
Recall that, in this task, the output of the predictor is a permutation matrix associating each keypoint location to a corresponding name. We report (i) the ratio of keypoint names that are mapped to the correct location and (ii) the ratio of keypoint locations that are mapped to the correct names. To the best of our knowledge, there are no prior works that associate keypoints with their names. We thus compare the result of this new task to (a) random choice and (b) a baseline where is obtained by cropping.
As seen in Table 1, prompting via visual marking (red circles) significantly outperforms the baselines, achieving almost twice the accuracy. Using the Sinkhorn-Knopp (SK) algorithm to normalize the matching score further boosts results, mainly improving results for points that are ambiguous and close to each other, e.g., mouth and nose.
What is the best visual marker?
We compare the use of (i) different shapes for highlighting a location: circle, rectangle, cross, arrow, (ii) different sizes, and (iii) different colors of the annotations, and show some examples in Fig. 4. We compare different shapes and colors in Table 2 and find that red circles perform best. Red is the best color despite the fact that it is a commonly occurring color in images, unlike colors like purple which can be found less often in nature and can thus be more distinctive, but lead to worse performance. We attribute this to the fact that this emergent capability of CLIP exists due to human-centric manipulations of its training data, and humans are likely to annotate using red circles, as shown next.
Are there visual markers in the training data?
To explore the hypothesis that CLIP can zero-shot classify annotations on images because of similar examples seen during training, we find images in YFCC15M that contain markers (YFCC15M is a subset of the CLIP training data). To this end, we train a binary classifier using an ensemble of a ViT-B/16 and RN50x16 CLIP vision encoders to classify images in YFCC15M that contain annotations. We then use this to filter a 6M subset of YFCC15M and take the top 10k images with the highest score. Finally, we manually examine the 10k images and find 70 images that have annotations drawn on top of them. We show 3 such images in Fig. 5. Hence, the training data contains examples of markers, but they are very rare (0.001%), suggesting that such behavior can only be learned from very large datasets by high-capacity models. This is further explored next.
How do different VLMs differ?
We compare a number of CLIP models in Fig. 6. In general, we observe that the performance of keypoint matching improves with (i) the size of the pretraining dataset and (ii) the size of the vision encoder. The former holds true for CLIP models trained on WIT-400M vs YFCC-15M (which is a subset of WIT-400M). However, using the LAION-2B dataset for pretraining leads to worse results. We suspect this result comes from differences in filtering when creating WIT-400M and LAION-2B, where in the latter, examples of annotations might have been discarded due to a stronger focus on aesthetic images for generative models. Similarly, we see big gains in performance as we increase the size of the vision encoder, and the gains do not seem to converge with the biggest available models. We emphasize on the dramatic increase in performance of the WIT-400M pretrained CLIP — the biggest models improve on the performance of the smallest by 250% on ZS keypoint matching, whereas the improvement on ZS ImageNet-1K classification is just 20%. We draw similarities between this task and tasks in the domain of NLP, such as zero-shot or one-shot arithmetic, where only the largest GPT models perform well . We argue that in a similar fashion, the vision encoder needs sufficient capacity and data in order to show this emergent behavior.
2 Localizing Keypoints
For this experiment, we use the same data and network architecture as for the previous one, but report the percentage of correct keypoints (PCK) as a metric, as the latter is widely used when evaluating semantic correspondences. Given a set of ground-truth points and predictions , PCK is given by:
Here, is a distance threshold given by , where is a ratio and is the bounding box size. For all datasets we use . Keypoint localization also utilizes an unsupervised saliency mask to ignore background locations.
Similarly to the naming task, we compare keypoint localization to random guessing and the crop-based baseline. As shown in Table 4, using red circles significantly outperforms both; as expected, results are further improved by using saliency to further filter keypoint locations. We show qualitative results in Fig. 3.
3 Referring Expression Comprehension
Referring expression comprehension is commonly evaluated on the RefCOCO , RefCOCO+ , and RefCOCOg datasets, all of which consist of images from the MS-COCO dataset together with expressions that refer to a unique object in the image, which are also annotated with a bounding box. RefCOCO+ only contains appearance-based expressions, whereas RefCOCO and RefCOCOg contain relation-based expressions (e.g., containing the words left/closer/bigger). The test sets of RefCOCO and RefCOCO+ are split in two, where “testA” and “testB” contain only people and non-people, respectively. We evaluate using the percentage of correct predictions, where a box is correctly predicted if its intersection-over-union with the ground-truth box is over 0.5.
For the referring expressions task, we use an ensemble RN50x16 and ViT-L/14@336 CLIP backbones. Following prior work , we score the bounding box proposals of MAttNet .
Results.
Using a red circle, we achieve state-of-the-art on most referring expressions comprehension baselines in the zero-shot setting, as shown in Table 5. Interestingly, this even outperforms ReCLIP , which is based on scoring image crops, followed by post-processing with manually designed relations rules. A red circle also outperforms Pseudo-Q on most benchmarks, even though Pseudo-Q explicitly trains for this task.
4 Model biases and ethics
While drawing circles on images can extract useful behaviors from a VLM for a wide variety of legitimate image analysis tasks, it can also extract unwanted ones and must not be used for the analysis of sensitive data.
To demonstrate this fact, in Fig. 8 we take a random image from COCO that contains a male-looking and a female-looking individual and zero-shot classify the image, as well as the image with circles over each individual, into 4 categories: male, female, missing person, and suspected murderer. While this leads to a correct resolution of the apparent gender of the annotated person, the annotated images are more likely to be classified as containing a missing person or a murderer. While we cannot know for certain, we hypothesize that this is due to the presence in the CLIP model training data of missing person reports, police footage, or similar, where people have been marked.
We further quantify these biases following , using synthetic faces from FaceSynthetics and person crops from COCO . For FaceSyntetics, we take 1000 random synthetic faces, and for COCO, we crop all bounding boxes for the class person from the validation set that have an area of at least 10% of the total area of the image, which comes down to 1352 crops. Following , we measure zero-shot classification rates into criminal categories. We introduce a “positive” category (honest man/woman/person), “neutral” category (man/woman/person) and a “criminal” category (criminal/thief/suspicious person). Finally, we zero-shot classify the original images and the images with circles. In Table 6 we present classification rates into criminal categories. We see that for all ViT encoders, the rate at which people are classified as criminals is significantly higher. This is problematic as such existing biases can lead to harmful consequences. Note that there are various limitations in this analysis, including the usage of binary gender.
Conclusions
We have shown that visual prompt engineering via marking can extract useful behavior from VLMs such as CLIP in a zero-shot manner, achieving state-of-the-art zero-shot referring expression comprehension performance, and significantly outperforming traditional techniques like image cropping. Our analysis suggests that this behavior emerges because relevant samples of marking exist in the training data of the VLMs, but these samples are very rare. As a consequence, the behavior can only be learned by very large models trained on very large datasets. The analysis also shows that VLMs acquire undesirable behaviors too, where the mere addition of a red circle to an image increases the model’s belief that the image has a negative connotation.
We use the RefCOCO, RefCOCO+, MS-COCO, FaceSynthetics, YFCC15M, CUB, SPair71k in a manner compatible with their terms. Some of these images may contain personal data (faces). In Sections 4.1, 4.2 and 4.3 there is no extraction of biometric data. In Section 4.4 we use MS-COCO to demonstrate that such a method cannot reliably extract information about people due to the bias in the pre-trained CLIP model (there is no identification). The FaceSynthetics, used for the same purpose, is a dataset of synthetic faces, so it does not raise privacy concerns. For further details on ethics, data protection, and copyright please see https://www.robots.ox.ac.uk/~vedaldi/research/union/ethics.html.
Acknowledgements.
We thank Luke Melas-Kyriazi, Tim Franzmeyer, Rhydian Windsor and Bruno Korbar for proofreading A. Shtedritski is supported by EPSRC EP/S024050/1. A. Vedaldi and C. Rupprecht are supported by ERC-CoG UNION 101001212. C. Rupprecht is also partially supported by VisualAI EP/T028572/1.
References
Datasets
As noted in the main paper, we contribute additional annotations to the Spair71k dataset for some of our experiments. We start from their keypoint annotations, which have no keypoint name annotations in the original dataset. We then manually name all keypoints of the animal classes in Spair71k, as shown in Table 8. We purposefully leave out some point annotations:
All animals have a left and right nostril annotated — we take the right one in all classes and annotate it as nose, and leave the left nostril out.
All tails have point annotations at the start of the tail (attached to the body) and end of the tail. Because of the lack of words to precisely describe both points, we take the point not attached to the body and annotate it as tail, and leave the other one out.
All ears have point annotations at the start of the ear (attached to the head), and at the pointy end. Because of the lack of words to precisely describe both points, we take the point not attached to the head and annotate it as ear, and leave the other one out.
Birds have annotations for (i) foot, (ii) ankle, (iii) knee, which are often ambiguous and very close together. We only keep the foot annotation.
Note we explicitly define different names for keypoints that can be ambiguous, e.g. eyes, ears, legs, etc. This ensures the role of questions and answers in Section 3.3 is satisfied.
Discovered annotations
Out of the discovered annotations in YFCC-15M, 44% contain red circles. Overall, 73% of the annotations were circles, and the rest were rectangles. 65% of all annotations were red, 10% yellow, 7% blue, 7% white, and the rest were black, green, and purple.
Additional implementation details
We base the evaluation of our method on ReCLIP , where an ensemble of two CLIP backbones is used — RN50x16 and ViT-B/32. We evaluated ReCLIP for all combinations of CLIP backbones in Table 7 and found that, on average, this is the highest-performing one. Similarly, for our method, we choose the ensemble of two backbones that lead to the highest performance — RN50x16 and ViT-L/14@336. Full comparison between the backbones can be found in Table 7.
Annotations
We experiment with different marker shapes, sizes, and colours, and present the results in Table 10. We find that, on average, a thin red circle leads to the best performance. We use an ensemble of the red circle annotation and two additional augmentations — blurring and gray-scaling the outside of the circle, for a total of three images per annotation, as shown. These augmentations were inspired by examples in YFCC15M we discovered that were annotated like that. We found that adding augmentations improves overall results. However, we did not explore including augmentations beyond these. We ablate these choices in Table 9.
Additional details
We augment the text queries by prepending “This is”. When subtracting the average with respect to other referring expressions, we use randomly sampled expressions.
2 Keypoint tasks
We evaluate different backbones in Table 3 in the main paper and find that ViT-L/14@336 performs best.
Annotations
We show examples of the markers we use in Fig. 4 in the main paper . We compare a large range of sizes and colors, as shown in Table 2 in the main paper. We find that a circle is the best marker, and drawing a cross over the point of interest is the worst. The best-performing marker out of all is a red circle, which is the one we end up using. In Fig. 10 we show a more detailed comparison of different colors, diameters, and thicknesses when using a circle annotation. We see that a thin red circle is the best-performing marker. We show what that circle looks like on an image in Fig. 11.
Given this, we draw red circles over the images, with radius and thickness , where H is the shorter side of the image. For the backbone we use, where the input size has px, this becomes px and px.
Additional details
For the keypoint localization task, we set , for a total of query locations before applying the pseudo mask. The templates we use are “This is the {part} of a bird” for CUB and “This image shows the {part} of the {animal}” for SPair71k. We use a temperature parameter .
Qualitative evaluations
We present qualitative evaluations on naming keypoints in Figs. 14 and 15, keypoint localization in Figs. 11 and 12 and referring expressions comprehension in Fig. 13.