ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

Yoad Tewel, Yoav Shalev, Idan Schwartz, Lior Wolf

Introduction

Deep learning has led to at least three major revolutions in computer vision: (i) machines that achieve, in multiple domains, what is considered a human level of performance earlier than anticipated , (ii) effective transfer learning, which supports rapid modeling of new domains , and (iii) a leap in unsupervised learning through the use of adversarial and self-supervised learning .

A fourth revolution that is currently taking place is that of zero-shot learning. A seminal work by OpenAI presented the transformer-based GPT-3 model . This model is trained on extremely large text corpora and can then generate text given a prompt. If the prompt contains an instruction, GTP-3 can often carry it out. For example, given the prompt “Translate English to French: typical →\rightarrow typique, house →…\rightarrow\dots ” would generate the word “maison.”

Impressive zero-shot capability was later on demonstrated, also by OpenAI, in computer vision. While state-of-the-art computer vision models are often trained as task-specific models that infer a fixed number of labels, Radford et al. have presented the CLIP image-text transformer model, which can perform tens of downstream tasks, without further training, with an accuracy comparable to the state of the art. This is done by selecting, given an image, the best match out of sentences of the form “This is an image of X.” Subsequently, Ramesh et al. presented a bi-modal Transformer termed DALL-E, which generates images that match a given description in unseen domains with unprecedented performance.

In this work, we employ CLIP to perform the inverse task of DALL-E, namely zero-shot image captioning. Given an image, we employ CLIP together with the GPT-2 language model (we do not have access to GPT-3) to generate a textual description of the input image. This adds a new image-analysis capability to CLIP, beyond the fixed-prompt zero-shot learning demonstrated by Radford et al.

As a zero-shot method, our approach does not involve any training. One can argue that the underlying CLIP model is trained with exactly the same type of supervision that image captioning methods are trained on, i.e., pairs of matching images and captions. However, image captioning methods are trained from curated sources, such as MS-COCO or Visual Genome , while CLIP is trained on WebImageText (WIT), which is an automatically collected web-scale dataset. Previous attempts to train a captioning model on WIT have led to poor performance in recognizing the objects in the image, see Sec. 2.

As a result of the difference in both methodology and underlying data, the captions produced by our method are very different from those obtained by the supervised captioning methods. While supervised methods can mimic human annotators and provide similar sentences, in terms of conventional NLP metrics (such as BLEU ) to the ground truth sentences, our results exhibit much more freedom and match the image better in the visual-semantic CLIP embedding space (ours is optimized for this). Moreover, the semantic knowledge incorporated into CLIP and GPT-2 is manifested in the resulting caption, see Fig. 1.

In addition to the different nature of the obtained captions, our method is also more flexible, since all the computing occurs at inference time. Specifically, we show the ability to perform semantic analysis in image space by using a new form of arithmetic. A well-known example for concept arithmetic in NLP is that of retrieving the word ‘queen’ as the closest word, in the embedding space, to the equation involving the embedding vectors associated with ‘king,’ ‘man,’ and ‘woman,’ after subtracting the 2nd from the 1st and adding the 3rd. We present the novel ability to do the same, only with images instead of words, such that the result is generated as a short sentences, and not just a word, see Fig. 1.

As a corollary, we can, for example, ask what the difference is between two scenes. This ability to compare two images semantically is a novel computer vision capability, which further demonstrates the power of zero-shot learning.

Related work

The first deep captioning methods applied RNNs to generate sequences of words . Attention was added to identify relevant salient objects . Graph neural networks and scene graphs incorporated spatial as well as semantic relationships between objects . Subsequently, Transformers modeled interactions among all image elements with self-attention . On the text modeling side of the problem, language models (LMs) have also advanced with the development of LSTMs , CNNs and Transformers . Language improvements include devising better image grounding , decoding non-visual words (e.g., ‘the,’ ‘and’) , generating fine, novel and diverse sentences , and incorporating information from different semantic taggers .

In recent years, significant improvements have been achieved by utilizing large-scale vision-language data sets. The unsupervised data is used as a pre-training phase, to initialize models with image-text correspondence . With this technique, millions of image and text pairs from the web can be adopted. Nevertheless, in previous work we are aware of, all captioning models employ human-annotated datasets, such as MS-COCO or the Visual Genome, in the last stage of training.

It is likely impossible to construct a database of curated captions that is large enough, to describe even a modestly large fraction of plausible images and objects. This results in biases . Several approaches focused on describing novel objects by conditioning the model on external unsupervised data during training . Alternatively, external object taggers can be used during different phases (e.g., pre-training, training, or inference) . Semi-supervised methods are also available . Unsupervised approaches can be achieved by training with a visual concept detector or by learning a joint image-language embedding space . In contrast, our method makes use of an existing image-text alignment score to direct an existing large-scale LM toward a given image without training.

CLIP is trained on 400M images/sentence pairs from the web , resulting with a powerful text-image matching score. Originally CLIP’s authors explored training an image-to-caption language model with this training set, but found that it struggled with zero-shot transfer. In a 16 GPU-day experiment, a language model only achieved 16% accuracy on ImageNet . CLIP achieves the same level of accuracy roughly 10x faster.

Using prompts, it is possible to imitate some capabilities of text generation. For example, CLIP-based applications exhibit zero-shot solving capabilities in various scenarios never seen before. With careful engineering of the prompt, one can, for example, improve detection of unseen objects . Zero-shot prompt engineering has also been used for higher-level tasks (e.g., VQA), but it is nowhere near the level of supervised methods .

CLIP also provides powerful means for supporting text-driven image manipulation with Generative Adversarial Networks (GANs) or other generative models . Our work explores the other direction: generating text using an image, by guiding a large-scale LM with CLIP.

Guided language modeling has become a primary challenge, as researchers strive to tune prior knowledge within large-scale LMs, such as GPT-2 . Fine-tuning is often accomplished by employing Reinforcement Learning or GANs for each attribute separately. Disentangling the latent representations into style and content is also relevant in terms of text style transfer . A controllable LM can also be formed using fixed control codes . Ideally, conditioning should be applied directly to the existing large-scale LM, without the need for fine-tuning. Several studies have explored the idea of steering an LM using small neural networks . Following that, PPLM demonstrated that a simple attribute classifier could steer a model without any further training. With our work, we present a novel visual LM guidance from visual cues.

Method

Visual captioning is the process of generating a descriptive sentence for an image. It can be formalized as a sequence generation problem given an input image II, i.e., as a conditional probability inference for the ii-th word xix_{i} of the sentence, i.e., p(xi∣[xt]t<i,I)p(x_{i}|[x_{t}]_{t<i},I).

This is typically accomplished in a supervised manner, by optimizing weights to reproduce ground truth sentences. However, since carefully curated datasets are small, and cannot adequately describe all images, the sentences generated often describe the content at the basic level of the objects present in the scene and sound artificial. Such problems can be mitigated with the use of web-scale datasets. We present a zero-shot method for guiding large-scale language models with a large-scale text-image alignment model.

Overview Our approach uses a transformer-based LM (e.g., GPT-2) to infer the next word from an initial prompt, such as “Image of a,” as illustrated in Fig. 2. To incorporate image-related knowledge to the auto-regression process, a calibrated CLIP loss LCLIP\mathcal{L}_{\text{CLIP}} stimulates the model to generate sentences that describe a given image. An additional loss term LCE\mathcal{L}_{\text{CE}} is used to maintain the next token distribution similar to the original language model. Optimization occurs during auto-regression, and repeated for each token.

Furthermore, the flexibility of our method enables the capturing of semantic relations through simple arithmetic of visual cues in CLIP’s embedding space. Finally, combining multi-modal encoders with our method allows knowledge to be extracted in a new way that mixes between text and images.

Language models In recent years, LMs have improved significantly and are getting closer to AI-complete capabilities, including broad external knowledge and solving a wide variety of tasks with limited supervision. A Transformer-based LM typically models interactions between the generated token and past tokens at each time-step.

Recall that the transformer block has three embedding functions K,Q,VK,Q,V . The first two, K,QK,Q, learn the token interactions that determine the distribution over VV. The attention mechanism pools values based on the similarity between queries and keys. Specifically, the pooled value for each token ii depends on the query associated with this token QiQ_{i}, which is computed using the function QQ over the current embedding of this token. The result is obtained as the weighted average of the value vectors, based on the cosine similarity between QiQ_{i} and the keys associated with all tokens KjK_{j}.

While KK and VV are functions, the obtained key and values KjK_{j} and VjV_{j} are used repeatedly when generating text, one word at a time. KjK_{j} and VjV_{j} can therefore be stored in what is called a context cache, in order to keep track of past embedding outputs of KK and VV. The sequence generation process can then be written as

where xix_{i} is the ii-th word of the generated sentence, Kjl,VjlK_{j}^{l},V_{j}^{l} are the context transformer’s key and value of the jj-th token, and ll indicates the index of the transformer layers, out of a total of LL layers. Our method employs GPT-2, which has L=24L=24 layers.

We next describe how we align our LM with the input image. We do so by modifying, during inference, the values of the context cache Ci=[(Kjl,Vjl)]j<i,1≤l≤LC_{i}=[(K_{j}^{l},V_{j}^{l})]_{j<i,1\leq l\leq L} leaving the LM unchanged.

CLIP-Guided language modelling Our goal is to guide the LM towards a desired visual direction with each generation step. The guidance we propose has two primary goals: (i) alignment with the given image; and (ii) maintaining language attributes. The first goal is obtained through CLIP, which is used to assess the relatedness of a token to an image and adjust the model (or, rather, the cache) accordingly. For the second goal, we regularize the objective to be similar to the original target output, i.e., before it was modified.

The solved optimization problem adjusts the context cache CiC_{i} at each time point and is formally defined as arg min⁡CiLCLIP(LM⁡(xi,Ci),I)\operatorname*{arg\,min}_{C_{i}}\mathcal{L}_{\text{CLIP}}(\operatorname{LM}\left(x_{i},C_{i}\right),I)

where x^i+1\hat{x}_{i+1} is the token distribution obtained using the original, unmodified, context cache. The second term employs CE loss to ensure that the probability distribution across words with the modified context is close to the one of the original LM. The hyperparameter λ\lambda balances the two loss terms. It was set to 0.2 early on in the development process and was unmodified since. Next, we explain how the CLIP loss term is calculated.

CLIP loss We calculate image relevance for the possible tokens at time ii. It is sufficient to compute potentials for the top 512 token candidates and set the rest to zero potential for efficiency. To this end, the corresponding candidate sentence sik=(x1,...,xi−1,xik)s_{i}^{k}=(x_{1},...,x_{i-1},x_{i}^{k}) for the kk-th candidate token is matched against the image II.

The clip potential of the kk-th token is computed as

where DCLIPD_{\text{CLIP}} is the cosine distance between CLIP’s embeddings of the text (i.e., ETextE_{\text{Text}}) and the image (i.e., EImageE_{\text{Image}}), and τc>0\tau_{c}>0 is a temperature hyperparameter that controls the sharpness of the target distribution. In all our experiments, we set τc\tau_{c} to 0.01.

The CLIP loss is defined as the cross-entropy loss between the clip potential distribution and the target distribution of the next token xi+1x_{i+1} obtained by the language model:

This loss encourages words that lead to higher CLIP matching scores between the image and the generated sentence.

Inference As a zero-shot method, no training takes place. At inference time one optimizes the problem in Eq. 2, which we denote as p(xi+1∣Ci)p(x_{i+1}|C_{i}), by conducting five steps of gradient descent, i.e.,

This update rule is simplified for brevity. With each newly-generated token, the optimization is re-done. In our implementation, the gradients are normalized with Euclidean normalization before each step, separately for each transformer layer. We set the learning rate α\alpha to 0.3.

Beam search The byte-level tokenizer used employs 256 bytes of base tokens to represent every word in existence . Any word can also be split into more than one subwords, e.g., the word ‘zebra’ is tokenized as ‘zeb’ and ‘ra’. As a result, we found that images of zebras are described as striped animals, since the token ‘zeb’ is not picked. Beam search inference helps solve this problem by enabling the search to be conducted in a less myopic way.

Visual-Semantic Arithmetic

Recent studies suggested that CLIP multi-modal representation holds an elaborate concept taxonomy . In accordance with this intuition, we find that our method can express CLIP’s embedding in a textual way. For instance, subtracting between CLIP-encoded images and applying our method transcribes a relationship between the two images. Furthermore, by summing vectors we can steer the generated caption towards a conceptual direction.

To perform arithmetic in CLIP’s embedding space, we first encode the image/text using CLIP’s image/text encoder. For instance, let I1,I2I_{1},I_{2} be two images. We encode the images with CLIP’s encoder, i.e., Eimage(I1),Eimage(I2)E_{\text{image}}(I_{1}),E_{\text{image}}(I_{2}). Next, we carry out the desired arithmetic, e.g., addition with Eimage(I1)+Eimage(I2)E_{\text{image}}(I_{1})+E_{\text{image}}(I_{2}). Finally, we use the obtained result instead of the image encoding Eimage(I)E_{\text{image}}(I) within Eq. 3 to steer the generated sentence.

Consequently, we can generate detailed knowledge of the external world by moving in conceptual directions. This way, our method can answer questions expressed visually, for example, “who is the president of Germany?” To achieve this, we subtract “America’s flag” from an image of “Obama” and obtain a presidential-direction, to which we can then add the image of a second country’s flag.

Our approach extends beyond visual interactions alone. Using CLIP’s textual encoder, interaction with a natural language is possible. In this case, one performs arithmetic operations in the embedding space such that the expression contains both image- and text-embeddings.

Experiments

For all the results reported in this section, we used a strategy for reducing repetitions, in which the probability for generating tokens that were generated at the last four time-steps was decreased by a factor of two. We also incorporated a mechanism that directly controls the length of the generated text by multiplying the probability of the end token by a factor of fef_{e}, starting from time-step tet_{e}. We use fe=1.04f_{e}=1.04 and te=3t_{e}=3 for image captioning, and fe=1.06f_{e}=1.06 and te=1t_{e}=1 for image arithmetic. On a single Titan X GPU, five beams and 512 candidate tokens can be generated in three seconds. Inference time is proportional to the number of candidates and beams.

We begin by studying our zero-shot method for caption generation. Notably, we find our captions to exhibit human-like characteristics, such as generating diverse captions, reading, exploiting a wide range of external knowledge, and coping with abstract concepts. In Tab. 1, we present our results for COCO’s test set . Two recent baselines that use CLIP’s embedding are compared to: ClipCap and CLIP-VL . In ClipCap, the image is encoded using CLIP and the representation is transferred and plugged as a token into a fine-tuned GPT-2. CLIP-VL incorporates spatial grid features from CLIP into a transformer network. Another method, VinVL is a state-of-the-art technique.

We first consider supervised metrics, i.e., metrics requiring human references. These metrics include the BLEU , METEOR , CIDEr , SPICE , and CLIPScoreRef that we discuss below. As can be seen, our method lags in these metrics in comparison to the supervised captioning methods. Since the ground truth human annotation is obtained similarly to the training set, with the same group of annotators using similar terms, there is a clear advantage for methods trained on COCO annotations.

We next consider diversity metrics. Our vocabulary over COCO’s test set is significantly larger than previous approaches (8681 vs. 2464). In addition, none of the generated sentences appear in the training set of COCO (100% on %Novel).

CLIPScore is a reference-free method for evaluating relatedness between an image and its caption, using CLIP’s alignment score. Evidently, our method is much better in this metric than the supervised method (87% vs. 77%). As an alternative to exact correspondence with human reference, we use CLIPScoreRef to measure the semantic distance from the references. Although supervised methods outperform our method in this score (similarity in the vocabulary and the sentence style still provide an advantage), the gap is narrower than in other supervised metrics.

Qualitative Analysis Fig. 3 compares our zero-shot approach with other baselines, demonstrating that our method can generate human-like captions, i.e., textually richer, better at image reasoning, and more effective at grounding objects. We discuss each image from left to right. First, as opposed to CLIP-VL, which assumes a toilet is in the bathroom, and VinVL, which disregards the background buildings and presumes it is on a sidewalk, our method determines it is on a rooftop. Next, our method attempts to generate the written text on a boat’s side. The following image describes a flight meal as a regular tray of food with the baselines, whereas our method describes it as a flight meal. We accurately describe the next image as a bar restroom with portraits and not a bathroom. Our method and VinVL specify specific birds in the following photo (red falcons and hawks are hard to tell apart). Next, the baselines repeat the same sentence, while our method mentions an interesting mesh tile pattern. In the next photo, our method identifies a family rather than a general group. Last, our method accurately describes a room’s interior, such as a bedroom with posters, and deduce that the posters depict bands. Note that the baselines’ captions are generally of the same pattern, while our method generates novel sentences. Also, note that the images are taken from COCO dataset, which was used to fine-tune CLIP-VL and VinVL.

OCR The ability of CLIP to classify text within an image from a closed set of possible prompts is impressive . We show in Fig. 4 that these capabilities can be exploited in a generative manner. To accomplish this, we change the prefix prompt we use in our method from “Image of a” to “Image of text that says.” Results include impressive understanding. e.g., “The president Kennedy’s death” from an image of a paper declaring it or generating “The University of Stanford” from a sign depicting its name.

External knowledge The generated captions comprise a wealth of real-world knowledge on a variety of topics. In Fig. 5 we show samples of famous people (e.g., Trump), animated shows (e.g., Simpsons), cities (e.g., Manhattan), movies (e.g., Avengers), games (e.g., Mario driving), and places (e.g., Stanford).

2 Visual-Semantic Arithmetic Study

We demonstrate how our method can generate text for subtraction to explain semantic directions. Next, we demonstrate that the summation operator allows guidance of the generated text through visual cues. One can then apply the above insights to solve visual analogy puzzles.

Subtraction Subtracting vectors intuitively represents a direction between the vectors. In Fig. 6 we demonstrate our method’s ability to express relations through several examples. “A caricature illustration” is the result of subtracting a real photo of an airplane from a caricature. To put it another way, adding the concept “A caricature illustration” to the right image of a real plane will match the image on the left of a caricature plane. Concepts of quantity and color can also be seen, for example, a comparison of a green apple versus a red apple yields ‘Red,’ and vice-versa, subtracting one basketball from many basketballs results in “a bunch.” Furthermore, we find directions related to a geographical area, e.g., ‘Snow,’ and ’Desert.‘ Further, a concept directly tied to day and night, and a concept of prison (i.e., ‘Jailed’). It should be noted that the operator is not symmetric, and cannot always be derived textually. For instance, on the right images, the concept direction from a skateboard to a skateboard tournament can be generated as “The event.” However, the direction from a skateboard tournament to a skateboard generated “schematic fossil view,” which is irrelevant.

Summation Through the addition operation, the generated text can be guided through visual semantics. In Fig. 7 we show examples of guidance. On the left side, with the addition of a police officer’s hat, the caption describes a man running as “A police officer…,” if we add a hammer to a man, we get “The judge.” On the right side, we show that a concept can be abstract. For example, the Apple company can be represented by an apple. Thus, adding an apple to a phone, results in the text “Apple’s iPhone released.” Additionally, a country’s concept can be represented visually with flags. If Canada’s flag is added to a tree, “Toronto Maple” results.

Guidance with Visual-Relations In the field of natural language processing, semantic relations have long been studied . Previous efforts studied visual relations with expensive annotated language-priors . With the introduction of CLIP, richer visual concepts from large-scale data became available . Through visual arithmetic, we are able to exploit this richer embedding space.

In Fig. 8, we show our proposed strategy. Using subtraction, we first determine the direction. For example, the concept of leadership is represented by an image of Obama minus the American flag. With this direction in hand, we can now manipulate the case of other nations. By adding the direction to the German flag, we obtain “Angela Merkel.” A different example is to examine the concept direction of CEO-to-company. With different images (e.g., Bill Gates and Microsoft, Jeff Bezos and Amazon), the direction can be summed to Mark Zuckerberg and Steve Jobs generating ‘Facebook’ and ‘Apple,’ respectively. On the right side, we study various interactions with country-related representations. We guide the image of a baguette to generate ‘France’ by taking photos of pizza and Italy and deriving the country-to-food direction.

The Visual Relations benchmark To further study the relation capabilities of our technique quantitatively, we introduce a new benchmark of visual relations, VR for short. This benchmark comprises 320 relations of the following templates: buildings→\rightarrowcountry, countries→\rightarrowcapital, foods→\rightarrowcountry, leaders→\rightarrowcountry, and CEO→\rightarrowcompany. These were chosen because they are roughly many to one, i.e., a country has many buildings, but a building only relates to one country. The benchmark is designed to measure both the ability to model relations visually and to apply real-world knowledge to perform the task.

We constructed the benchmark through the following steps: (i) we created semantic directions by subtracting visual pairs and (ii) we then used each direction and added it to a visual element in another pair to create its corresponding text companion. As an example, we used images of (’japan,’ ’sushi’) to convey the direction of food→\rightarrowcountry, and then we added this direction to an image of a pizza and examined the appearance of Italy in the generated text.

We focused on single-word answers. The three evaluation metrics we find relevant to this setting are (1) BLEU-1, which measures unigram precision; (2) Recall@5, which indicates a word’s appearance within the first five words generated; and (3) CLIP-score, which indicates semantic relatedness. To calculate the CLIP-score, we first add “Image of” as a prefix to the ground truth. Using CLIP’s textual encoder, we then use a cosine distance. More details are provided in the supplementary material.

In Tab. 2, we show performance for each relation. While this task is challenging, our approach resulted in a significant success rate of 30% at R@5 in most relations. Note that, since the benchmark lacks multiple references, it is still limited, e.g., we mark a miss if the generated word is ‘US,’ while the ground truth is ‘USA’. Observing the returned answers reveals that some mistakes are understandable, e.g., answering Sydney instead of Canberra or the Sinai province instead of the country Egypt. However, other cases return truncated sentences, e.g., returning ‘flag’ instead of a country name or returning general concepts such as “flickr image”. See supplementary for a discussion. When employing the softer CLIP-Score metric, which is based on a semantic distance, a correlation of 70% is observed.

We compared our results with ClipCap that encodes the image with CLIP’s image encoder and uses it as an initial token for GPT-2. The method is fine-tuned based on COCO dataset. As can be seen, this method fails to retrieve the correct response, despite employing the same large-scale models as we do and performing arithmetic in the same CLIP embedding space. CLIP-VL and the supervised captioning methods cannot be tested on this benchmark since it uses spatial grid features as embedding.

Multi-modal Arithmetic Our method enables multi-modal reasoning, which involves manipulating images and text simultaneously in the same embedding space. Using CLIP’s textual encoder, ETextE_{\text{Text}}. In Fig. 9, we show that a day-to-night direction can be obtained with text inputs, i.e., “image of a night,” and “image of a day.” The direction steers an image of breakfast to “Nighttime dinner.”

Discussion and Limitations

The zero-shot capabilities presented by CLIP pave a new path for computer vision. However, these are limited to multiclass classification. DALL-E presents an impressive ability to generate images that are very different from its training images in what is termed zero-shot generation ability. However, this ability is exactly the generative task DALL-E was trained to do, only in new domains. No previous computer vision work, as far as we can ascertain, has presented a generative semantic zero-shot capability of the sort that is revolutionizing the NLP world with transformers, such as GPT-3 . Our work is the first to present a generative visual-semantic work.

While the ability to rely on pre-trained models such as GPT-2 and CLIP allows us to achieve such new capabilities, they also highlight the uneven playing field AI has become. GPT-2 is far inferior to GPT-3 and other recent LMs in which resources far beyond the reach of most research labs are invested.

On a similar note, it is likely that combining zero-shot with supervised training would lead to a method that outperforms the baselines in all captioning metrics. However, the amount of resources currently used to train supervised captioning methods is becoming a deterring factor from pursuing this direction. For instance, UNITER uses 3645 hours of a V100 GPU .

The use of an LM and an image-language matching model trained on large corpora of collected data inevitably leads to biases. For example, the models we employ are clearly oriented towards Western knowledge and can recognize people, places, objects and concepts that are popular in Western media, while being much less knowledgeable about other cultures. For example, our model fails to form relations with the president of China, Xi Jinping.

Conclusions

The marriage between a language model and a visual-semantic matching model is a powerful union, with the potential to provide zero-shot captioning that brings together real-world variability in text, recognition abilities that are unrestricted by categories, and real-world knowledge that is embedded in the models through web-scale datasets.

We propose a zero-shot method for combining the two models, which does not involve optimizing over the weights of the models. Instead, we modify, for all layers and attention heads, the key-value pairs of the tokens generated by the language model up to each inference step.

As a captioning model, our method produces results that are less restrictive than those provided by the human annotators on the datasets used by supervised captioning methods. While this lowers the word-to-word metrics, the captions generated seem to be a good match to the image at the semantic level and exhibit real-world information. Moreover, the flexibility of using an embedding-space zero-shot method enables us to perform visual-semantic arithmetic.

We show how we can describe in words the difference between two images and how we can combine concepts from multiple images. Both are novel high-level recognition tasks. Combining these two capabilities, a powerful image analogy machine is obtained, which answers, by providing a text string, questions of the form “A is to B as C is to X” (X∼C+B−AX\sim C+B-A), in which A, B, and C can each be either textual or visual.

Acknowledgments

This project has received funding from the European Research Council (ERC) under the European Unions Horizon 2020 research, innovation programme (grant ERC CoG 725974). The contribution of the first author is part of a PhD thesis at Tel Aviv University.

References

Appendix A Supplementary Material: ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

This supplementary material describes our experimental setup (see Appendix B), provides additional ablation study (see Appendix C), provides additional qualitative results (see Appendix D), explores the limitations of our approach (see Appendix E), and discusses visual relation benchmark failure cases (see Appendix F).

Appendix B Experimental Setup

As part of our experiments, we used COCO’s validation set (Karpathy splits) for both qualitative and quantitative evaluations. We report the beam with the lowest CLIP loss score among the five beams. Our model has several hyperparameters: (i) λ\lambda (see Eq. (2)), which was set to 0.2; (ii) τc\tau_{c} (see Eq. (3)), which was set to 0.01; (iii) α\alpha (see Eq. (5)), which was set to 0.3; (iv) We decreased the likelihood of repeated tokens by a factor of two in order to mitigate repetitions. Based on a human assessment, these parameters produced concise, fluent, and image-related captions. We use the PyTorch framework .

Pre-trained models: As part of our approach, we use two large-scale pre-trained models: (i) GPT-2, using HuggingFace’s gpt2-medium implementationhttps://huggingface.co/transformers/model_doc/gpt2.html, with 24 attention models and 345M trainable parameters. This model was trained on an 8M web-page dataset with a causal language modeling (CLM) objective; (ii) CLIP, trained on 400M (images, text) crawled from the web. We use the OpenAI implementationhttps://github.com/openai/CLIP. We employed a version of CLIP with a vision transformer image encoding architecture that is equivalent to ViT-B/32 .

Prompt engineering: Our method begins with an initial prompt. In the majority of our experiments, we used “Image of a”. We determine the caption from the words generated after the initial prompt. We did observe that the prompt affected output results, e.g., “Image of text that says,” is much better if the caption is intended for OCR.

Appendix C Ablation Study

Effect of CLIP-based optimization: A further ablation was performed, in which CLIP’s score is used directly to optimize the LM. In Fig. 10, we show two variants: (A1) selecting tokens one by one to maximize the CLIP score, and (A2) doing so on a score that combines CLIP score with an LM-score. Evidently, the captions are not competitive with our method. We also assessed the differences in language fluency (perplexity measured with GPT Neo) and image correspondence (measured with CLIP Score). Despite a higher CLIP score (Tab. 3), our method has improved language fluency. It is worth noting that higher CLIP doesn’t necessarily translate to better wording.

A human study further supports this, conducted to determine which method is perceived as the best one. The study included 50 images randomly selected from COCO and 40 annotators. Our caption was selected 70.5% , (A1) 8.9%, and (A2) 20.6%.

Effect of regularizer coefficient: As shown below, an increase in the regularizer coefficient results in a decrease in the perplexity score measured with GPT Neo (i.e., language fluency improves) while it decreases the clip similarity. We find λ=0.2\lambda=0.2 to be a good trade-off point.

Human evaluation: We conducted an additional human study on 50 images. We picked the images from the web (e.g., video-game screenshot, real-world knowledge; specifically, the subreddit ‘i took a picture’). We asked the annotators to score between 1 to 5 two properties: human-like and visual grounding. We compared against a supervised method ClipCap. On human-like, our approach got 3.79 vs. 3.17 of ClipCap. On image grounding, our method got 3.98 vs. 3.21 of ClipCap.

Appendix D Additional Qualitative Results

Image Captioning: In Fig. 17 (shown at the end of the document due to size), we present our results on 200 randomly-selected images along with baselines. For baselines, we use ClipCap , CLIP-VL , and VinVL . Our method generates original captions that are completely different both in vocabulary and pattern from the baselines’ captions.

Appendix E Limitations

We detail both the caption quality issues and the biases resulting from the noisy web-scale data used to train CLIP and GPT-2 in the following sections.

Web-scale noise: The captions we generate are influenced by CLIP’s training data. Due to its extraction from the web without special care, it contains noise. This leads to two undesirable outcomes: 1) Generating entities related to the data source (e.g., Flickr) or irrelevant entities (e.g., the name of the photographer). We solve this problem by adding a negative prior regularization to any capitalized subword. Consequently, a more generic caption will be created, but at the expense of world-knowledge capabilities. We show samples with and without the mechanism in Fig. 11; and 2) At times, the captions become irrelevant because they fail to remain focused. This can be controlled using two hyperparameters. We multiply the probability of the end token by a factor of fef_{e}, starting from time-step tet_{e}. In our method we used fe=1.04f_{e}=1.04, and te=3t_{e}=3. In Fig. 12, several random examples are shown, and the length control mechanism is ablated.

Bias and Fairness: It is common for web-scale data to contain biased sources (e.g., news), resulting in an unintended bias against some ethnic groups. In Fig. 13, an abstract illustration of a terrorist is described as Palestinian. Another example, racial characteristics are used to portray a child as an immigrant. Additionally, a caption implies homosexual orientation for an image of two men.

Appendix F Visual Relations Benchmark Study

Our benchmark combines real-world knowledge with the ability to represent visual relationships. In Fig. 14(b). we show at typical mistakes. Samples are referred to by their character counter: (a) Unpopular real-world knowledge. GPT-2 and CLIP training are based on web crawled data. Consequently, it may choose words based on popularity on the Internet. Sydney is a more popular city than Canberra worldwide (we validate this with Google Trends); (b) Synonyms. The relationship between the president and his or her country leads to ”Canadian” rather than ”Canada;” (c) Closely related. Rather than relating the pyramids to Egypt, this sample refers to Sinai, an area in Egypt; (d), (e),(f) Relation mistake. Subtracting Australia from Canberra conveys a relationship relevant to a university. It appears that adding the relationship to the UK led to ‘Berkeley.’ A ‘Chinese university’ is generated by adding it to China, and a ‘German university’ is generated by adding it to Germany. This might be due to Canberra being known for its university. Since we use the same relation (pair subtraction) for multiple triplet of images, inferring the wrong relation can lead to many errors in the benchmark.