Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality
Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, Kyle Mahowald
Introduction
Despite the success of large pretrained transformer models on a wide variety of tasks, the extent to which they are compositional (e.g., Kim and Linzen, 2020; Soulos et al., 2020; Hewitt and Manning, 2019; Sinha et al., 2021a; Clouatre et al., 2021) and grounded (Bender and Koller, 2020; Bisk et al., 2020) is debated. Taking compositionality and groundedness as key desiderata, the recent Winoground dataset (Thrush et al., 2022) provides a clever way to test multimodal vision and language models. Given two images and two captions, the goal is to pair them correctly. The key insight is that, inspired by the Winograd schema (Levesque et al., 2012), the two captions contain the same set of words/morphemes, only in a different order. Figure 1(A) shows a representative example.
Pretrained multimodal transformer models (e.g. Radford et al., 2021; Tan and Bansal, 2019; Chen et al., 2019) have achieved impressive performance in multimodal tasks like image retrieval, image captioning, and visual question answering, as measured on a variety of datasets (e.g., Johnson et al., 2017; Suhr et al., 2017; Bitton et al., 2021). But, on Winoground, they all fall down: not one performs meaningfully better than random chance—despite the fact that humans can easily do the task.
Citing evidence from Sinha et al. (2021a) that large language models don’t need word order information to do well on tasks (see also Sinha et al., 2021b; Hessel and Schofield, 2021; Pham et al., 2021; Gupta et al., 2021; O’Connor and Andreas, 2021), the Winoground authors suggest that models track word co-occurrences, thus giving “the illusion of an understanding of word order” without actually achieving that understanding (Thrush et al., 2022). Indeed, given that information about semantic meaning can be uncovered without word order information (e.g., Papadimitriou et al., 2022) and that seemingly syntactic and semantic tasks can be solved with lexical heuristics (e.g., McCoy et al., 2019; Sinha et al., 2021b), Winoground failures may offer another evidence that language models solve complex tasks in a relatively superficial way.
To assess this possibility, we examine the Winoground task and conduct a series of novel experiments on the dataset, testing three models (CLIP, Radford et al. 2021; UNITER, Chen et al. 2019; LXMERT, Tan and Bansal 2019) that reflect three broad categories of Transformer-based vision-and-language architectures. First, we test these models on the more general (and standard) text-to-image and image-to-text Recall@K task using the Winoground images and captions. Some models fail at this simpler task, suggesting that failure on Winoground may not be just because of a failure in semantic composition but due to broader difficulty with atypical images in the dataset. We show that even fine-tuning probes specifically on Winoground does not help, implying a potential absence of information necessary to succeed at the task.
Second, to understand what the source of failure might be, we develop a new taxonomy of Winoground examples consisting of six classes (Section 4). Our taxonomy reflects various abilities required to solve the task, and model’s performances vary significantly among our classes. Given an explosion of interest in testing image generation models (e.g., DALL-E 2 and Imagen) on their compositional ability (e.g., Marcus et al., 2022), Winoground can be a crucial benchmark, which motivates the need for a deeper analysis of its properties; prior work has performed such deep analyses for other benchmarks (Alt et al., 2020; Luccioni and Rolnick, 2022; Luo et al., 2022). We tag every example in Winoground with our scheme (see Appendix E for full reporting) and provide a performance breakdown of models. Figure 1(B) shows these tags. We show high variability in performances based on our proposed tags and observe low performance on tags that are challenging for reasons beyond compositional language understanding (e.g., a low-res version of the image simply lacks the visual detail necessary for answering the question). Thus, we conclude that not all Winoground items test what they aim to, and identify a subset of 171 items which directly measure compositionality.
Third, we run a series of probing experiments to better understand whether the failure arise because of failures in visual discrimination, in linguistic compositionality, or in the fusion of vision and language. Specifically, we augment the original captions with a set of textual variants Dhole et al. (2021). While these textual variants are indeed highly separable in embedding space, using them fails to improve the task performance. Figure 1(C) shows these textual variants.
Taken together, our results suggest that failures found on Winoground reflect meaningful model failures. While some Winoground items may be ill-suited to evaluate compositionality, even the most straightforward items pose a challenge. Our evidence suggests that the source of these robust failures lies in fusing visual and linguistic information, not strictly in complex language understanding. We hope our analysis will help future endeavors in interpreting emerging models’ Winoground performance.
Background
The Winoground dataset (Thrush et al., 2022) contains 400 items (each consisting of two image+text pairs with overlapping lexical content). The items were categorized linguistically based on whether the text swaps an object, a relation, or both. The items were further categorized based on if they involved: a Pragmatics tag indicating non-literal/pragmatic reasoning required, a Symbolic tag indicating reasoning about something in symbolic space (e.g., children’s drawing), and a Series tag (indicating whether the items come from the same, as opposed to from unrelated, photos).
Evaluated models see one image/caption pair at a time for a given item, where an item consists of two pairs: and its paired caption , and and its paired caption . They then compute an Image Score, Text Score, and Group Score (by scoring each item as either 1 or 0 and then aggregating). For a given pair, the Image Score is 1 if and only if for image a higher score is assigned to caption than and for a higher score is assigned to than . Similarly, the Text Score is 1 if and only if for text a higher score is assigned to than (and vice versa for ). Thus, for both the Image and Text Score, random chance is . An item’s Group Score is 1 if and only if both its Text and Image Scores are 1. The random chance for Group Score is .
Relaxing Winoground Constraints
These metrics are relatively harsh in two respects. First, they require perfect matching between the images and captions, implicitly evaluating an unusual variant of Recall: Recall @ 1 over candidates (i.e., the best image must be ranked first between the two candidates). Further, they do not allow any adaptation to the task (i.e., zero-shot transfer is required). We therefore relax each of these two constraints in turn.
We evaluate using a standard Recall at (R@) metric for retrieval, which asks whether the correct caption (for Image-to-Text, or I2T, retrieval) or image (for Text-to-Image, or T2I, retrieval) is present in the top candidates as ranked by the model. We consider R@1, R@2, R@5, and R@10 for CLIP, UNITER, and LXMERT (see Appendix A for further model details).
Crucially, R@1 requires discriminating both within semantic minimal pairs and between unrelated Winoground items, while all other metrics can be solved without having to differentiate the semantic minimal pairs (i.e., for R@2, the model can simply return both relevant items).
Each model is used to compute a similarity score for all possible pairs of any image from Winoground with any caption from Winoground. As in the original Winoground methodology, we do not finetune the models. In I2T retrieval, we score each image in turn and retrieve the top highest-ranked captions. In T2I retrieval, we score each caption in turn and retrieve the highest-ranked images. In either case, we then compute R@ as the percentage of image or caption prompts for which the correct match is among the top candidates.
Table 1 presents the results. CLIP performs well on the less harsh R@5 and R@10 metrics, while LXMERT performs poorly across all values of , with UNITER’s performance falling about halfway in between. Since neither UNITER nor CLIP clearly outperforms the other on the Winoground metrics (Thrush et al., 2022), the stark difference in overall R@ that we see between them here is surprising. One plausible explanation for this pattern is that LXMERT sees only about K unique images during pretraining (despite seeing between M and M captions), while UNITER sees about M and CLIP sees M. We hypothesize that CLIP’s larger training set size means that it can more easily adapt to unusual texts and images. Our results suggests that while the strict evaluation metric of Winoground leaves the three models at similar baseline performance, they clearly exhibit different levels of understanding Winoground captions in easier setting.
2 Task Adaptation
Thrush et al. (2022) evaluate models on Winoground zero-shot (with no fine-tuning to allow it to adapt to the task) and in such a way that the model is fed one caption and one image at a time (meaning, in choosing the best image match for , it does not get to simultaneously compare and in the way that a human does). To test whether performance is helped by addressing both factors, we train probes to select between two concatenated cross-modal embeddings as to which represents the better match for a given reference item. This amounts to a binary classification task, where the output is if the first embedding is a better match, or if the second is better.
We first divide the 400 Winoground items into 300 for training and 100 for testing. Stratified sampling is used to ensure that the original ratios of each Winoground tag (Pragmatic, Symbolic, etc.) are preserved in each subset. Our probes are 4-layer MLPs with a hidden dimension of 1024 trained for 200 epochs on the embeddings of the training items. We consider the Pooled Output embeddings produced by both UNITER and LXMERT, which are generated by applying a linear projection and Tanh activation to the hidden state of the CLS token at the last layer of each model; these are the embeddings used to predict similarity scores in the retrieval setting. Two variants of each probe are learned: one which picks between embeddings of the same caption with two different images (roughly corresponding to Text Score or I2T retrieval), and one which picks between embeddings of the same image with two different captions (roughly corresponding to Image Score or T2I retrieval). We report additional methodological details in Appendix C.
In addition to our target task of picking the correct match within each Winoground item, we train another set of probes which learn a control task. For our control task, we randomly pick of the training items and of the testing items and flip their labels, then train the probes the same way. All probes are trained and evaluated 11 times with different random seeds, and the min and max score across trials is recorded.
Probing results are reported in Table 2. None of the probes achieve an appreciably higher accuracy than either chance () or the control on the test set (although the UNITER text and image probe test accuracies trend somewhat higher than the UNITER control accuracies). This implies that the representations produced by LXMERT or UNITER may not contain the information required to succeed on Winoground, although it is possible that a different probe design or probing technique may be able to extract such information.
Characterizing the Challenges Presented by Winoground Items
The results of our more traditional evaluation suggest that the Winoground text/image pairs are, even without focusing on semantic minimal pairs, interestingly different from other visuolinguistic datasets. In this section, we seek to characterize what makes the Winoground task challenging. See Appendix B for details on our annotation method and Table 5 for tag to dataset item mappings. We introduce our taxonomy below, and present examples of each new tag in Figure 2.
While these items are textual minimal pairs, they are actually not semantically compositional variants of one another. This may be because the swapped words appear in a compound (e.g. “banana split” in WG #133, “downfall” in WG #325), because they are part of an idiom (e.g. “fishing for compliments” in WG #333), or because they are two different lexemes exhibiting polysemy. Items with this tag do not require compositional reasoning to resolve, since they don’t contain the same semantic entities.
2 Potentially Difficult Pairs: In-Domain
We identify two challenging categories of examples that are in-domain, but involve additional challenges beyond visual or linguistic understanding.
These items can be resolved when both images and both captions are considered together, but when considered separately, at least one of the captions is either a correct description of both images or not quite a correct description of either. SOTA Transformer-based VL models are trained to distinguish valid captions from invalid captions, but not to select the best caption from a set of valid candidates. Humans, while capable of making such fine-grained judgments, were queried differently than models in Thrush et al. (2022): rather than rating the quality of an image-caption pair along a continuum (analogous to models’ similarity scores), humans were asked for a binary judgment. Even a perfect respondent, if asked to evaluate some of these image/text pairs in isolation (without seeing the competitor pair), could receive zero Winoground scores since the correct answer is only discernible when both competitors are present.
For items given this tag, at least one element required to correctly sort the images is small, blurry, in the background, out-of-focus, indistinct, blends with the background, or otherwise difficult to detect. Since most VL models have low input image resolution, they may simply be unable to detect visual elements which are key to resolving these Winoground items.
3 Potentially Difficult Pairs: Out-of-Domain
We also identify three kinds of out-of-domain reasoning required to solve the Winoground task: either because the image is unusual, the text is unusual, or because they require extensive real-world knowledge or reasoning ability. While humans can adapt to out-of-domain tasks and it is desirable to build systems that can as well, this goes beyond mere compositionality.
Items that we tag UnusualImage have at least one image which is either entirely unrealistic or highly unusual and therefore likely out-of-distribution for most VL models. UnusualText captions may be difficult for models to resolve because they include a misspelled word (only found in WG #327); because non-standard capitalization is used in one of the captions (found in Winoground items); because they’re ungrammatical in Standard English (found in Winoground items); or, most commonly, because the wording of the caption is awkward. These may be descriptions a human would be highly unlikely to generate (e.g. WG #10, which captions an image of a boat “the water rests below the sail”) or phrases which are difficult to parse.
This category encompasses any item which requires common-sense reasoning or world knowledge to resolve. This may be numerical reasoning, as in WG #396 (which requires counting to 3 and 8 and identifying even and odd numbers); understanding of non-English languages, as in WG #298 (which requires the model to first perform OCR, then understand French text sufficiently to know “chaud” is hot and “froid” is cold); recognition of scientific terminology, as in WG #303 (which requires the model to know that a lizard is cold-blooded while a polar bear is warm-blooded); or causal inference regarding the ongoing events depicted, as in the example in Figure 2.
4 Results on New Tags
We compare Text, Image, and Group Score (as in Thrush et al. 2022) over the splits corresponding to each of our new tags, as well as on the 171 items which don’t receive any tag. Results for CLIP, LXMERT and UNITER are reported in the table in Figure 2, with scores beating random chance in bold.
As predicted, all of the potentially difficult tags are harder than the NonCompositional tag, in some cases strikingly so. For CLIP, performance on the 38 VisuallyDifficult tags is actually 0 for the Image and Group Score metrics, suggesting that for at least some items there may just not be sufficient visual information available for the model to make an accurate judgement. CLIP performs above random chance on all three metrics only for the NonCompositional tag, which tests the models’ response to highly similar texts without testing their compositional reasoning. CLIP also performs better on the AmbiguouslyCorrect tag than it does on the full dataset: it appears that CLIP is able to discriminate between multiple valid or multiple invalid captions for an image to some extent, even if distinguishing between multiple valid or multiple invalid images for a caption remains out of reach. CLIP’s much higher scores on the NonCompositional split compared to all other splits, including the NoTag split, implies that it is compositional reasoning in particular which makes Winoground so difficult, at least for the CLIP model evaluated here.
For LXMERT, the AmbiguouslyCorrect and UnusualText tags appear to be particularly challenging, and the VisuallyDifficult tag doesn’t appear to present much of a problem. However, it’s worth noting that all LXMERT scores are below random chance–we therefore cannot be certain that any particular score difference is not a coincidence. LXMERT’s failure to perform coarse-grained retrieval over the full Winoground dataset makes it unsuprising that it cannot correctly match even the potentially easy NonCompositional tag.
UNITER is able to beat random chance on Text Score in all cases except for the UnusualImage and UnusualText, suggesting that out-of-domain samples are a particularly salient challenge for UNITER. In terms of Image and Group Score, UNITER is only able to beat random chance on the NonCompositional tag. This again implies that it is not just textual minimal pairs that cause catastrophic failure, but specifically textual and semantic minimal pairs.
Generating Non-Minimal Winoground Data with Textual Variants
In this section, we look in depth at whether the minimal textual pairs are simply not sufficiently distingushable with existing vision and language models. That is, at the level of text, does the model not understand that “grass in the mug” is distinguishable from “mug in the grass”? Or is the problem instead that the images are not distinguishable—or that the fusion of the visual and linguistic information is too difficult?
To tease apart these hypotheses, we run experiments using caption variants: we modify each caption in each Winoground item so that the captions are no longer minimally contrastive. We obtain caption variants by using manually selected augmentation strategies from NLAugmenter (Dhole et al., 2021) and categorize them by the type of modification they make (see Table 3 for an example). For a given Winoground item , the caption variants are denoted by and . For more details about these augmentation strategies, refer to Appendix D.
We first investigate the separability of textual variants of from textual variants of in model embedding space for the three models (LXMERT, UNITER, CLIP) in Section 5.1. Then, we test whether providing models access to textual variants helps performance on the Winoground task in Section 5.2. Finally, we analyze the ability of models to distinguish the right caption conditioned on its textual variant in Section 5.3.
Our core question in this experiment is whether textual variants of and textual variants of are effectively partitioned in each model’s embedding space. If semantic differences aren’t captured by the language branch, then no matter how well fine-grained semantics are extracted from images and no matter how well text semantics and image semantics are aligned, these models cannot be expected to succeed on Winoground. On the other hand, if there’s a clear linear division between the caption groups, then the semantic distinctions between and are already easily retrievable from a model’s text branch, and the model’s overall failure cannot be resolved by improvements to its ability to discriminate text.
For each Winoground item, we construct four sets of CLS @ embeddings (the embedding for the [CLS] token at layer ): variants of caption 0 conditioned on image 0 (), variants of caption 0 conditioned on image 1 (), variants of caption 1 conditioned on image 0 (), and variants of caption 1 conditioned on image 1 (). We fix the image input, and compare the target task of distinguishing variants of caption 0 from variants of caption 1 with a control task where variants of both captions are randomly assigned to one of two arbirary sets.
Separately, for each of the 400 Winoground items with textual variants, we use a Linear Support Vector Classifier probe to measure separability, with hyperparameter to prioritize complete separation over margin width. We obtain two key measures: the binary variable of whether the sets are linearly separable (true if and only if every variant is correctly labeled by the learned probe), and the width of the discovered margin (computable by ). We train one SVC over CLS @ embeddings for each combination of task, layer, Winoground item, model, and (for LXMERT and UNITER) which image is input alongside the text, then average across items and images to analyze high-level trends.
The first row of Figure 3 shows the results. For LXMERT, we find that embeddings only become separable with a margin size of at least at layer 6 and remain separable for the rest of the layers. The introduction of cross-modal attention at layer 9 is followed by a slowing of margin growth and an eventual decrease. Control sets are much less likely than target sets to be linearly separable. For UNITER, which is always cross-modal, the margin increases steadily across layers until dropping sharply for the last two. The control task peak for UNITER is similar to that of LXMERT, but the target task peak is lower, suggesting that LXMERT’s representations of fine-grained semantic distinctions are slightly more linearly separable than UNITER’s. For CLIP, we find that neither target-task nor control-task captions can be linearly separated in the embedding space after the first layer. We investigate the possibility of non-linear representations in 5.3.
Across all LXMERT layers, target task probes find a linear decision boundary which perfectly separates variants of one caption from variants of the other caption of the time, while control task probes able to find a perfect decision boundary only of the time. Among perfect decision boundaries discovered for the target task, the average margin width is , while among perfect decision boundaries discovered for the control task, the average margin width is only .
Across UNITER layers, target task probes find a perfect decision boundary 84.7% of the time with an average margin width of 1.018 while control task probes only do so 19.0% of the time with an average margin width of 0.48. CLIP target task probes only find a perfect linear decision boundary 3.5% of the time with an average margin width of only 0.13, while CLIP control task probes never succeed in perfectly separating variant embeddings.
This gap suggests that LXMERT’s and UNITER’s language layers are in principle able to learn easily-extractable representations of Winoground captions, which may capture the differences between semantic minimal pairs. Given this result, we ask (in Section 5.2) whether using these separable caption sets (as opposed to the original textual minimal pairs) could be used to improve Winoground performance. If so, it would suggest that state-of-the-art VL models stumble only on lexically overlapping captions; if not, it would suggest that these models struggle with fine-grained semantic distinctions even in the absence of significant lexical overlap.
2 Do Caption Variants Help with the Winoground Task?
To assess whether the separable captions help on the main task, we develop a new test of the Winoground task, using our augmented captions. Using different variant-generation methods as defined in Section 5, we can obtain sets of captions and . We then obtain augmentation-aware similarity scores between an image and a set of caption variants of , as whereas the similarity score interpolates, with a hyper parameter , between the original similarity score and an aggregate score across all variants; we experiment with both the max and the mean as possible aggregation functions. We use these scores to see if performance is improved on Winoground by using caption variants. We conduct a hyperparameter search over the aggregation function and the value for , as described in Appendix D.1.
Results are presented in Table 4. Augmentation does not improve the Text or Group Score by much, indicating that these interventions to increase textual discriminability do not make it easier for the model to pick the correct text given the image. This suggests the high lexical similarity between the caption pairs is unlikely to be the main challenge, since models fail to pick between semantically-similar, lexically-different caption candidates.
3 Distinguishing Captions Conditioned on Caption Variants
Finally, we evaluate whether the partitions of the embedding space found in Section 5.1 are meaningful by training MLP probes to select between two captions conditioned on a different variant of one of the captions (all paired with the same reference image). These probes, unlike the SVC ones, have the ability to identify and use non-linear patterns in the embeddings. Intuitively, we are asking: can a probe over the text embeddings produced by each model correctly identify that the caption “a human viewing a cat on a screen” is correctly paired with the paraphrase “a person who looks at a cat on a screen” (which is a semantic match) and not with a variant of its semantic minimal pair (e.g., “a cat who looks at a person on a screen”)? In this experiment we do not train a separate SVC for each Winoground item, but use a single MLP across all Winoground items, increasing the task difficulty significantly. If the partitions found for each Winoground item are arbitrary, then a probe trained to distinguish between caption variants should fail on any Winoground item not seen during training. On the other hand, if a probe is able to distinguish between caption variants for unseen Winoground items, then it must have learned a semantically meaningful partition of the embedding space. We use the same train/test splits as in Section 3.2, and a similar control task, in which the labels of a fixed random of Winoground items are swapped.
Our results are depicted in the second row of Figure 3. Performance for LXMERT and UNITER falls between the catastrophic failure of the cross-modal probes in Section 3.2 and the clear success of the unimodal probes in Section 5.1. Performance on the test set is never higher than for any probe size or embedding layer. However, target task probes clearly outperform control task probes on test set accuracy, as shown in Figure 3. On the other hand, MLP probes over CLIP’s text branch are more successful than the linear SVC probes over CLIP from Section 5.1, beating chance by about 10% accuracy on the test set for the target task. This suggests that CLIP may in fact be encoding some semantic distinctions, but that the representations produced by CLIP layers are non-linear.
Test set accuracy clearly improves with layer depth for all three models in early layers, but begins decreasing when cross-modal attention is introduced at layer 9 in LXMERT, and for the final two layers of UNITER. This finding mirrors our results from Section 5.1. The findings for CLIP differ from Section 5.2, with performance peaking at layer 3 and remaining similar across all subsequent layers. Performance above chance on this task constitutes some evidence for our hypothesis that text processing is not the primary cause of failure on Winoground for the best current VL models.
Conclusion
We initially asked whether failures on Winoground occur because SOTA models rely more on bag-of-words than they let on and cannot tell the difference between sentences that contain the same words but differ in meaning. We found that the story is more complicated: high lexical overlap between captions is not the only—or even the most likely—cause of failure.
First, we showed that it’s not only the textual difference between “a mug in some grass” and “a grass in some mug” that makes Winoground hard. Indeed, using Recall@k, we showed that models struggle to identify that either minimally different caption matches a particular image.
Next, we re-categorized the Winoground dataset using a set of tags that identify significant challenges beyond semantic compositionally. For instance, we identified 38/400 items as VisuallyDifficult, meaning they require identifying a subtle visual feature of the image such as the eye color of a person in an image. Performance is very low on this subset, for reasons that may have nothing to do with language. Moreover, some of the images (56/400) and captions (50/400) are unusual or hard to parse: these images and captions are challenging for reasons having nothing to do with their inclusion in a minimal pair.
Even ignoring these cases, we still found that performance on the 171 vanilla Winoground items was low. To determine whether this is due to the particular zero-shot evaluation setting used by Thrush et al. (2022), we trained small probes to distinguish between LXMERT or UNITER embeddings of correct matches and incorrect matches. These probes’ performance was not consistently better than those trained on a parallel control task, suggesting that zero-shot evaluation is not the source of model failure.
So are these examples hard because models do not understand word order? We ran a set of experiments in which we made the textual minimal pairs more different from each other: by augmenting the Winoground dataset with variants of each caption, we produced sets of captions which were semantic but not lexical minimal pairs. Probing the embeddings of these variants, we found that semantic distinctions were linearly separable from LXMERT and UNITER layer representations and non-linearly separable to some extent by LXMERT, UNITER and CLIP representations. Even still, all three models fail to match each set of caption variants with the correct image. Thus, we observe robust failure on the task even when we use caption variants known to be distinguishable. It seems that the problem is not simply that the model cannot distinguish between captions with overlapping text, but likely lies in associating those distinctions with images.
Overall, Winoground remains a challenging and promising way to test visuolinguistic ability. We would encourage future work to report results on each of the tags we introduce separately, given the clear performance differences across tags we found for CLIP, UNITER, and LXMERT. And we urge care in drawing conclusions about the compositional abilities of vision-and-language models.
Limitations
Like the original Winoground dataset, we evaluate only English. Because English is highly word-order dependent, less word-order dependent languages may behave very differently, and, in fact, constructing a Winoground-like dataset in such a language would be non-trivial. Thus, we should not assume these results generalize to all languages.
We test only 3 types of multimodal models. While we chose our models to be representative and amenable to the kinds of experiments we were running, we cannot guarantee that our findings apply to all multimodal models.
Also, we focus here mainly on the separability of embeddings in text space. There are a parallel set of experiments that could be done for the visual space, but we did not conduct such experiments here. Therefore, our conclusions should be limited to what can be concluded from text augmentations.
Finally, we draw some conclusions based on failures to improve models. While we believe these negative results are informative, it is of course possible that a better method could be used that would give different results and so one should remain open to this possibility.
Acknowledgements
We gratefully thank the Winoground authors for sharing data and helpful conversations, particularly Candace Ross and Adina Williams. We thank the students in the UT Austin LIN 393 “What do neural networks know about linguistic structure?” seminar for input and comments in the early stages of this project. We also thank Gauri Kambhatla, Vanya Cohen, Jierui Li, and Ray Mooney for helpful discussions. This work was supported by National Science Foundation Grants No. 2104995 to KM.
References
Appendix A Details on Evaluation Methods
Specifically, we use UNITER-base pretrained on COCO Captions, Visual Genome, Conceptual Captions, and SBU Captions as described in Chen et al. (2019); LXMERT-base pretrained on COCO Captions, Visual Genome, VQA v2.0, GQA balanced, and VG-QA as descrbed in Tan and Bansal (2019); and CLIP with a ViT-B/32 image encoder pretrained on WebImageText as described in Radford et al. (2021).
Appendix B Method for tagging
The development of these tags occurred in four stages, or “passes” through the dataset. In our first pass, we looked briefly at all items to get a broad sense of the dataset, reflecting on and discussing with colleagues any items which caught our interest. The second pass through the dataset looked at each Winoground item carefully one at a time, taking notes on the type of swap performed and any particular challenges or interesting features present in that item. From the pages of notes produced during the second pass, our final set of tags was selected to encapsulate broad patterns found throughout the dataset (this number was not fixed in advance, but determined by the number of unique patterns identified). A third pass over the dataset was performed to assign these tags to each image. Finally, since the second and third passes were initially performed by one annotator to ensure consistency, a fourth pass was performed by the other authors to verify the tagging, and Winoground items for which any two annotators disagreed about the tagging were carefully examined and discussed by all annotators to reach a consensus.
After retagging the dataset, the frequency of each new tag was computed, and tag-level performance was measured.
Appendix C Task Adaptation methods
We consider the following embeddings as potential inputs to our probes:
Pooled Outputs: The hidden state for the CLS token at the last layer of the model, further processed by a linear projection followed by a Tanh activation.
CLS @ : The hidden state for the CLS token at layer
Mean @ : The vector produced by mean-pooling across all hidden states at layer
Max @ : The vector produced by max-pooling across all hidden states at layer
Let the function be application of a given model to input image and caption , followed by extraction of the embedding being probed over. Our probes are then given a pair of concatenated embeddings for two candidate image-caption pairs, and asked to output if the first pair is a better match or if the second pair is a better match. Specifically, we define the following task corresponding to Text Score for a probe :
An equivalent probing task corresponding to Image Score is defined:
Our control task is formulated nearly identically to the target task, except that a random 50% of the Winoground items in each split are chosen at the start of training, and the labels for these items are flipped. That is, if our target task has the label and the given item is not flipped, the control task also has the label ; if our target task has the label and the given item is flipped, the control task has the label .
Since the flipped items are selected before training, a probe which simply memorizes the data will perform well on this control task for items it has already seen. We therefore split the data into a training and a testing set, the latter of which is never seen during training. This means we are dividing the Winoground items into four new splits: training samples whose labels are not flipped in the control task, training samples whose labels are flipped, testing samples whose labels are not flipped, and testing samples whose labels are. In order to ensure each split is representative of overall the Winoground benchmark, we perform stratified sampling, where are buckets are any combination of the “Pragmatics”, “Symbolic”, and “Morpheme-Level” visual tags and the “Both” linguistic tag. We subsequently confirm that the ratio of each visual and linguistic tag, as well as of each new tag introduced here, is similar across each split.
We test a variety of small Multi-Layer Perceptron (MLP) probes over these extracted embeddings, each of which maps from the input dimension of to a single output. ReLU activation is applied at intermediate layers, and Sigmoid activation is used at the final layer to ensure the output is in the range $10240.00014200$ epochs of training to use for every probe, after ablating each of these hyperparameters for all combinations of probe and embedding types.
We use a single NVIDIA RTX 8000 GPU for all our experiments. All probes took no more than hours to run.
Appendix D Textual Variants Methods
We generate a maximum of variants using the PropbankSRLRoles augmentation. This augmentation extracts semantic role labels for the provided sentence using the AllenNLP implementation https://demo.allennlp.org/semantic-role-labeling of SRL BERT (Shi and Lin, 2019) and applies its hardcoded syntactic rules (if applicable) to generate a new sentence.
D.0.2 Semantic, word-based augmentations
We use different augmentation methods, each of which randomly replace words in the sentence with new approximately meaning-preserving words. All methods use SpaCy (Honnibal et al., 2020) to parse the sentence to perform POS tagging.
ReplaceHyponyms, ReplaceHypernyms. The first augmentation replaces a noun with a hyponym and the second replaces a noun with a hypernym. We generate a maximum of variants per augmentation. This method uses CheckList (Ribeiro et al., 2020) for the list of hyponyms/hypernyms.
Slangificator. It replaces a word with a slang word. This uses a manually curated list of word -> slang word mappings. We generate a maximum of variants.
SynonymSubstitution. It replaces a word with a synonym based on WordNet (Miller, 1998) via NLTK (Bird, 2006). We generate a maximum of variants.
D.0.3 Paraphrasing
Backtranslation. It translates a sentence to German and back using FSMT (Ng et al., 2019). We generate a maximum of variant.
DiverseParaphrase. It generates diverse paraphrases using DiPS (Kumar et al., 2019) equipped with Diverse Beam Search (Vijayakumar et al., 2018). We generate a maximum of variants.
ProtAugmentDiverseParaphrase. It generates diverse paraphrases using ProtAugment (Dopierre et al., 2021). We generate a maximum of variants.
D.0.4 Identity
We also define the original input text as a ‘variant’ that has undergone the identity transformation.
D.1 Discriminable Caption Pair Experiment
To assess whether the separable captions help on the main task, we develop a new test of the Winoground task, using our augmented captions. Using different variant-generation methods as defined in Section 5, we can obtain sets of captions and . Then, every multimodal model under consideration outputs a similarity score given an image and text as input. We define augmentation-aware similarity scores between a given image and a set of caption variants of , as follows:
where the choice of using vs. and the value of are hyperparameters. This similarity score interpolates, using , between the original similarity score and an aggregated (max/mean) score across all variants. We can use these scores to see if performance is improved on Winoground by using caption variants.
We conduct a hyperparameter search over the similarity function and the value for . We found that works best for LXMERT while works best for UNITER and CLIP. We picked the best value for by testing every value between and in steps of and picking the value that maximizes the group score. works best for LXMERT, for UNITER and CLIP.
Appendix E Winoground: New Tags
Our new tags appear in Table 5. For the full Winoground dataset, see https://huggingface.co/datasets/facebook/winoground.