Equivariant Similarity for Vision-Language Foundation Models
Tan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, Lijuan Wang
Introduction
Vision-language (VL) training is all about learning “good” features for each modality, such that the features should faithfully represent the underlying semantics. Thanks to the large-scale image-text pairs on the Web, we have abundant multimodal supervision for the two features with the same semantic meaning —each matched image-text pair should have “similar” visual and textual features, and each unmatched pair should have “dissimilar” ones. Thus, the image-text similarity plays a crucial role to define the feature quality in training VL foundation models (VLMs) .
Has the prevailing “matched vs. unmatched” similarity fulfilled its duty? Yes and no. On the one hand, recent VLMs have demonstrated impressive results in various downstream VL tasks such as image-text retrieval. However, on the other hand, it is acknowledged by the community that the VLMs still fall short in nuanced and complex semantic compositions . In this regard, we present a text-to-image retrieval example on LAION400M with the most recent SOTA VLM FIBER . As shown in Figure 1(a), given the query text “the house on the right side of the road”, we first invite 5 graduate students to rank 25 candidate images from most similar to least similar. The continuously decreasing ranking from human judges () is served as the oracle semantic similarity measure. We then compared this ranking with the ones from FIBER (). Although FIBER correctly retrieved the top- image (image#1, ranks 1), some semantically incorrect images (e.g., image#25, ranks 17) are falsely ranked higher than the correct ones (e.g., image#2, ranks 20). Furthermore, when modifying the query text with a slight semantic change (“right” “left”), the rankings remain almost the same. Clearly, the similarity changes in FIBER do not faithfully reflect the semantic changes in images (#1 #25) or text queries (“right” “left”).
To quantitatively measure the above inconsistency between semantic and similarity score changes, we consider two matched image-text pairs and that are semantically similar but only different in the number of clocks in Figure 1 (b). With a slight change of clock counts in caption (“2”“3”), FIBER mistakenly assigns a higher similarity score to rather than ( v.s. ). Furthermore, the changes in similarity scores guided by the semantic change (“2”“3”) are highly inconsistent ( v.s. ). Ideally, an equivariant image-text similarity measure should faithfully reflect the semantic change, i.e., the same semantic changes should lead to a similar amount of similarity changes (e.g., v.s. of ours in Figure 1(b)).
Equivariance Loss. To address this non-equivariance issue, we propose Equivariant Similarity Learning (EqSim), which imposes additional equivariance regularization on image-text pairs for VLM learning without additional supervision. Figure 2 illustrates the underlying semantics perceived by human, where each matched pair demonstrates the image and text corresponding to the underlying semantic. Given two matched image-text pairs as semantic 1 and as semantic 2, we can obtain four similarity scores , , , and . We define Equivariant Similarity to be an image-text similarity function, whose output value should correspond to the underlying semantic change, which can be measured by text or image change.
(Equivariant Similarity) The similarity between image and text is equivariant if and only if the following equations hold:
where () denotes the measure in image (text) space, i.e., an infinitesimal unit of visual (textual) change. Based on Definition 1, we formally derive EqSim, an equivariance loss for a hybrid learning strategy on both semantically close and distant training pairs (Section 3). Specifically, EqSim directly enforces and for semantically close samples; while for semantically distant samples, we derive a simplified formulation of . We show that adding EqSim as a regularization term improves existing similarity training objectives significantly on challenging datasets (e.g., over on Winoground ) and tricky tasks (e.g., around on VALSE ). EqSim can also retain or even improve retrieval performance on Flickr30K dataset.
Equivariance Benchmark. To further facilitate the proper evaluation of equivariance in VL community, we present a novel evaluation benchmark dubbed EqBen (Section 4). Motivated by the examples in Figure 1(b), EqBen features “slightly” mis-matched pairs with a minimal semantic drift from the matched pairs, as opposed to “very different” matched and unmatched pairs that are easily distinguishable by both non-equivariant and equivariant similarities. Unlike recent efforts focusing on minimal semantic changes in captions, EqBen pivots on diverse visual-minimal changes, automatically curated from time-varying visual contents in natural videos and synthetic engines with more precise control. We benchmark a full spectrum of VLMs on EqBen, and reveal that the non-equivariant similarity in existing VLMs fails easily. On this new test bed, EqSim can serve as a remedy and bring a large performance gain of 3% on average.
Our contributions are summarized as follows: (1) We comprehensively study the problem of similarity equivariance in VLMs. We propose EqSim for equivariant training and EqBen for diagnostic evaluation; (2) EqSim is not only theoretically grounded but also simple, effective and easily pluggable; and (3) EqBen clearly diagnoses that conventional evaluation is not responsive to equivariance. Furthermore, EqSim can significantly improve VLMs on EqBen, as well as other challenging benchmarks.
Related Work
Pre-training VL Models. Early object detector (OD)-based methods utilized the offline image region features from a pre-trained object detector . More recent methods mainly learn from image pixels directly in an end-to-end manner . Researchers further categorize VLMs into (1) Dual-Encoder (e.g., CLIP and ALIGN ) and (2) Fusion-Encoder (e.g., METER , FIBER , and ALBEF ). It is worth noting that our proposed EqSim is model-agnostic, and can be easily plugged into the image-text alignment objectives such as Image-Text Matching (ITM) and Image-Text Contrastive (ITC) loss.
Diagnosing VL Models. Years of VL research have spawned a series of VL evaluation kits, from classical VL tasks (e.g., VQA and image captioning ), to more complex contexts, such as adversarial examples , robustness and counterfactual reasoning . However, these benchmarks require manual annotation and their evaluation relies on task-specific model fine-tuning. Another line of work probe VLMs on similarity measure with minimal caption semantic changes while keeping images intact. While our EqBen tries to test whether the inherent image-text similarity measure in existing VLMs is sensitive to visual semantic changes. The most relevant work is ImageCoDe which leverages video frames toward fine-grained image-text retrieval. However, ImageCoDe requires additional human crowdsourcing and is limited to real-world video sources. In contrast, EqBen explores both natural and synthetic ways to generate image pairs with minimal semantic change, making the data generation process inclusive, automatic, and extensive.
Equivariance Learning. Unlike the wide usage of invariance in deep neural networks (e.g., shift invariance achieved by convolutional layers), strict group equivariance is hard to apply in practice. However, the equivariance property still plays an important role in various fields, such as self-supervised learning , representation learning , and language understanding . In this paper, we point out the significance of the equivariant similarity measure in VLMs. Based on this, we further propose a novel loss EqSim for the regularization of equivariance, as well as a new challenging benchmark EqBen to diagnose the equivariance of existing VLMs. We notice that the recent CyCLIP delivers a similar idea but with different motivation, implementation and evaluation settings. In Table 6, we compare with CyCLIP-equivalent baseline as EqSim. Check more detailed comparison in Appendix.
Improving VLMs with EqSim
Recall that VLMs adopt the image-text similarity as the training proxy for learning multimodal feature representation . Therefore, the ultimate goal of pursuing equivariant similarity is to learn an equivariant feature map between image space and text space.
Let and be two continuous feature spaces. Let be a group whose group action on is defined by , and that on is defined by . Then, is an equivariant feature map if and only if for all the group actions and . The commutativity for , and is shown in Figure 3.
In Definition 1, the measure can be considered as the semantic group acting on an infinitesimal region in image or text space. Thus, by applying the commutativity of Definition 2 in Definition 1 to change the sum from image space to text space. Without loss of generality, we only show the results of sum :
This implies that the equivariant map establishes an isometry for the measure in image space and in text space. Thus, they only differ by a constant scale , i.e., :
By combining Eq. (1), (2), (3), and (4), we have the following ratio equality as our EqSim constraint:
Note that can be derived by using the fact that and . By simplifying Eq. (5) further, we have the following two regularizations:
Note that the viable space of EqSim is a subset of EqSim, because EqSim is exactly equivalent to Eq. (5) while EqSim further requires . Empirically, we find that EqSim is more suitable to the semantically close pairs and ; and EqSim to distant pairs. Figure 4 illustrates such hybrid training loss within a training batch. Semantically “close” and “distant” are determined by the similarity score , where we regard samples with top- as “close” samples. For dual encoder VLMs with ITC loss, is the cosine similarity between image and text features. For fusion encoder VLMs with ITM, is the scoring output from the ITM head.
In our implementation, we adopt Mean Square Error (MSE) loss to regularize the equation of similarities. In addition, motivated by the hinge loss , we utilize a margin parameter to control the strength of regularization. EqSim can be written as , where and denotes the L2 norm. EqSim can be implemented similarly. In practice, given a retrieval fine-tuning objective , the final loss can be written as: , where is EqSim (EqSim) for semantically distant (close) samples, is the balancing factor. Experiments in Section 5.4 validate that the hybrid training is better than only using EqSim.
Diagnosing VLMs with EqBen
We argue that standard VL evaluation kits are too coarse to evaluate the equivariance of VLMs similarity. Existing VLMs can easily distinguish most samples in conventional retrieval benchmarks, e.g., images of a group of people against those with cars, given a caption of “people standing on the street”. Therefore, we propose EqBen to focus on visual minimal semantic changes to check whether VLMs can faithfully respond, i.e., the equivariance of the similarity measure in VLMs. Specifically, EqBen contains 5 sub-datasets, covering diverse image domains, from real-life scenarios to synthetic well-controlled scenes. And it is designed to stress test VLMs with accurate semantic changes in action, location, and attribution (e.g., color, count and size). Figure 5 presents an overview of EqBen.
Next, we introduce our design principle for constructing EqBen. Each sample in EqBen consists of a pair of images and a pair of captions . A valid EqBen sample must satisfy: (1) is preferred to be used as the description for ; (2) and are visual-minimally different. The former one requires that and should be semantically distinguishable without confusion, while the latter one limits the extent of the distinction – “visual-minimal change”. Previous work defines “minimal” semantic change in the caption space as the same words but in a different order. However, due to the continuity and entanglement of image pixels, “minimal” semantic change in visual space is hard to determine. In this paper, we roughly define it as changes in the foreground (e.g., attribute, action, location, etc.) while sharing the same scene and background.
In practice, we source image pairs with “visual-minimal change” in two ways: (1) from natural videos and (2) from synthetic engines, where we adopt different construction pipelines, as shown at the bottom of Figure 5. For the former one, we directly leverage the continuity of scene changes along the temporal dimension in natural videos, which can provide massive image pairs with minimal visual changes. Specifically, we leverage the existing video-language datasets to construct EqBen samples. To more precisely control the varying component in images, we further explore the photo-realistic scene generator (Kubric ) and the open-source diffusion model (Stable Diffusion ) to synthetically generate pairs of images by providing two captions that are minimally different from each other. In what follows, we introduce the construction pipeline for each sub-dataset in detail.
Let’s define a video with caption annotations as , where is the number of sampled frames. We assume the “visual-minimal change” is naturally guaranteed between any two frames and ideally, where , , as we limit the source video to be either short segment or capturing a fixed scene . However, we find that it is hard to ensure the validity of all the video frame pairs in practice. Therefore, we utilize a frame filter to filter out invalid samples automatically for different video sources. Below, we briefly introduce the dataset construction process and delay the details to Appendix.
We construct three sub-datasets based on real images from natural videos, including Eq-AG, Eq-GEBC and Eq-YouCook2. We construct Eq-AG by leveraging the scene graph annotations from Action Genome (AG) , which capture detailed changes between objects and their pairwise relationships while action occurs. We first use a slot-filling template to translate scene graphs to captions. As videos usually come with redundant frames, we avoid the nearly duplicated frames by sampling frames and if and only if at least 2 of 3 pairwise relationships are different. Furthermore, if is chosen for previous samples, we empirically skip the subsequent 2 frames to . Eq-GEBC is built on GEBC that contains captions describing the event before and after an event boundary. We adopt the frames before and after the boundary as our visual minimally different images. Similarly, we avoid temporal redundancy via sparse sampling across multiple boundaries. We construct Eq-YouCook2 based on YouCook2 , which is sparsely annotated with captions for each cooking step. We construct the dataset by sampling the middle frame of a short video segment as with its annotated caption as . We then apply off-the-shelf object detectors to filter scene changes. Please note that EqBen can be easily extended to other video-language datasets by applying the same construction pipeline.
2 Construction from Synthetic Engine
Synthetic engine may provide more precise and controllable visual changes in the generated images, to allow more accurate diagnosis in terms of model failure when evaluating with EqBen. We assume that a synthetic engine can faithfully generate images based on a text prompt describing the image content. Based on this assumption, given a pair of semantic-minimally different captions and , we expect the generated and to be correspondingly visual-minimally different. In the following, we briefly introduce the utilized engine and how to construct semantic-minimally different captions for each sub-dataset and leave more details to Appendix.
Eq-Kubric takes advantages of Kubric , an open-source graphics engine to generate photo-realistic scenes. Here we adopt Google Scanned Objects (GSO) for scene construction and categorize caption change into three aspects: attribute, counting, and location. For each aspect, we construct image-text pairs by intervening corresponding phrases of sentences while leaving other words unchanged. Eq-SD is inspired by the recent advances in diffusion models for text-to-image generation . We utilize the open-source checkpoint v1.4 of Stable Diffusion (SD) with prompt-to-prompt image editing framework to translate two semantic-minimally different captions to a pair of images. Specifically, we elaborately design a set of textual semantic-minimal editing: 1) object change (e.g., “dog”“cat”); 2) scene change (e.g., + “in the winter”); 3) attribute change (e.g., + “with a sunglasses”). Finally, we perform a human evaluation to filter out poor-quality generations. Notably, we can adopt more rendered objects (e.g., rendered animals) and various synthetic engine (e.g., better generative models) to further extend our EqBen following the proposed pipeline.
3 Comparisons with Other Datasets
In Table 1, we conduct a direct comparison of EqBen against two widely adopted retrieval benchmarks (Flickr30K and COCO ) and two recent datasets with textual-minimal change (VALSE and Winoground ) from four aspects. 1) On the dataset characteristics, to the best of our knowledge, EqBen is the first diagnosing benchmark to examine the equivariance of VLMs in terms of minimal visual semantic change. 2) For evaluation setting, pairwise setting asks VLMs to select the correct counterpart within a pair of slightly different samples rather than thousands of very different samples in conventional retrieval datasets. The minimal semantic drift between the pair of samples makes the evaluation of equivariant similarity measure more effective. 3) For domain diversity, our EqBen contains rich visual contents collected from different video domains as well as synthetic domains, as opposed to the common image-text datasets which existing diagnosing kits are built upon. 4) In terms of scalability, EqBen is highly scalable as our automatic pipeline can be easily applied to other video-language datasets and synthetic engines, as opposed to manual annotation for building traditional retrieval datasets. While VALSE only focuses on linguistic editing of the captions, EqBen can be further scale up with more diverse visual contents.
Experiments
We first introduce our experimental setting in Section 5.1, followed by evaluation of EqSim on existing benchmarks in Section 5.2. Section 5.3 benchmarks SOTA VLMs on EqBen to show their insensitivity to minimal visual semantic changes, and we further validate EqSim on EqBen. Section 5.4 presents additional ablation studies to examine the design of EqSim.
Training Details. Recent efforts on diagnosing benchmarks only provide testing data and directly evaluate models after VL pre-training. The low performance reported on these benchmarks can mainly be attributed to two factors: 1) the inherent weaknesses of VLMs, e.g., non-equivariant similarity measure; and 2) the domain gap between training and testing. To better validate the effectiveness of our method, we fine-tune the VLMs on limited image-text pairs from conventional retrieval dataset Flickr30K with or without the regularization term of EqSim, and then test the fine-tuned VLMs on the challenging Winoground , VALSE and our EqBen. Under a fair comparison, we argue that the absolute performance improvements from EqSim thus would suggest that the gain is entirely from the remedy of model weaknesses.
To validate the effectiveness and genraliazability of our proposed method, we apply EqSim to two SOTA end-to-end methods with different architectures and retrieval losses. Specifically, FIBER supports the dual encoder with ITC loss for fast retrieval, which computes similarities for image-text pairs with only forwarding. In contrast, the SOTA fusion-encoder model METER , optimized with ITM task during pre-training, computes the similarity by forwarding the concatenation of each pair of image and text, resulting in time complexity. Fine-tuning details for each model can be found in Appendix.
Evaluation Metric. On Winoground , given two image-text pairs and , a VLM measures similarity between image and text . Three metrics are computed based on : 1) Text score measures whether the model can select the correct text for a given image. The model wins one point if . 2) Image score evaluates if VLMs can select the correct image for a given text and the model wins one point when . 3) Group score combines the previous two, such that the VLMs win one point if and only if both text score and image score are 1, meaning the following condition must be satisfied: and . On VALSE , the two image-text pairs share a common image, i.e., (correct) and (foil). We follow to report the following metrics: 1) acc is the overall accuracy on both correct and foil image-text pairs; and 2) min() is the minimum of precision and foil precision , where () measures how well models identify the correct (foil) pair. We also report performance on the conventional image-text retrieval task, where recall R@K (K=1,5,10) is used as the evaluation metric.
2 Evaluation of EqSim
In Table 2, we compare model performance under three settings: () direct evaluation after pre-training (the first rows of each block); () standard fine-tuning (FT) on Flickr30K training data (the second rows of each block); and () fine-tuning with EqSim regularization (the third rows of each block). We observe that standard fine-tuning can somewhat bring a little performance improvement on both Winoground and VALSE benchmarks, indicating that some domain overlap between Flickr30K training data and testing samples. It is difficult to entirely rule out the domain influence, but comparing fine-tuning with EqSim against standard fine-tuning, our method brings consistent and significant performance improvements on both of the challenging Winoground and VALSE across METER and FIBER models. Specifically, EqSim improves the group score over standard fine-tuning by for METER and for FIBER on Winoground, respectively. While for VALSE, the performance improvement on min is as large as , further validating the effectiveness of our EqSim. In addition, we observe that the equivariance regularization from EqSim does not sacrifice retrieval performance. On Flickr30K, EqSim can mostly retain the retrieval performance, and sometimes even yield performance gain, e.g., on R@1 for image-to-text retrieval with FIBER.
3 Benchmarking VLMs with EqBen
We evaluate a wide range of VLMs with different configurations on EqBen in a zero-shot manner, to examine the equivariance of their similarity measures for distinguishing visually-minimal different samples. We consider representative VLMs, including () LXMERT , ViLBERT for OD-Based models; and () CLIP variants, FLAVA , ViLT , ALBEF , BLIP and METER and FIBER as prominent examples of end-to-end SOTA methods. Full results on more VLMs can be found in Appendix A.5. For evaluation metrics, we adopt text score, image score and group score to compare model performance, similar to Winoground .
Table 3 presents the evaluation results of existing VLMs on EqBen and we summarize our observations below.
Regardless of the subsets, end-to-end VLMs generally achieve better performance as it is not constrained by the fixed visual representation from a pre-trained object detector , as in OD-based methods.
Among all subsets, VLMs obtain evidently higher performance on Eq-SD. The stable diffusion model is pre-trained on similar VL corpus to these VLMs. Hence, the generated images can be biased towards the same underlying data distribution, much easier for VLMs to tell the differences. Besides, the generated images maybe visually minimally different to human eyes, but it is unclear whether in the pixel space, they are minimally different w.r.t. the model input. It is worth noting that LXMERT and ViLBERT are the exception due to the totally different distribution with the off-the-shelf object detector.
Interestingly, a larger pre-training corpus (e.g., CLIP and FLAVA ) does not always guarantee better results. This implies training loss may be more critical in learning equivariant similarity measure.
In Table 4, we further conduct a fine-grained examination with the synthetic subset EQ-Kuric, where we focus on specific visual changes in location, counting and attribute. VLMs fail substantially in terms of location and counting, while being sensitive to attribute changes. Similar findings are also observed by from the text side.
We again equip the two strong baseline models (METER and FIBER) with EqSim and fine-tune on Flickr30K. As EqBen covers diverse domains, standard fine-tuning on Flickr30K can hardly improve or even hurt model performance, compared with direct evaluation after pre-training (with and performance drop for METER and FIBER, respectively). However, by enforcing equivariant constraint with EqSim, we observe significant performance improvements than standard fine-tuning, with an absolute gain of for METER and for FIBER.
4 Ablation Study
In this section, we conduct ablation studies to validate the scalability, design and effectiveness of EqSim in terms of enforcing equivariant similarity.
Scalability of EqSim. Table 5 evaluates the scalability of EqSim and standard fine-tuning baseline on the natural subsets of EqBen and Winoground by gradually including more training data. Under the same fine-tuning data, EqSim achieves consistent and significant improvements ( - ) over the baseline. Interestingly, there is no remarkable correlation between the corpus size and model performance. This may be due to the distribution of standard VL data is far away from that of EqBen and Winoground. Note that for the 4M experiment, we fine-tune the models for 10K steps due to computational constraints. Our results demonstrate the potential of EqSim to benefit VL pre-training on large-scale data. Additionally, we validate the generalizability of EqSim in other relevant downstream tasks. Further details are provided in Appendix A.6.
Ablation on EqSim design. Table 6 compares EqSim against the four ablated instances on Eq-Kubric and Winoground , including 1) fine-tuning with hard negative sampling (HardNeg); 2) applying EqSim to all samples in the training batch (EqSim-all); 3) applying EqSim to all samples in the training batch (EqSim-all); and 4) applying EqSim for only semantically close samples (EqSim-close). The final EqSim is equivalent to EqSim-all + EqSim-close, which achieves the best performance. Notably, enforcing EqSim on all (EqSim-all) even degrades the performance by -0.58% on average, compared to applying only to semantically close samples (EqSim-close). This validates our claim in Section 3 that EqSim is better suited for semantically close samples.
Validation of equivariance via EqSim. Given the similarity scores calculated by a VLM, we can define the equivariance score as the derivation of EqSim (headline of Figure 6) to measure the degree of equivariance (the smaller, the better). In Figure 6, we plot the distribution of EqSim values across all samples in EQ-Youcook2 dataset for FIBER and its variants, attached with their group scores. A tighter curve indicates smaller derivation, hence better equivariance similarity measure. Full results on other EqBen subset are presented in Appendix A.7. Compared with pre-training only (PT), fine-tuning on Flickr30K (FT) can improve the group score while being more equivariant in the similarity measure. Adding EqSim (Ours) obtains additional improvements on both similarity equivariance and group score, indicating EqSim indeed enforces equivariant similarity measure. Additionally, due to the space limitation, we leave more visualizations in Appendix A.8.
5 Pilot Study of MLLM on EqBen
Powered by the remarkable capabilities of the large language model (LLM), the community has witnessed an emergent interest in developing Multimodal Large Language Model (MLLM) very recently. Instead of accepting the pure text as the input, MLLM additionally sees the image and provides the response, which can be regarded as another line of VLMs. Here we conduct a pilot study of the performance of MLLM on our EqBen. We adopt LLaVa-7B as our base model with Vicuna as the LLM backend. Given two matched image-text pairs and , we concatenate and horizontally as the single input image. We build the question prompt with the template: “There are two images (left and right). Now you have two captions: caption 1: ; caption 2: . Please indicate which caption corresponds to the left image and which caption corresponds to the right one. The answer should follow the format: ”#index for the left image; #index for the right image”. For example, ”1;2” represents that caption 1 corresponds to image left.” Since it is hard to reformat the MLLM free-form textual output to the label space, we randomly collect 20 samples from each subset of EqBen and manually compare the MLLM output and the ground-truth label. The results are shown in Table 7. Interestingly, by comparing two rows, we can find that the performance of MLLM is quite sensitive to the order of the input caption and ( v.s ). This indicates that the MLLM does NOT truly understand how to distinguish two semantically similar image-text pairs but just follows the given sequence of the captions.
Conclusion
In this study, we investigated the non-equivariant similarity issue in VLMs, hidden behind their excellent performances on standard evaluation benchmarks. To address this issue, we proposed Equivariance Similarity Learning (EqSim), an elegant and effective regularization method that can be easily integrated into the fine-tuning process of existing VLMs. Meanwhile, to better diagnose the equivariance of VLMs, we further introduced a new challenging benchmark EqBen, the first to focus on “visual-minimal change”. Our proposed EqSim is backed by the strong results on both challenging benchmarks (e.g., Winoground, VALSE, EqBen) and the conventional Flickr30K dataset. In future work, we plan to explore the application of EqSim in VL pre-training and instruction tuning. Acknowledgement. We thank Ziyi Dou, Xuejiao Zhao for valuable discussions and help, and all anonymous reviewers for constructive suggestions. This work is partly supported by AI Singapore AISG2-RP-2021-022.
Section A includes full illustrations or more experimental results on EqBen and detailed analysis of EqSim.
Section B provides construction details for EqBen.
Section C presents the implementation details of EqSim.
Section D visualizes more examples in EqBen.
Appendix A More Results
In this section, we include full illustrations and additional experimental results, due to the space limitation of the main paper.
In Figure 1 of the main paper, we perform a toy experiment on LAION400M to compare the similarity measure of FIBER and the human oracle. Due to the space limitation, we only show partial ranking results in the main paper. Here we illustrate the full ranking in Figure 7. With the full ranking results, the observation we summarize in the main paper becomes more clear. That is, the similarity changes in FIBER do not faithfully reflect the semantic changes in images (#1 #25) or text queries “righ” “left”).
A.2 Retrieval Results on COCO dataset
We report the retrieval performance of FIBER variants on COCO 5K test split in Table 8. We observe similar trends on COCO to that on Flickr30K in Table 2. The results suggest the effectiveness of the proposed EqSim, which brings large performance gain across all metrics.
A.3 Full Results of Table 5 and Table 6
We show the full results of ablation studies in Table 9 and Table 10, with group scores across all 5 subsets of EqBen and Winoground. The observation is similar to the main paper. From Table 9, we can find that EqSim is scalable in terms of training data, showing the potential to benefit VL pre-training. The solitary exception happens on Eq-SD, where EqSim cannot consistently obtain the improvements. We hypothesize this is probably because Eq-SD is biased towards the same underlying distribution with the VLMs, as discussed in the main paper. With Table 10, we can find that EqSim (the hybrid combination of EqSim-all and EqSim-close) is the best-performing one, which validates our claim in Section 3. Meanwhile, EqSim-all and EqSim-close also achieve good results (compared with EqSim-all), where both of them are supported by the claim in Section 3.
A.4 Computation Cost of EqSim
We present the computation cost of adding EqSim in the table below. The forward time is measured with the average of 100 times of forward passes on a single GPU. First, EqSim is added as a regularization loss, without additional overhead on # of parameters. On time cost, we observe an acceptable overhead for fusion-encoder (i.e., METER) due to the similarity calculation on negative pairs. While for dual-encoder (i.e., FIBER), which calculates the similarity for each image-text pair, the extra time needed for EqSim is almost negligible. Additionally, we show the forward time consumption v.s. the batch size in the figure below. The computation cost of EqSim linearly scales with batch size, which is only slightly higher than the baseline for each data point.
A.5 More Benchmarking Results on EqBen
We comprehensively report the model performance of existing VLMs on EqBen in Table 11. In addition to the observations drawn in the main paper, we can also find that: 1) When comparing the results of ALBEF/BLIP and their variants with contrastive loss (indicated by ), utilizing cosine similarity as the similarity measure as in ITC often leads to inferior accuracy compared to score computed by the ITM head. As ITC is usually implemented without cross-attention, making it hard to perform the fine-grained semantic recognition required in EqBen. 2) Fine-tuning on Flickr30K (F30K) results in a better performance. In contrast to the noisy samples of the pre-training data, F30K contains high-quality captions that describe images in detail, hence helpful for the equivariant similarity learning of VLMs. 3) The recent method BLIP2 shows strong capacity on our EqBen. Compared to other baselines, it is pre-trained on a much larger vision-language corpus (with 129 million image-text pairs), and thus shows better generalizability.
A.6 Generalization to Video Grounding
To further validate the generalization ability of the proposed EqSim, we conduct additional experiments on a very different but relevant downstream task, zero-shot video boundary grounding task , where the model is required to accurately predict the video boundary indicating event status change, given the before and after query captions. To adapt a pre-trained VLM to this video-language task, we extract video frames at fps=5 first and measure the similarity between each frame and the two query captions. Then given the two adjacent frames () and the two query captions (), we define a boundary grounding score for boundary grounding, where is the similarity produced by VLMs. actually measures whether the boundary is located between frame and . The larger means that and are more likely to be a simultaneous match, thus indicating the boundary between the before and after captions. Results are reported in Table 12 on metrics following . We compute the accuracies based on the absolute distance between ground truth time boundaries and the predicted time boundaries, with the threshold varying from 0.1s to 3s. Across all compared baselines, our EqSim can attain consistent performance improvements on the average accuracy, suggesting that EqSim is effective to identify fine-grained shot changes in videos.
A.7 Distribution Curves on More Subsets
We present the distribution curves of the equivariant score on more EqBen subsets in Figure 13 as the complement to Figure 6 in the main paper. We can find that our EqSim (indicated by “Ours”) indeed achieves the most equivariant similarity (i.e., the tightest curve) across different datasets. Meanwhile, it is worth noting that the equivariance of similarity scores are not always positively correlated to the accuracy. For example, on Eq-SD, EqSim (Ours) is similarly tight as the vanilla fine-tuning (FT), but the accuracy slightly drops.
A.8 More Visualizations
Visualizations of Similarity Scores on Specific Examples. The distribution curves in Figure 6 of the main paper depict the equivariant scores across the whole data. While in Figure 8, we explicitly visualize and compare the similarity scores (blue squares) for specific examples between FIBER baseline and our EqSim. We can clearly observe that current SoTA VLM still falls short in the similarity measure when facing two visually similar images. On one hand, the matching results are not even correct (red cross); On the other hand, regardless of the correctness of the matching results, the similarity scores are not yet equivariant, similar to the Figure 1(b) of the main paper. As shown in the left part of Figure 8, for the same visual semantic change (redgrey), the corresponding similarity change of FIBER is very different from . While our model can produce much more equivariant similarity measure.
Visualizations of Retrieval Results. We visualize the retrieval results on the commonly used Flickr30K in Figure 9. In addition to the better top- retrieval accuracy, our EqSim can produce much more reasonable similarity measurements for the whole retrieval sequence. For example, for the baseline model, the rank and images are not in line with the text of “a young man” and “ throw”. While our top- images are clearly more relevant.
A.9 EqSim vs. CyCLIP [17]
We notice this related contemporaneous work . We compare and discuss more detailed differences here.
Different motivation and implementation. Given the two image-text pairs and , CyCLIP regularizes the CLIP cosine similarity score with the in-modal consistency (forcing to be close to ) and the cross-modal consistency (forcing to be close to ). While our EqSim steps from the motivation that the similarity score change should faithfully respect to the semantic change and derive to the two regularization terms in Eq. (6). Our final objective is the weighted combination of such two terms.
Different evaluation settings and tasks. CyCLIP is solely built on dual-encoder architecture (e.g. CLIP) and evaluates the effectiveness on the zero-shot image classification task. While our EqSim can adapt to both dual-encoder and fusion-encoder architectures (e.g. METER and FIBER) and achieve improvements across various VL benchmarks and downstream tasks, e.g., image-text retrieval, vision-language compositionality, and video boundary grounding.
Better performance of EqSim. The closest CyCLIP counterpart to our EqSim is the cross-modal consistency, which we implemented as EqSim-all in Table 6. As we compare EqSim-all against +EqSim, we clearly observe the superior performance of our EqSim.
Appendix B Construction Details of EqBen
In addition to the general construction pipeline of EqBen in Section 5.3 of the main paper, here we include more details specific to each subset. For all subsets built on natural videos, we denote and as two different frames of the video, while () represents the immediate next frame following ().
Source Dataset. Action Genome (AG) captures changes between objects and their pairwise relationships while action occurs. It contains nearly 10K videos with 1.7M visual relationships which can be used for caption generation. Given the scene graph person - attention relationship - spatial relationship - object, we first create the caption with the template “The person is attention relationship object which is spatial relationship him/her.”
Invalid Samples. In AG, we find that sometimes it is hard to tell apart the two adjacent frames due to the continuity of the video data. This results in two problems for the dataset construction as shown in Figure 10 (top): 1) The two images and are too similar, and may be described by the same caption; 2) The two sample pairs and are too similar, leading to many duplicates.
Frame Filter. To solve this problem, we adopt a sparse sampling strategy to select frames as candidates (Figure 10 (bottom)). Specifically, we only choose frames and if and only if at least 2 of 3 relationships are different. This will make sure the distinction between two images in a single sample, thus solving problem #1. Furthermore, for problem #2, we assume that given a chosen frame (), the immediate next frame () is too similar to (). Therefore, if () is chosen, we will skip the subsequent frame (red cross in Figure 10 (bottom)), and move to ().
B.2 Eq-GEBC
GEBC consists of over 170k boundaries associated with captions describing the events before and after the boundaries. It is built upon 12K videos from Kinetic-400 dataset. We construct Eq-GEBC examples based on annotations from the training and validation splits of GEBC. Intuitively, we can directly adopt the frames before and after the boundaries (i.e., and ) as our visual minimally different images, and the provided GEBC annotation before and after the boundaries can be naturally leveraged as the captions.
Invalid Samples. However, similar to AG, we find that it is hard to tell apart the two images separated by a boundary in practice (see Figure 11 (top)). The reason behind is that the boundary of GEBC is annotated as the status change between two video segments (e.g., from “walking” to “running”). Such action words can be hard to recognize from the sampled static frames.
Frame Filter. As shown in Figure 11 (bottom), we propose to skip an additional boundary to choose and as the twin images to enlarge the semantic gap. Meanwhile, we filter out images with captions containing action words (e.g., “up”, “down”, “upward”, “downward” and “towards”), which are hard to infer without temporal information. Finally, to ensure data quality, we perform a manual screening process with 10 graduate students to filter out invalid samples.
Invalid Samples. As shown in Figure 12 (top), we find that for the cooking video, the chosen frame may contain the view of the chef rather than accurately capturing the objects described in the cooking step. This leads to the mismatch between the image and the caption.
B.3 Eq-YouCook2
We utilize YouCook2 as the data source which contains 2K YouTube videos with average duration of 5.3 minutes, summing to a total of 176 hours. The videos have been manually annotated with segmentation boundaries and captions. On average there are 7.7 segments/captions per video, and 8.8 words per caption. We construct Eq-YouCook2 examples based on annotations from the training and validation splits of YouCook2. For each video with segments, we directly select the middle frame as and its annotated caption as , .
Frame Filter. To solve this problem, we adopt a simple yet effective solution with the face detector https://github.com/ageitgey/face_recognition for frame filtering. Specifically, we directly discard the frames with human faces.
B.4 Eq-Kubric
As introduced in the main paper, Eq-Kubric takes advantage of an open-source graphics engine to faithfully generate photo-realistic scene for the given captions. Therefore, the visual-minimal images generation has been translated into the semantic-minimally different captions construction. We categorize the caption change into three aspects: attribute, counting and location. Figure 14 presents the caption construction details. The semantic-minimally difference is ensured by only intervening the corresponding part in the template while leaving other words unchanged.
B.5 Eq-SD
Similar process can be applied to Eq-SD, for which we summarize the construction details in Figure 15. We similarly categorize the textual semantic-minimal editing into three aspects: object change, scene change and attribute change. We randomly select from the aforementioned three aspects to construct the semantic-minimally different captions. However, in contrast to the Kubric engine, the generation quality of the stable diffusion model is heavily correlated to the given textual prompt. Therefore, we design a more fine-grained template selection for SD. We select scene and attribute from a more restricted subset based on object. For example, given the object of “horse” and the category of “object change” (first row of Figure 15), the changed object will be selected from the animals from the same subset (i.e., “cattle”, “elephant”, “goat”, “deer”, “camel” and “zebra”). Meanwhile, the scene shared across two captions will be selected from the first subset of scene (i.e., “standing on the grass”, “in the desert”, “near the river” and “in the zoo”) for rationality.
Appendix C Implementation Details of EqSim
We fine-tune the models on 8 NVIDIA V100 GPUs. The regularization margin and the balancing factor are selected from and . We adopt an image resolution as due to computational constraints. For FIBER, which is implemented with ITC loss for fast retrieval, we adopt the cosine similarity between image and text features as , and then normalize it by a softmax function. The images and text with top- are regarded as the semantically “close” samples to apply EqSim. METER is designed with ITM loss, which does not compute all pairwise similarities in the training batch. Therefore, we leverage the pre-trained METER model to pre-compute and cache all pairwise similarities in Flickr30K training split, prior to fine-tuning. However, the computation of the ITM similarity for each image-text pair of the training set still takes a long time (more than two weeks on 8 V100 GPUs in practice). To further reduce the computation, we apply a “coarse-to-fine” strategy. For a given image, we first select the top-128 similar images, based on the image feature extracted from the METER vision encoder. Assuming each image is associated with 5 captions, we then utilize the ITM head to compute a fine-grained similarity measure for image-text pairs (leading to 1 hour on 8 V100 GPUs). During retrieval fine-tuning, we follow the original METER to sample captions as negatives and additionally sample their counterpart images for EqSim. Furthermore, of items (i.e., ) are selected as hard negative (i.e., semantically close) samples based on the pre-computed similarity matrix. In METER, the similarity score is normalized with a sigmoid activation. We apply other model-specific hyper-parameters (e.g., training epochs and learning rates) following the original METER and FIBER paper .
Appendix D More Examples of EqBen
Figure 16 visualizes examples in EqBen. We can clearly find that the two images from one data sample are visually similar, indicating that our EqBen indeed focuses on visual-minimal change.