Equivariant Similarity for Vision-Language Foundation Models

Tan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, Lijuan Wang

Introduction

Vision-language (VL) training is all about learning “good” features for each modality, such that the features should faithfully represent the underlying semantics. Thanks to the large-scale image-text pairs on the Web, we have abundant multimodal supervision for the two features with the same semantic meaning —each matched image-text pair should have “similar” visual and textual features, and each unmatched pair should have “dissimilar” ones. Thus, the image-text similarity plays a crucial role to define the feature quality in training VL foundation models (VLMs) .

Has the prevailing “matched vs. unmatched” similarity fulfilled its duty? Yes and no. On the one hand, recent VLMs have demonstrated impressive results in various downstream VL tasks such as image-text retrieval. However, on the other hand, it is acknowledged by the community that the VLMs still fall short in nuanced and complex semantic compositions . In this regard, we present a text-to-image retrieval example on LAION400M with the most recent SOTA VLM FIBER . As shown in Figure 1(a), given the query text “the house on the right side of the road”, we first invite 5 graduate students to rank 25 candidate images from most similar to least similar. The continuously decreasing ranking from human judges () is served as the oracle semantic similarity measure. We then compared this ranking with the ones from FIBER (). Although FIBER correctly retrieved the top-11 image (image#1, ranks 1), some semantically incorrect images (e.g., image#25, ranks 17) are falsely ranked higher than the correct ones (e.g., image#2, ranks 20). Furthermore, when modifying the query text with a slight semantic change (“right” →\rightarrow “left”), the rankings remain almost the same. Clearly, the similarity changes in FIBER do not faithfully reflect the semantic changes in images (#1 →\rightarrow #25) or text queries (“right” →\rightarrow “left”).

To quantitatively measure the above inconsistency between semantic and similarity score changes, we consider two matched image-text pairs {I1,T1}\{I_{1},T_{1}\} and {I2,T2}\{I_{2},T_{2}\} that are semantically similar but only different in the number of clocks in Figure 1 (b). With a slight change of clock counts in caption (“2”→\rightarrow“3”), FIBER mistakenly assigns a higher similarity score to {I1,T2}\{I_{1},T_{2}\} rather than {I1,T1}\{I_{1},T_{1}\} (3.833.83 v.s. 3.793.79). Furthermore, the changes in similarity scores guided by the semantic change (“2”↔\leftrightarrow“3”) are highly inconsistent (+0.04+0.04 v.s. −1.81-1.81). Ideally, an equivariant image-text similarity measure should faithfully reflect the semantic change, i.e., the same semantic changes should lead to a similar amount of similarity changes (e.g., −0.22-0.22 v.s. −0.17-0.17 of ours in Figure 1(b)).

Equivariance Loss. To address this non-equivariance issue, we propose Equivariant Similarity Learning (EqSim), which imposes additional equivariance regularization on image-text pairs for VLM learning without additional supervision. Figure 2 illustrates the underlying semantics perceived by human, where each matched pair demonstrates the image and text corresponding to the underlying semantic. Given two matched image-text pairs {I1,T1}\{I_{1},T_{1}\} as semantic 1 and {I2,T2}\{I_{2},T_{2}\} as semantic 2, we can obtain four similarity scores s11s_{11}, s12s_{12}, s22s_{22}, and s21s_{21}. We define Equivariant Similarity to be an image-text similarity function, whose output value should correspond to the underlying semantic change, which can be measured by text or image change.

(Equivariant Similarity) The similarity ss between image and text is equivariant if and only if the following equations hold:

where μ(I)\mu(I) (μ(T)\mu(T)) denotes the measure in image (text) space, i.e., an infinitesimal unit of visual (textual) change. Based on Definition 1, we formally derive EqSim, an equivariance loss for a hybrid learning strategy on both semantically close and distant training pairs (Section 3). Specifically, EqSim directly enforces s11−s12=s22−s21s_{11}-s_{12}=s_{22}-s_{21} and s11−s21=s22−s12s_{11}-s_{21}=s_{22}-s_{12} for semantically close samples; while for semantically distant samples, we derive a simplified formulation of s12=s21s_{12}=s_{21}. We show that adding EqSim as a regularization term improves existing similarity training objectives significantly on challenging datasets (e.g., over 4%4\% on Winoground ) and tricky tasks (e.g., around 30%30\% on VALSE ). EqSim can also retain or even improve retrieval performance on Flickr30K dataset.

Equivariance Benchmark. To further facilitate the proper evaluation of equivariance in VL community, we present a novel evaluation benchmark dubbed EqBen (Section 4). Motivated by the examples in Figure 1(b), EqBen features “slightly” mis-matched pairs with a minimal semantic drift from the matched pairs, as opposed to “very different” matched and unmatched pairs that are easily distinguishable by both non-equivariant and equivariant similarities. Unlike recent efforts focusing on minimal semantic changes in captions, EqBen pivots on diverse visual-minimal changes, automatically curated from time-varying visual contents in natural videos and synthetic engines with more precise control. We benchmark a full spectrum of VLMs on EqBen, and reveal that the non-equivariant similarity in existing VLMs fails easily. On this new test bed, EqSim can serve as a remedy and bring a large performance gain of ∼\sim3% on average.

Our contributions are summarized as follows: (1) We comprehensively study the problem of similarity equivariance in VLMs. We propose EqSim for equivariant training and EqBen for diagnostic evaluation; (2) EqSim is not only theoretically grounded but also simple, effective and easily pluggable; and (3) EqBen clearly diagnoses that conventional evaluation is not responsive to equivariance. Furthermore, EqSim can significantly improve VLMs on EqBen, as well as other challenging benchmarks.

Related Work

Pre-training VL Models. Early object detector (OD)-based methods utilized the offline image region features from a pre-trained object detector . More recent methods mainly learn from image pixels directly in an end-to-end manner . Researchers further categorize VLMs into (1) Dual-Encoder (e.g., CLIP and ALIGN ) and (2) Fusion-Encoder (e.g., METER , FIBER , and ALBEF ). It is worth noting that our proposed EqSim is model-agnostic, and can be easily plugged into the image-text alignment objectives such as Image-Text Matching (ITM) and Image-Text Contrastive (ITC) loss.

Diagnosing VL Models. Years of VL research have spawned a series of VL evaluation kits, from classical VL tasks (e.g., VQA and image captioning ), to more complex contexts, such as adversarial examples , robustness and counterfactual reasoning . However, these benchmarks require manual annotation and their evaluation relies on task-specific model fine-tuning. Another line of work probe VLMs on similarity measure with minimal caption semantic changes while keeping images intact. While our EqBen tries to test whether the inherent image-text similarity measure in existing VLMs is sensitive to visual semantic changes. The most relevant work is ImageCoDe which leverages video frames toward fine-grained image-text retrieval. However, ImageCoDe requires additional human crowdsourcing and is limited to real-world video sources. In contrast, EqBen explores both natural and synthetic ways to generate image pairs with minimal semantic change, making the data generation process inclusive, automatic, and extensive.

Equivariance Learning. Unlike the wide usage of invariance in deep neural networks (e.g., shift invariance achieved by convolutional layers), strict group equivariance is hard to apply in practice. However, the equivariance property still plays an important role in various fields, such as self-supervised learning , representation learning , and language understanding . In this paper, we point out the significance of the equivariant similarity measure in VLMs. Based on this, we further propose a novel loss EqSim for the regularization of equivariance, as well as a new challenging benchmark EqBen to diagnose the equivariance of existing VLMs. We notice that the recent CyCLIP delivers a similar idea but with different motivation, implementation and evaluation settings. In Table 6, we compare with CyCLIP-equivalent baseline as EqSimv1{}_{\textrm{v1}}. Check more detailed comparison in Appendix.

Improving VLMs with EqSim

Recall that VLMs adopt the image-text similarity as the training proxy for learning multimodal feature representation . Therefore, the ultimate goal of pursuing equivariant similarity is to learn an equivariant feature map between image space and text space.

Let I\mathcal{I} and T\mathcal{T} be two continuous feature spaces. Let G\mathcal{G} be a group whose group action on I\mathcal{I} is defined by g:I→Ig:\mathcal{I}\to\mathcal{I}, and that on T\mathcal{T} is defined by g′:T→Tg^{\prime}:\mathcal{T}\to\mathcal{T}. Then, ϕ:I→T\phi:\mathcal{I}\to\mathcal{T} is an equivariant feature map if and only if g′⋅ϕ(I)=ϕ(g⋅I)g^{\prime}\cdot\phi(I)=\phi(g\cdot I) for all the group actions and I∈II\in\mathcal{I}. The commutativity for ϕ\phi, gg and g′g^{\prime} is shown in Figure 3.

In Definition 1, the measure μ\mu can be considered as the semantic group acting on an infinitesimal region in image or text space. Thus, by applying the commutativity of Definition 2 in Definition 1 to change the sum from image space to text space. Without loss of generality, we only show the results of sum 1→21\to 2:

​​​This implies that the equivariant map establishes an isometry for the measure μ(I)\mu(I) in image space and ϕ(μ(I))\phi(\mu(I)) in text space. Thus, they only differ by a constant scale C>0C>0, i.e., ϕ(μ(I))=Cμ(I)\phi(\mu(I))=C\mu(I):

​​By combining Eq. (1), (2), (3), and (4), we have the following ratio equality as our EqSim constraint:

​Note that C=1C=1 can be derived by using the fact that s11>s12s_{11}>s_{12} and s22>s21s_{22}>s_{21}. By simplifying Eq. (5) further, we have the following two regularizations:

​​Note that the viable space of EqSimv2{}_{\textrm{v2}} is a subset of EqSimv1{}_{\textrm{v1}}, because EqSimv1{}_{\textrm{v1}} is exactly equivalent to Eq. (5) while EqSimv2{}_{\textrm{v2}} further requires s11=s22s_{11}=s_{22}. Empirically, we find that EqSimv2{}_{\textrm{v2}} is more suitable to the semantically close pairs (I1,T1)(I_{1},T_{1}) and (I2,T2)(I_{2},T_{2}); and EqSimv1{}_{\textrm{v1}} to distant pairs. Figure 4 illustrates such hybrid training loss within a training batch. Semantically “close” and “distant” are determined by the similarity score ss, where we regard samples with top-kk ss as “close” samples. For dual encoder VLMs with ITC loss, ss is the cosine similarity between image and text features. For fusion encoder VLMs with ITM, ss is the scoring output from the ITM head.

In our implementation, we adopt Mean Square Error (MSE) loss to regularize the equation of similarities. In addition, motivated by the hinge loss , we utilize a margin parameter α\alpha to control the strength of regularization. EqSimv1{}_{\textrm{v1}} can be written as [∣∣s12−s21∣∣22−α]+[||s_{12}-s_{21}||_{2}^{2}-\alpha]_{+}, where [x]+=max(x,0)[x]_{+}=\text{max}(x,0) and ∣∣⋅∣∣2||\cdot||_{2} denotes the L2 norm. EqSimv2{}_{\textrm{v2}} can be implemented similarly. In practice, given a retrieval fine-tuning objective LRet\mathcal{L}_{Ret}, the final loss can be written as: L=LRet+βL\textscEq\mathcal{L}=\mathcal{L}_{Ret}+\beta\mathcal{L}_{\textsc{Eq}}, where L\textscEq\mathcal{L}_{\textsc{Eq}} is EqSimv1{}_{\textrm{v1}} (EqSimv2{}_{\textrm{v2}}) for semantically distant (close) samples, β\beta is the balancing factor. Experiments in Section 5.4 validate that the hybrid training is better than only using EqSimv1{}_{\textrm{v1}}.

Diagnosing VLMs with EqBen

We argue that standard VL evaluation kits are too coarse to evaluate the equivariance of VLMs similarity. Existing VLMs can easily distinguish most samples in conventional retrieval benchmarks, e.g., images of a group of people against those with cars, given a caption of “people standing on the street”. Therefore, we propose EqBen to focus on visual minimal semantic changes to check whether VLMs can faithfully respond, i.e., the equivariance of the similarity measure in VLMs. Specifically, EqBen contains 5 sub-datasets, covering diverse image domains, from real-life scenarios to synthetic well-controlled scenes. And it is designed to stress test VLMs with accurate semantic changes in action, location, and attribution (e.g., color, count and size). Figure 5 presents an overview of EqBen.

Next, we introduce our design principle for constructing EqBen. Each sample in EqBen consists of a pair of images (I1,I2)(I_{1},I_{2}) and a pair of captions (T1,T2)(T_{1},T_{2}). A valid EqBen sample must satisfy: (1) TiT_{i} is preferred to be used as the description for IiI_{i}; (2) I1I_{1} and I2I_{2} are visual-minimally different. The former one requires that {I1,T1}\{I_{1},T_{1}\} and {I2,T2}\{I_{2},T_{2}\} should be semantically distinguishable without confusion, while the latter one limits the extent of the distinction – “visual-minimal change”. Previous work defines “minimal” semantic change in the caption space as the same words but in a different order. However, due to the continuity and entanglement of image pixels, “minimal” semantic change in visual space is hard to determine. In this paper, we roughly define it as changes in the foreground (e.g., attribute, action, location, etc.) while sharing the same scene and background.

In practice, we source image pairs with “visual-minimal change” in two ways: (1) from natural videos and (2) from synthetic engines, where we adopt different construction pipelines, as shown at the bottom of Figure 5. For the former one, we directly leverage the continuity of scene changes along the temporal dimension in natural videos, which can provide massive image pairs with minimal visual changes. Specifically, we leverage the existing video-language datasets to construct EqBen samples. To more precisely control the varying component in images, we further explore the photo-realistic scene generator (Kubric ) and the open-source diffusion model (Stable Diffusion ) to synthetically generate pairs of images by providing two captions that are minimally different from each other. In what follows, we introduce the construction pipeline for each sub-dataset in detail.

Let’s define a video with caption annotations as V={Ii,Ti}i=1N\mathcal{V}=\{I_{i},T_{i}\}_{i=1}^{N}, where NN is the number of sampled frames. We assume the “visual-minimal change” is naturally guaranteed between any two frames IiI_{i} and IjI_{j} ideally, where Ii,Ij∈VI_{i},I_{j}\in\mathcal{V}, i≠ji\neq j, as we limit the source video to be either short segment or capturing a fixed scene . However, we find that it is hard to ensure the validity of all the video frame pairs in practice. Therefore, we utilize a frame filter to filter out invalid samples automatically for different video sources. Below, we briefly introduce the dataset construction process and delay the details to Appendix.

We construct three sub-datasets based on real images from natural videos, including Eq-AG, Eq-GEBC and Eq-YouCook2. We construct Eq-AG by leveraging the scene graph annotations from Action Genome (AG) , which capture detailed changes between objects and their pairwise relationships while action occurs. We first use a slot-filling template to translate scene graphs to captions. As videos usually come with redundant frames, we avoid the nearly duplicated frames by sampling frames IiI_{i} and IjI_{j} if and only if at least 2 of 3 pairwise relationships are different. Furthermore, if IjI_{j} is chosen for previous samples, we empirically skip the subsequent 2 frames to Ij+2I_{j+2}. Eq-GEBC is built on GEBC that contains captions describing the event before and after an event boundary. We adopt the frames before and after the boundary as our visual minimally different images. Similarly, we avoid temporal redundancy via sparse sampling across multiple boundaries. We construct Eq-YouCook2 based on YouCook2 , which is sparsely annotated with captions for each cooking step. We construct the dataset by sampling the middle frame of a short video segment as IiI_{i} with its annotated caption as TiT_{i}. We then apply off-the-shelf object detectors to filter scene changes. Please note that EqBen can be easily extended to other video-language datasets by applying the same construction pipeline.

2 Construction from Synthetic Engine

Synthetic engine may provide more precise and controllable visual changes in the generated images, to allow more accurate diagnosis in terms of model failure when evaluating with EqBen. We assume that a synthetic engine can faithfully generate images based on a text prompt describing the image content. Based on this assumption, given a pair of semantic-minimally different captions TiT_{i} and TjT_{j}, we expect the generated IiI_{i} and IjI_{j} to be correspondingly visual-minimally different. In the following, we briefly introduce the utilized engine and how to construct semantic-minimally different captions for each sub-dataset and leave more details to Appendix.

Eq-Kubric takes advantages of Kubric , an open-source graphics engine to generate photo-realistic scenes. Here we adopt Google Scanned Objects (GSO) for scene construction and categorize caption change into three aspects: attribute, counting, and location. For each aspect, we construct 20002000 image-text pairs by intervening corresponding phrases of sentences while leaving other words unchanged. Eq-SD is inspired by the recent advances in diffusion models for text-to-image generation . We utilize the open-source checkpoint v1.4 of Stable Diffusion (SD) with prompt-to-prompt image editing framework to translate two semantic-minimally different captions to a pair of images. Specifically, we elaborately design a set of textual semantic-minimal editing: 1) object change (e.g., “dog”→\to“cat”); 2) scene change (e.g., + “in the winter”); 3) attribute change (e.g., + “with a sunglasses”). Finally, we perform a human evaluation to filter out poor-quality generations. Notably, we can adopt more rendered objects (e.g., rendered animals) and various synthetic engine (e.g., better generative models) to further extend our EqBen following the proposed pipeline.

3 Comparisons with Other Datasets

In Table 1, we conduct a direct comparison of EqBen against two widely adopted retrieval benchmarks (Flickr30K and COCO ) and two recent datasets with textual-minimal change (VALSE and Winoground ) from four aspects. 1) On the dataset characteristics, to the best of our knowledge, EqBen is the first diagnosing benchmark to examine the equivariance of VLMs in terms of minimal visual semantic change. 2) For evaluation setting, pairwise setting asks VLMs to select the correct counterpart within a pair of slightly different samples rather than thousands of very different samples in conventional retrieval datasets. The minimal semantic drift between the pair of samples makes the evaluation of equivariant similarity measure more effective. 3) For domain diversity, our EqBen contains rich visual contents collected from different video domains as well as synthetic domains, as opposed to the common image-text datasets which existing diagnosing kits are built upon. 4) In terms of scalability, EqBen is highly scalable as our automatic pipeline can be easily applied to other video-language datasets and synthetic engines, as opposed to manual annotation for building traditional retrieval datasets. While VALSE only focuses on linguistic editing of the captions, EqBen can be further scale up with more diverse visual contents.

Experiments

We first introduce our experimental setting in Section 5.1, followed by evaluation of EqSim on existing benchmarks in Section 5.2. Section 5.3 benchmarks SOTA VLMs on EqBen to show their insensitivity to minimal visual semantic changes, and we further validate EqSim on EqBen. Section 5.4 presents additional ablation studies to examine the design of EqSim.

Training Details. Recent efforts on diagnosing benchmarks only provide testing data and directly evaluate models after VL pre-training. The low performance reported on these benchmarks can mainly be attributed to two factors: 1) the inherent weaknesses of VLMs, e.g., non-equivariant similarity measure; and 2) the domain gap between training and testing. To better validate the effectiveness of our method, we fine-tune the VLMs on limited image-text pairs from conventional retrieval dataset Flickr30K with or without the regularization term of EqSim, and then test the fine-tuned VLMs on the challenging Winoground , VALSE and our EqBen. Under a fair comparison, we argue that the absolute performance improvements from EqSim thus would suggest that the gain is entirely from the remedy of model weaknesses.

To validate the effectiveness and genraliazability of our proposed method, we apply EqSim to two SOTA end-to-end methods with different architectures and retrieval losses. Specifically, FIBER supports the dual encoder with ITC loss for fast retrieval, which computes similarities for N2N^{2} image-text pairs with only O(N)O(N) forwarding. In contrast, the SOTA fusion-encoder model METER , optimized with ITM task during pre-training, computes the similarity by forwarding the concatenation of each pair of image and text, resulting in O(N2)O(N^{2}) time complexity. Fine-tuning details for each model can be found in Appendix.

Evaluation Metric. On Winoground , given two image-text pairs {I1,T1}\{I_{1},T_{1}\} and {I2,T2}\{I_{2},T_{2}\}, a VLM measures similarity sijs_{ij} between image IiI_{i} and text TjT_{j} (i,j∈{0,1},i≠j)(i,j\in\{0,1\},i\neq j). Three metrics are computed based on sijs_{ij}: 1) Text score measures whether the model can select the correct text for a given image. The model wins one point if sii>sijs_{ii}>s_{ij}. 2) Image score evaluates if VLMs can select the correct image for a given text and the model wins one point when sii>sjis_{ii}>s_{ji}. 3) Group score combines the previous two, such that the VLMs win one point if and only if both text score and image score are 1, meaning the following condition must be satisfied: sii>sijs_{ii}>s_{ij} and sii>sjis_{ii}>s_{ji}. On VALSE , the two image-text pairs share a common image, i.e., {I1,T1}\{I_{1},T_{1}\} (correct) and {I1,T2}\{I_{1},T_{2}\} (foil). We follow to report the following metrics: 1) acc is the overall accuracy on both correct and foil image-text pairs; and 2) min(pc,pfp_{c},p_{f}) is the minimum of precision pcp_{c} and foil precision pfp_{f}, where pcp_{c} (pfp_{f}) measures how well models identify the correct (foil) pair. We also report performance on the conventional image-text retrieval task, where recall R@K (K=1,5,10) is used as the evaluation metric.

2 Evaluation of EqSim

In Table 2, we compare model performance under three settings: (ii) direct evaluation after pre-training (the first rows of each block); (iiii) standard fine-tuning (FT) on Flickr30K training data (the second rows of each block); and (iiiiii) fine-tuning with EqSim regularization (the third rows of each block). We observe that standard fine-tuning can somewhat bring a little performance improvement on both Winoground and VALSE benchmarks, indicating that some domain overlap between Flickr30K training data and testing samples. It is difficult to entirely rule out the domain influence, but comparing fine-tuning with EqSim against standard fine-tuning, our method brings consistent and significant performance improvements on both of the challenging Winoground and VALSE across METER and FIBER models. Specifically, EqSim improves the group score over standard fine-tuning by 4%4\% for METER and 4.5%4.5\% for FIBER on Winoground, respectively. While for VALSE, the performance improvement on min(pc,pf)(p_{c},p_{f}) is as large as 31.6%31.6\%, further validating the effectiveness of our EqSim. In addition, we observe that the equivariance regularization from EqSim does not sacrifice retrieval performance. On Flickr30K, EqSim can mostly retain the retrieval performance, and sometimes even yield performance gain, e.g., 3.1%3.1\% on R@1 for image-to-text retrieval with FIBER.

3 Benchmarking VLMs with EqBen

We evaluate a wide range of VLMs with different configurations on EqBen in a zero-shot manner, to examine the equivariance of their similarity measures for distinguishing visually-minimal different samples. We consider representative VLMs, including (ii) LXMERT , ViLBERT for OD-Based models; and (iiii) CLIP variants, FLAVA , ViLT , ALBEF , BLIP and METER and FIBER as prominent examples of end-to-end SOTA methods. Full results on more VLMs can be found in Appendix A.5. For evaluation metrics, we adopt text score, image score and group score to compare model performance, similar to Winoground .

Table 3 presents the evaluation results of existing VLMs on EqBen and we summarize our observations below.

Regardless of the subsets, end-to-end VLMs generally achieve better performance as it is not constrained by the fixed visual representation from a pre-trained object detector , as in OD-based methods.

Among all subsets, VLMs obtain evidently higher performance on Eq-SD. The stable diffusion model is pre-trained on similar VL corpus to these VLMs. Hence, the generated images can be biased towards the same underlying data distribution, much easier for VLMs to tell the differences. Besides, the generated images maybe visually minimally different to human eyes, but it is unclear whether in the pixel space, they are minimally different w.r.t. the model input. It is worth noting that LXMERT and ViLBERT are the exception due to the totally different distribution with the off-the-shelf object detector.

Interestingly, a larger pre-training corpus (e.g., CLIP and FLAVA ) does not always guarantee better results. This implies training loss may be more critical in learning equivariant similarity measure.

In Table 4, we further conduct a fine-grained examination with the synthetic subset EQ-Kuric, where we focus on specific visual changes in location, counting and attribute. VLMs fail substantially in terms of location and counting, while being sensitive to attribute changes. Similar findings are also observed by from the text side.

We again equip the two strong baseline models (METER and FIBER) with EqSim and fine-tune on Flickr30K. As EqBen covers diverse domains, standard fine-tuning on Flickr30K can hardly improve or even hurt model performance, compared with direct evaluation after pre-training (with −0.62%-0.62\% and −1.46%-1.46\% performance drop for METER and FIBER, respectively). However, by enforcing equivariant constraint with EqSim, we observe significant performance improvements than standard fine-tuning, with an absolute gain of 3.26%3.26\% for METER and 1.92%1.92\% for FIBER.

4 Ablation Study

In this section, we conduct ablation studies to validate the scalability, design and effectiveness of EqSim in terms of enforcing equivariant similarity.

Scalability of EqSim. Table 5 evaluates the scalability of EqSim and standard fine-tuning baseline on the natural subsets of EqBen and Winoground by gradually including more training data. Under the same fine-tuning data, EqSim achieves consistent and significant improvements (2%2\% - 3%3\%) over the baseline. Interestingly, there is no remarkable correlation between the corpus size and model performance. This may be due to the distribution of standard VL data is far away from that of EqBen and Winoground. Note that for the 4M experiment, we fine-tune the models for 10K steps due to computational constraints. Our results demonstrate the potential of EqSim to benefit VL pre-training on large-scale data. Additionally, we validate the generalizability of EqSim in other relevant downstream tasks. Further details are provided in Appendix A.6.

Ablation on EqSim design. Table 6 compares EqSim against the four ablated instances on Eq-Kubric and Winoground , including 1) fine-tuning with hard negative sampling (HardNeg); 2) applying EqSimv1{}_{\textrm{v1}} to all samples in the training batch (EqSimv1{}_{\textrm{v1}}-all); 3) applying EqSimv2{}_{\textrm{v2}} to all samples in the training batch (EqSimv2{}_{\textrm{v2}}-all); and 4) applying EqSimv2{}_{\textrm{v2}} for only semantically close samples (EqSimv2{}_{\textrm{v2}}-close). The final EqSim is equivalent to EqSimv1{}_{\textrm{v1}}-all + EqSimv2{}_{\textrm{v2}}-close, which achieves the best performance. Notably, enforcing EqSimv2{}_{\textrm{v2}} on all (EqSimv2{}_{\textrm{v2}}-all) even degrades the performance by -0.58% on average, compared to applying only to semantically close samples (EqSimv2{}_{\textrm{v2}}-close). This validates our claim in Section 3 that EqSimv2{}_{\textrm{v2}} is better suited for semantically close samples.

Validation of equivariance via EqSim. Given the similarity scores ss calculated by a VLM, we can define the equivariance score as the derivation of EqSimv2{}_{\textrm{v2}} (headline of Figure 6) to measure the degree of equivariance (the smaller, the better). In Figure 6, we plot the distribution of EqSimv2{}_{\textrm{v2}} values across all samples in EQ-Youcook2 dataset for FIBER and its variants, attached with their group scores. A tighter curve indicates smaller derivation, hence better equivariance similarity measure. Full results on other EqBen subset are presented in Appendix A.7. Compared with pre-training only (PT), fine-tuning on Flickr30K (FT) can improve the group score while being more equivariant in the similarity measure. Adding EqSim (Ours) obtains additional improvements on both similarity equivariance and group score, indicating EqSim indeed enforces equivariant similarity measure. Additionally, due to the space limitation, we leave more visualizations in Appendix A.8.

5 Pilot Study of MLLM on EqBen

Powered by the remarkable capabilities of the large language model (LLM), the community has witnessed an emergent interest in developing Multimodal Large Language Model (MLLM) very recently. Instead of accepting the pure text as the input, MLLM additionally sees the image and provides the response, which can be regarded as another line of VLMs. Here we conduct a pilot study of the performance of MLLM on our EqBen. We adopt LLaVa-7B as our base model with Vicuna as the LLM backend. Given two matched image-text pairs {I1,T1}\{I_{1},T_{1}\} and {I2,T2}\{I_{2},T_{2}\}, we concatenate I1I_{1} and I2I_{2} horizontally as the single input image. We build the question prompt with the template: “There are two images (left and right). Now you have two captions: caption 1: {T1}\{T_{1}\}; caption 2: {T2}\{T_{2}\}. Please indicate which caption corresponds to the left image and which caption corresponds to the right one. The answer should follow the format: ”#index for the left image; #index for the right image”. For example, ”1;2” represents that caption 1 corresponds to image left.” Since it is hard to reformat the MLLM free-form textual output to the label space, we randomly collect 20 samples from each subset of EqBen and manually compare the MLLM output and the ground-truth label. The results are shown in Table 7. Interestingly, by comparing two rows, we can find that the performance of MLLM is quite sensitive to the order of the input caption T1T_{1} and T2T_{2} (∼90%\sim 90\% v.s ∼0%\sim 0\%). This indicates that the MLLM does NOT truly understand how to distinguish two semantically similar image-text pairs but just follows the given sequence of the captions.

Conclusion

In this study, we investigated the non-equivariant similarity issue in VLMs, hidden behind their excellent performances on standard evaluation benchmarks. To address this issue, we proposed Equivariance Similarity Learning (EqSim), an elegant and effective regularization method that can be easily integrated into the fine-tuning process of existing VLMs. Meanwhile, to better diagnose the equivariance of VLMs, we further introduced a new challenging benchmark EqBen, the first to focus on “visual-minimal change”. Our proposed EqSim is backed by the strong results on both challenging benchmarks (e.g., Winoground, VALSE, EqBen) and the conventional Flickr30K dataset. In future work, we plan to explore the application of EqSim in VL pre-training and instruction tuning. Acknowledgement. We thank Ziyi Dou, Xuejiao Zhao for valuable discussions and help, and all anonymous reviewers for constructive suggestions. This work is partly supported by AI Singapore AISG2-RP-2021-022.

Section A includes full illustrations or more experimental results on EqBen and detailed analysis of EqSim.

Section B provides construction details for EqBen.

Section C presents the implementation details of EqSim.

Section D visualizes more examples in EqBen.

Appendix A More Results

In this section, we include full illustrations and additional experimental results, due to the space limitation of the main paper.

In Figure 1 of the main paper, we perform a toy experiment on LAION400M to compare the similarity measure of FIBER and the human oracle. Due to the space limitation, we only show partial ranking results in the main paper. Here we illustrate the full ranking in Figure 7. With the full ranking results, the observation we summarize in the main paper becomes more clear. That is, the similarity changes in FIBER do not faithfully reflect the semantic changes in images (#1 →\rightarrow #25) or text queries “righ” →\rightarrow “left”).

A.2 Retrieval Results on COCO dataset

We report the retrieval performance of FIBER variants on COCO 5K test split in Table 8. We observe similar trends on COCO to that on Flickr30K in Table 2. The results suggest the effectiveness of the proposed EqSim, which brings large performance gain across all metrics.

A.3 Full Results of Table 5 and Table 6

We show the full results of ablation studies in Table 9 and Table 10, with group scores across all 5 subsets of EqBen and Winoground. The observation is similar to the main paper. From Table 9, we can find that EqSim is scalable in terms of training data, showing the potential to benefit VL pre-training. The solitary exception happens on Eq-SD, where EqSim cannot consistently obtain the improvements. We hypothesize this is probably because Eq-SD is biased towards the same underlying distribution with the VLMs, as discussed in the main paper. With Table 10, we can find that EqSim (the hybrid combination of EqSimv1{}_{\textrm{v1}}-all and EqSimv2{}_{\textrm{v2}}-close) is the best-performing one, which validates our claim in Section 3. Meanwhile, EqSimv1{}_{\textrm{v1}}-all and EqSimv2{}_{\textrm{v2}}-close also achieve good results (compared with EqSimv2{}_{\textrm{v2}}-all), where both of them are supported by the claim in Section 3.

A.4 Computation Cost of EqSim

We present the computation cost of adding EqSim in the table below. The forward time is measured with the average of 100 times of forward passes on a single GPU. First, EqSim is added as a regularization loss, without additional overhead on # of parameters. On time cost, we observe an acceptable overhead for fusion-encoder (i.e., METER) due to the similarity calculation on negative pairs. While for dual-encoder (i.e., FIBER), which calculates the similarity for each image-text pair, the extra time needed for EqSim is almost negligible. Additionally, we show the forward time consumption v.s. the batch size in the figure below. The computation cost of EqSim linearly scales with batch size, which is only slightly higher than the baseline for each data point.

A.5 More Benchmarking Results on EqBen

We comprehensively report the model performance of existing VLMs on EqBen in Table 11. In addition to the observations drawn in the main paper, we can also find that: 1) When comparing the results of ALBEF/BLIP and their variants with contrastive loss (indicated by ‡\ddagger), utilizing cosine similarity as the similarity measure as in ITC often leads to inferior accuracy compared to score computed by the ITM head. As ITC is usually implemented without cross-attention, making it hard to perform the fine-grained semantic recognition required in EqBen. 2) Fine-tuning on Flickr30K (F30K) results in a better performance. In contrast to the noisy samples of the pre-training data, F30K contains high-quality captions that describe images in detail, hence helpful for the equivariant similarity learning of VLMs. 3) The recent method BLIP2 shows strong capacity on our EqBen. Compared to other baselines, it is pre-trained on a much larger vision-language corpus (with 129 million image-text pairs), and thus shows better generalizability.

A.6 Generalization to Video Grounding

To further validate the generalization ability of the proposed EqSim, we conduct additional experiments on a very different but relevant downstream task, zero-shot video boundary grounding task , where the model is required to accurately predict the video boundary indicating event status change, given the before and after query captions. To adapt a pre-trained VLM to this video-language task, we extract video frames at fps=5 first and measure the similarity between each frame and the two query captions. Then given the two adjacent frames (I1,I2I_{1},I_{2}) and the two query captions (T1,T2T_{1},T_{2}), we define a boundary grounding score sbg=s(I1,T1)+s(I2,T2)s_{bg}=s(I_{1},T_{1})+s(I_{2},T_{2}) for boundary grounding, where ss is the similarity produced by VLMs. sbgs_{bg} actually measures whether the boundary is located between frame I1I_{1} and I2I_{2}. The larger sbgs_{bg} means that (I1,T1)(I_{1},T_{1}) and (I2,T2)(I_{2},T_{2}) are more likely to be a simultaneous match, thus indicating the boundary between the before and after captions. Results are reported in Table 12 on metrics following . We compute the accuracies based on the absolute distance between ground truth time boundaries and the predicted time boundaries, with the threshold varying from 0.1s to 3s. Across all compared baselines, our EqSim can attain consistent performance improvements on the average accuracy, suggesting that EqSim is effective to identify fine-grained shot changes in videos.

A.7 Distribution Curves on More Subsets

We present the distribution curves of the equivariant score on more EqBen subsets in Figure 13 as the complement to Figure 6 in the main paper. We can find that our EqSim (indicated by “Ours”) indeed achieves the most equivariant similarity (i.e., the tightest curve) across different datasets. Meanwhile, it is worth noting that the equivariance of similarity scores are not always positively correlated to the accuracy. For example, on Eq-SD, EqSim (Ours) is similarly tight as the vanilla fine-tuning (FT), but the accuracy slightly drops.

A.8 More Visualizations

Visualizations of Similarity Scores on Specific Examples. The distribution curves in Figure 6 of the main paper depict the equivariant scores across the whole data. While in Figure 8, we explicitly visualize and compare the similarity scores (blue squares) for specific examples between FIBER baseline and our EqSim. We can clearly observe that current SoTA VLM still falls short in the similarity measure when facing two visually similar images. On one hand, the matching results are not even correct (red cross); On the other hand, regardless of the correctness of the matching results, the similarity scores are not yet equivariant, similar to the Figure 1(b) of the main paper. As shown in the left part of Figure 8, for the same visual semantic change (red↔\leftrightarrowgrey), the corresponding similarity change of FIBER s11−s21=−0.07s_{11}-s_{21}=-0.07 is very different from s22−s12=1.4s_{22}-s_{12}=1.4. While our model can produce much more equivariant similarity measure.

Visualizations of Retrieval Results. We visualize the retrieval results on the commonly used Flickr30K in Figure 9. In addition to the better top-11 retrieval accuracy, our EqSim can produce much more reasonable similarity measurements for the whole retrieval sequence. For example, for the baseline model, the rank 22 and 33 images are not in line with the text of “a young man” and “ throw”. While our top-33 images are clearly more relevant.

A.9 EqSim vs. CyCLIP [17]

We notice this related contemporaneous work . We compare and discuss more detailed differences here.

Different motivation and implementation. Given the two image-text pairs {I1,T1}\{I_{1},T_{1}\} and {I2,T2}\{I_{2},T_{2}\}, CyCLIP regularizes the CLIP cosine similarity score ss with the in-modal consistency (forcing s(I1,I2)s(I_{1},I_{2}) to be close to s(T1,T2)s(T_{1},T_{2})) and the cross-modal consistency (forcing s(I1,T2)s(I_{1},T_{2}) to be close to s(I2,T1)s(I_{2},T_{1})). While our EqSim steps from the motivation that the similarity score change should faithfully respect to the semantic change and derive to the two regularization terms in Eq. (6). Our final objective is the weighted combination of such two terms.

Different evaluation settings and tasks. CyCLIP is solely built on dual-encoder architecture (e.g. CLIP) and evaluates the effectiveness on the zero-shot image classification task. While our EqSim can adapt to both dual-encoder and fusion-encoder architectures (e.g. METER and FIBER) and achieve improvements across various VL benchmarks and downstream tasks, e.g., image-text retrieval, vision-language compositionality, and video boundary grounding.

Better performance of EqSim. The closest CyCLIP counterpart to our EqSim is the cross-modal consistency, which we implemented as EqSimv1{}_{\textrm{v1}}-all in Table 6. As we compare EqSimv1{}_{\textrm{v1}}-all against +EqSim, we clearly observe the superior performance of our EqSim.

Appendix B Construction Details of EqBen

In addition to the general construction pipeline of EqBen in Section 5.3 of the main paper, here we include more details specific to each subset. For all subsets built on natural videos, we denote IiI_{i} and IjI_{j} as two different frames of the video, while Ii+1I_{i+1} (Ij+1I_{j+1}) represents the immediate next frame following IiI_{i} (IjI_{j}).

Source Dataset. Action Genome (AG) captures changes between objects and their pairwise relationships while action occurs. It contains nearly 10K videos with 1.7M visual relationships which can be used for caption generation. Given the scene graph ⟨\langleperson - attention relationship - spatial relationship - object⟩\rangle, we first create the caption with the template “The person is ⟨\langleattention relationship⟩\rangle ⟨\langleobject⟩\rangle which is ⟨\langlespatial relationship⟩\rangle him/her.”

Invalid Samples. In AG, we find that sometimes it is hard to tell apart the two adjacent frames due to the continuity of the video data. This results in two problems for the dataset construction as shown in Figure 10 (top): 1) The two images IiI_{i} and IjI_{j} are too similar, and may be described by the same caption; 2) The two sample pairs {Ii,Ij}\{I_{i},I_{j}\} and {Ii,Ij+1}\{I_{i},I_{j+1}\} are too similar, leading to many duplicates.

Frame Filter. To solve this problem, we adopt a sparse sampling strategy to select frames as candidates (Figure 10 (bottom)). Specifically, we only choose frames IiI_{i} and IjI_{j} if and only if at least 2 of 3 relationships are different. This will make sure the distinction between two images in a single sample, thus solving problem #1. Furthermore, for problem #2, we assume that given a chosen frame IiI_{i} (IjI_{j}), the immediate next frame Ii+1I_{i+1} (Ij+1I_{j+1}) is too similar to IiI_{i} (IjI_{j}). Therefore, if IiI_{i} (IjI_{j}) is chosen, we will skip the subsequent frame (red cross in Figure 10 (bottom)), and move to Ii+2I_{i+2} (Ij+2I_{j+2}).

B.2 Eq-GEBC

GEBC consists of over 170k boundaries associated with captions describing the events before and after the boundaries. It is built upon 12K videos from Kinetic-400 dataset. We construct Eq-GEBC examples based on annotations from the training and validation splits of GEBC. Intuitively, we can directly adopt the frames before and after the boundaries (i.e., IiI_{i} and Ii+1I_{i+1}) as our visual minimally different images, and the provided GEBC annotation before and after the boundaries can be naturally leveraged as the captions.

Invalid Samples. However, similar to AG, we find that it is hard to tell apart the two images separated by a boundary in practice (see Figure 11 (top)). The reason behind is that the boundary of GEBC is annotated as the status change between two video segments (e.g., from “walking” to “running”). Such action words can be hard to recognize from the sampled static frames.

Frame Filter. As shown in Figure 11 (bottom), we propose to skip an additional boundary to choose IiI_{i} and Ii+2I_{i+2} as the twin images to enlarge the semantic gap. Meanwhile, we filter out images with captions containing action words (e.g., “up”, “down”, “upward”, “downward” and “towards”), which are hard to infer without temporal information. Finally, to ensure data quality, we perform a manual screening process with 10 graduate students to filter out invalid samples.

Invalid Samples. As shown in Figure 12 (top), we find that for the cooking video, the chosen frame may contain the view of the chef rather than accurately capturing the objects described in the cooking step. This leads to the mismatch between the image and the caption.

B.3 Eq-YouCook2

We utilize YouCook2 as the data source which contains 2K YouTube videos with average duration of 5.3 minutes, summing to a total of 176 hours. The videos have been manually annotated with segmentation boundaries and captions. On average there are 7.7 segments/captions per video, and 8.8 words per caption. We construct Eq-YouCook2 examples based on annotations from the training and validation splits of YouCook2. For each video with NN segments, we directly select the middle frame as IiI_{i} and its annotated caption as TiT_{i}, i∈{1,2,...,N}i\in\{1,2,...,N\}.

Frame Filter. To solve this problem, we adopt a simple yet effective solution with the face detector https://github.com/ageitgey/face_recognition for frame filtering. Specifically, we directly discard the frames with human faces.

B.4 Eq-Kubric

As introduced in the main paper, Eq-Kubric takes advantage of an open-source graphics engine to faithfully generate photo-realistic scene for the given captions. Therefore, the visual-minimal images generation has been translated into the semantic-minimally different captions construction. We categorize the caption change into three aspects: attribute, counting and location. Figure 14 presents the caption construction details. The semantic-minimally difference is ensured by only intervening the corresponding part in the template while leaving other words unchanged.

B.5 Eq-SD

Similar process can be applied to Eq-SD, for which we summarize the construction details in Figure 15. We similarly categorize the textual semantic-minimal editing into three aspects: object change, scene change and attribute change. We randomly select from the aforementioned three aspects to construct the semantic-minimally different captions. However, in contrast to the Kubric engine, the generation quality of the stable diffusion model is heavily correlated to the given textual prompt. Therefore, we design a more fine-grained template selection for SD. We select ⟨\langlescene⟩\rangle and ⟨\langleattribute⟩\rangle from a more restricted subset based on ⟨\langleobject⟩\rangle. For example, given the object of “horse” and the category of “object change” (first row of Figure 15), the changed object will be selected from the animals from the same subset (i.e., “cattle”, “elephant”, “goat”, “deer”, “camel” and “zebra”). Meanwhile, the scene shared across two captions will be selected from the first subset of scene (i.e., “standing on the grass”, “in the desert”, “near the river” and “in the zoo”) for rationality.

Appendix C Implementation Details of EqSim

We fine-tune the models on 8 NVIDIA V100 GPUs. The regularization margin α\alpha and the balancing factor β\beta are selected from {0,0.04,0.1}\{0,0.04,0.1\} and {0.2,0.5,1.0}\{0.2,0.5,1.0\}. We adopt an image resolution as 288×288288\times 288 due to computational constraints. For FIBER, which is implemented with ITC loss for fast retrieval, we adopt the cosine similarity between image and text features as ss, and then normalize it by a softmax function. The images and text with top-88 ss are regarded as the semantically “close” samples to apply EqSimv2{}_{\textrm{v2}}. METER is designed with ITM loss, which does not compute all pairwise similarities in the training batch. Therefore, we leverage the pre-trained METER model to pre-compute and cache all pairwise similarities in Flickr30K training split, prior to fine-tuning. However, the computation of the ITM similarity for each image-text pair of the training set still takes a long time (more than two weeks on 8 V100 GPUs in practice). To further reduce the computation, we apply a “coarse-to-fine” strategy. For a given image, we first select the top-128 similar images, based on the image feature extracted from the METER vision encoder. Assuming each image is associated with 5 captions, we then utilize the ITM head to compute a fine-grained similarity measure for 128×5128\times 5 image-text pairs (leading to 1 hour on 8 V100 GPUs). During retrieval fine-tuning, we follow the original METER to sample 1515 captions as negatives and additionally sample their counterpart images for EqSim. Furthermore, 88 of 1515 items (i.e., k=8k=8) are selected as hard negative (i.e., semantically close) samples based on the pre-computed similarity matrix. In METER, the similarity score ss is normalized with a sigmoid activation. We apply other model-specific hyper-parameters (e.g., training epochs and learning rates) following the original METER and FIBER paper .

Appendix D More Examples of EqBen

Figure 16 visualizes examples in EqBen. We can clearly find that the two images from one data sample are visually similar, indicating that our EqBen indeed focuses on visual-minimal change.

References