VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment

Shraman Pramanick, Li Jing, Sayan Nag, Jiachen Zhu, Hardik Shah, Yann LeCun, Rama Chellappa

Introduction

Inspired by the escalating unification of transformer-based modeling in vision (Dosovitskiy et al., 2021; Liu et al., 2021; Chen et al., 2021a) and language (Devlin et al., 2019; Liu et al., 2019) domains, coupled with readily available large-scale image-caption pair data, vision-language pre-training (VLP) (Lu et al., 2019; Li et al., 2020a; Kim et al., 2021; Kamath et al., 2021; Zhang et al., 2021) has recently been receiving increasing attention. VLP has not only been proven the de-facto for several vision-language tasks, but it has also been beneficial for traditional vision-only tasks, such as image classification and object detection. Such wide-range applications of VLP can broadly be categorized into two groups: (i)(i) tasks requiring image-level understanding, e.g., image classification, image & text retrieval (Plummer et al., 2015), image captioning (Zhou et al., 2020), visual question answering (Antol et al., 2015), and (ii)(ii) tasks requiring region-level understanding, e.g., object detection, instance segmentation, and referring expression comprehension (Kazemzadeh et al., 2014; Yu et al., 2016). Most existing VLP methods address only one group of application, leaving the question of a generalizable and unified VL framework under-explored.

Traditional VLP methods with image-level understanding (Li et al., 2021a; Wang et al., 2021b; Dou et al., 2022b) utilize large-scale image-caption pair datasets and are commonly trained with image-text contrastive objectives computed on global features. Hence, it is not trivial to extend such methods to region-level applications. On the other hand, VLP methods with region-level understanding (Kamath et al., 2021; Li et al., 2022c; Zhang et al., 2022) use image-text-box grounding data and are designed to predict bounding boxes during pre-training. Consequently, they do not support image-level tasks. Furthermore, accurate bounding box annotations require high-resolution input images, which are often expensive to collect, annotate and use for pre-training at scale. Recently, FIBER (Dou et al., 2022a) addressed the problem of such unified VLP and proposed a two-stage pre-training algorithm requiring fewer box annotations than previous region-level pre-training methods. Moving a step forward, as shown in Figure 1, we aim to eliminate the use of costly box annotations and ask the challenging but natural question: Can we attain region-level understanding from global image-caption annotations and unify image- and region-level tasks in a single VL framework?

Subsequently, we focus on achieving region-level fine-grained understanding by weakly-supervised alignment of image patches and text tokens. Previous VLP methods (Chen et al., 2020d; Kim et al., 2021) in this direction use Wasserstein distance (WD) (Peyré et al., 2019), a.k.a Earth Mover’s distance (EMD)-based optimal transport (OT) algorithms for such alignment problems. However, we argue that WD is not optimum for images with multiple similar entities. Thus, we propose to jointly utilize Gromov-Wasserstein distance (GWD) (Peyré et al., 2016) and Wasserstein distance (WD) in a setup known as graph optimal transport (Chen et al., 2020a). Moreover, instead of using a commonly deployed contrastive objective, we propose to use redundancy reduction from Barlow Twins (Zbontar et al., 2021), which is less data-intensive and does not require hard-negative mining. We also follow Dou et al. (2022a) and incorporate deep multi-modal fusion into the uni-modal backbones, removing the need for costly fusion-specific transformer layers. These steps when integrated yield VoLTA, Vision-Language Transformer with weakly-supervised local-feature Alignment, a unified VLP paradigm only utilizes image-caption annotations but achieves fine-grained region-level image understanding, eliminating the need for expensive box annotations. Figure 4 visualizes the feature-level image-text alignment generated by VoLTA, which can attend text tokens to the corresponding visual patches without relying on low-level supervision.

In summary, our contributions are three-fold. (i)(i) We propose to use graph optimal transport for weakly-supervised feature-level patch-token alignment in VLP. (ii)(ii) We introduce VoLTA, a unified VLP paradigm for image-level and region-level applications, but pre-trained only using image-caption pairs. VoLTA is memory, compute, and time-efficient and can easily be scaled up with readily available large-scale image-caption data harvested from the web. (iii)(iii) We present the results of a wide range of vision- and vision-language coarse- and fine-grained downstream experiments to demonstrate the effectiveness of VoLTA compared to strong baselines pre-trained with significantly more caption and box annotations.

Related Works

Uni-modal Self-supervised Pre-training: In recent years, the machine learning community has observed a boom in self-supervised pre-training. In the language domain, representations learned by BERT (Devlin et al., 2019), RoBERTa (Liu et al., 2019) have become the default setting for many downstream tasks. Generative models such as GPT (Radford et al., 2019; Brown et al., 2020) have also achieved impressive few-shot/zero-shot performances on novel applications. SimCSE (Gao et al., 2021) uses contrastive learning to help learn useful sentence representations.

In the vision domain, several contrastive/joint-embedding methods (He et al., 2020; Chen et al., 2020c; 2021b; b; Grill et al., 2020; Chen & He, 2021; Caron et al., 2021; Zbontar et al., 2021; Bardes et al., 2022; Shah et al., 2022; Assran et al., 2022) have outperformed supervised counterparts. Recently, generative models such as BEiT (Bao et al., 2021) and MAE (He et al., 2022) have also achieved impressive performances with much more scalable potential.

Vision-Language Pre-training (VLP): Vision-language pre-training mainly relies on image-text pair datasets to learn joint visual-language representations. One line of work is to train separate vision and language encoders and only fuse in the representation space. CLIP (Radford et al., 2021), UniCL (Yang et al., 2022a), and ALIGN (Jia et al., 2021) use the image-text contrastive loss to learn aligned representations. SLIP (Mu et al., 2021) combines self-supervised visual representation learning and contrastive multi-modal learning. M3AE (Geng et al., 2022), FLAVA (Singh et al., 2022) combines masked image modeling and masked language modeling. Another line of work uses cross attention to fuse vision and language information in the early stage (Kamath et al., 2021; Dou et al., 2022b; Lu et al., 2019; Li et al., 2020b; Kiela et al., 2019; Kim et al., 2021; Zhang et al., 2021; Li et al., 2022b; Wang et al., 2022c; Pramanick et al., 2023; Park & Han, 2023; Li et al., 2023a; Jang et al., 2023; Wang et al., 2023a). These works focus on learning semantic-level aligned vision-language representations. In addition, UniTAB (Yang et al., 2022c), OFA (Wang et al., 2022b), GLIP (Li et al., 2022c), and FIBER (Dou et al., 2022a) use expensive grounding image-text-box annotations to learn the fine-grained aligned representations. Our work uses representation space alignment and cross-attention fusion, but we do not use any box annotation to learn robust feature-level alignments.

Unsupervised Representation Alignment: Unsupervised multi-modal alignment typically relies on specific metrics. Wasserstein distance (Peyré et al., 2019), a.k.a EMD-based optimal transport (OT) algorithms have been widely adopted to various domain alignment tasks, including sequence-to-sequence learning (Chen et al., 2019), few-shot learning (Zhang et al., 2020), knowledge distillation (Balaji et al., 2019), unsupervised domain adaptation (Balaji et al., 2019), generative networks (Han et al., 2015; Genevay et al., 2018; Mroueh et al., 2018; 2019), and multi-modal learning (Yuan et al., 2020; Chen et al., 2020d; Kim et al., 2021; Li et al., 2022d; Pramanick et al., 2022). Previous VLP methods (Chen et al., 2020d; Kim et al., 2021), which use OT-based patch-word alignment, only utilize the Wasserstein distance. However, we argue that jointly modeling GWD (Peyré et al., 2016) and WD results in a superior multi-modal alignment for intricate images. To the best of our knowledge, this is the first work to apply WD and GWD-based optimal transport for feature-level alignment in VLP.

Proposed System - VoLTA

In this section, we present our proposed approach, VoLTA, which contains three broad modules - (i)(i) intra- and inter-modality redundancy reduction, (ii)(ii) weekly-supervised cross-modal alignment of local features, and (iii)(iii) cross-modal attention fusion (CMAF). Next, we introduce the fine-tuning strategies for various uni- and multi-modal downstream tasks as supported by VoLTA. An overview of the different modules of VoLTA is presented in Figure 2.

We use Barlow Twins (BT) (Zbontar et al., 2021), a non-contrastive covariance regularization as the foundational objective of VoLTA. The recent success of contrastive vision-language pre-training (Radford et al., 2021; Li et al., 2021b; Jia et al., 2021; Kim et al., 2021; Yang et al., 2022a; Dou et al., 2022a; b) has already shown that, compared to a single modality, image-caption pairs offer a significantly higher-level of abstractive and semantic concepts about the training samples. However, common contrastive VLP objectives, like InfoNCE (Oord et al., 2018), are data-hungry, as they require large batch sizes and well-mined hard negatives. On the other hand, the BT objective operates on the dimensions of the embeddings across the two views of training samples. Hence, it is more robust to batch size and can be trained using lower memory resources. In this work, we extend the BT objective for a multi-modal setup.

The original BT algorithm, which operates on joint embeddings of distorted samples, was proposed only for image modality. Specifically, for each image of a batch X\mathcal{X}, two distorted views are obtained using a distribution of data augmentation T\mathcal{T} with disparate probabilities. These distorted images are then fed into a shared image encoder containing a feature extraction network (e.g., ResNet (He et al., 2016)) cascaded with trainable linear projection layers, producing a batch of parallel embeddings zAz^{A} and zBz^{B}. The BT loss computed using the encoded embeddings can be denoted as:

λ\lambda is a positive weighting factor; CC is the cross-correlation matrix computed between zAz^{A} and zBz^{B} along the batch dimension; bb stands for sample indices in a batch; i,ji,j refers to the dimension indices of zAz^{A} and zBz^{B}. The first term in Equation 1 is the invariance term which attempts to equate the diagonal elements of the cross-correlation CC matrix to 11, whereas the second term is the redundancy reduction term which pushes the off-diagonal elements of CC matrix to .

In this work, we use BT for image-caption pairs. Specifically, we use stochastic data augmentations for both images and textAugmentation details are provided in Appendix D.1., and directly apply the BT objective for all the 2×22\times 2 pairs, resulting in additional supervision. Note this simple, straightforward, and instinctive extension enables us to apply redundancy reduction in between and across modalities, which intuitively results in superior visual representation. Moreover, in this bi-modal setting, we can pre-train a text encoder in parallel with the image encoder and, thus, can generalize our system to a broader range of uni- and multi-modal downstream applications.

Intra-modal Objective: Intra-modal objective refers to applying the BT loss in-between pairs of image and text embeddings. Given an image-caption pair, we first have two augmented views (I,I′)(I,I^{\prime}) for each image, and two augmented views (T,T′)(T,T^{\prime}) for each text. Then, we resort to Equation 1 individually for the image and text pairs.

Inter-modal Objective: Inter-modal objective refers to applying the BT loss across image and text embeddings. Since the image and text encoders can output features with different shapes, we design the projector layers with same output dimension. Hence, in addition to the original BT loss between (I,I′)(I,I^{\prime}) in Zbontar et al. (2021), we get three more loss terms - (T,T′)(T,T^{\prime}), (I,T′)(I,T^{\prime}), (I′,T)(I^{\prime},T), leading to 3×3\times diverse and high-quality additional supervision. The inter-modal BT losses can be directly computed following Equation 1.

2 Alignment of Local Features

Though the inter-modal redundancy reduction provides high-quality semantic supervision, it is computed on the global image- and text features and, thus, only simulates implicit and non-interpretable multi-modal alignment. However, fine-grained region-level downstream applications like detection, segmentation, and reference expression comprehension require local visual feature descriptors with specific spatial information. To achieve this, most existing top-performing VLP methods, including UniTAB (Yang et al., 2022c), OFA (Wang et al., 2022b), GLIP (Li et al., 2022c), and FIBER (Dou et al., 2022a), use high-resolution image-text-box data for fine-grained pre-training. However, bounding box annotations are expensive to collect and use for supervision. Hence, we seek an alternate weekly-supervised solution for local feature-level alignment using global image-caption annotations.

Importance of GOT in Patch-Word Alignment: As mentioned previously, GOT adopts two types of OT distances - WD for node matching and GWD for edge matching. In contrast, previous vision-language pre-training algorithms using OT for patch-word alignment only considered WD (Chen et al., 2020d; Kim et al., 2021). However, we argue that intricate images with multiple objects with similar shapes and colors require both WD and GWD for accurate, fine-grained matching. For example, in Figure 3, there are multiple “orange" present in the image. WD can only match nodes in the graph, and will treat all “orange" entities as identical and will ignore neighboring relations like “on the laptop". However, by using proper edge matching with GWD, we can preserve the graph’s topological structure. We can correctly identify which “orange" in the image the sentence is referring to. Hence, we couple WD and GWD mutually beneficially and use a joint transport plan for accurate patch-word matching.

Once Gx\mathcal{G}_{x} and Gy\mathcal{G}_{y} are computed, we follow Chen et al. (2020a) to compute WD and GWD.

Gromov-Wasserstein Distance assists in edge matching and preserves graph topology by calculating distances between pairs of nodes in each domain and measuring how these distances compare to the counter domain. In the same discrete graph matching setting, GWD between ϕ\phi and ψ\psi can be mathematically represented as:

where intra-graph structural similarity between two node pairs (xi,xi′)(x_{i},x_{i}^{\prime}) and (yj,yj′)(y_{j},y_{j}^{\prime}) is represented as L(xi,yj,xi′,yj′)=∥c1(xi,xi′)−c2(yi,yi′)∥L(x_{i},y_{j},x_{i}^{\prime},y_{j}^{\prime})=\|c_{1}(x_{i},x_{i}^{\prime})-c_{2}(y_{i},y_{i}^{\prime})\|, cic_{i} being cosine similarity between a node pair in any graph Gi\mathcal{G}_{i}. Transport plan T^\hat{\mathbf{T}} is periodically updated to align the edges in different graphs belonging to disparate modalities.

We further follow Chen et al. (2020a) to combine WD and GWD transport plans, leading to a unified GOT objective given as:

where γ\gamma regulates the importance of two loss terms.

3 Cross-Modal Attention Fusion (CMAF)

BT and GOT losses are computed in a dual encoder setting, which does not contain cross-modal interactions and is not suitable for complex multi-modal feature representation. Most existing methods, including UNITER (Chen et al., 2020d), ViLT (Kim et al., 2021), METER (Dou et al., 2022b), and GLIP (Li et al., 2022c) design cross-modal fusion by stacking additional transformer layers on top of uni-modal encoders, introducing a large number of added parameters during pre-training. We follow a more efficient solution proposed by FIBER (Dou et al., 2022a), which inserts cross-modal fusion into the uni-modal backbones with a gating mechanism. Specifically, at the top MM transformer layers in the vision and language backbone, cross-attention signals, weighted by a gating scalar α\alpha, are added to self-attention:

where α\alpha is a trainable parameter initialized to . Following existing literature (Li et al., 2021a; Wang et al., 2021a; Dou et al., 2022b; a), we use masked language modeling (MLM) and image-text matching (ITM) to pre-train the cross-attention parameters. For MLM, we randomly mask 15%15\% text tokens, and the loss aims to reconstruct the masked tokens. We feed the network with randomly sampled image-caption pairs for ITM, and the loss predicts whether they are matched. The gating mechanism is a good choice for CMAF because (i)(i) cross-attention parameters can easily be switched off by setting the gating scalar α\alpha to when computing the BT and GOT losses. Thus, we can learn the cross-attention parameters without affecting the original computational flow of uni-modal backbones. (ii)(ii) gating mechanism is more lightweight and memory-efficient than adding fusion-specific layers (GLIP and METER use 4×\times more fusion parameters than VoLTA).

Overall, VoLTA training pipeline can be summarized in the following three steps:

The overall VoLTA pipeline for computation of different training objectives is shown in Figure 2. The pseudo-code for VoLTA is presented in Appendix A.

4 Finetuning For Downstream Tasks

We adopt VoLTA to various vision- and vision-language downstream tasks. We switch off the inserted cross-attention modules for the vision-only tasks and use the image encoder. We utilize the learned cross-attention parameters as required for the vision-language tasks, following Dou et al. (2022a). For example, VQA and visual reasoning employ all cross-attention modules, whereas captioning requires only image-to-text cross-attention. Again, during IRTR, we switch off all cross-attentions and use VoLTA in a dual encoder setting. We keep all cross-attention parameters during multi-modal object detection and referring expression comprehension and train an object detection head from scratch using the language-aware image features.

Experiments, Results, and Analysis

Following Chen et al. (2020d) and Huang et al. (2021), we perform pre-training by appending the VG dataset (Krishna et al., 2017) with COCO20172017 (Lin et al., 2014), together consisting of 231231k images. We divide our downstream tasks into three categories - (i)(i) Uni-modal tasks such as image classification on ImageNet (Deng et al., 2009), VOC0707 (Everingham et al., 2010), COCO; object detection on VOC07+1207+12, COCO, and instance segmentation on COCO. (ii)(ii) Multi-modal fine-grained tasks such as region-level VL tasks - referring expression comprehension (REC) on RefCOCO, RefCOCO++, RefCOCOg (Kazemzadeh et al., 2014; Yu et al., 2016), and language-conditioned object detection on COCO and LVIS (Gupta et al., 2019). (iii)(iii) Multi-modal coarse-grained tasks such as image-level VL tasks - visual question answering on VQAv22 (Antol et al., 2015), visual reasoning on NLVR2 (Suhr et al., 2019), image- and text retrieval on Flicker3030k (Plummer et al., 2015) and captioning on COCO. We exclude any overlap between our pre-training and downstream validation/test splits. Detailed statistics of all downstream datasets are given in Appendix C.

2 Network Architectures

Following FIBER (Dou et al., 2022a), we adopt Swin-Base (Liu et al., 2021) and RoBERTa-Base (Liu et al., 2019) as our vision and text encoders, which are initialized with weights from uni-modal pre-training. We collect patch- and token features from the last transformer layers, feed them into the local projector network, and compute GOT loss. Furthermore, we apply AvgPool on patch and token features, feed them into the global projector network, and compute BT loss. Both local and global projector networks have three linear layers with dimensions 20482048-20482048-10241024, with batch normalization and ReLU after the first two layers. Section 4.7 gives an ablation on projector dimension. We use the image and text features after the AvgPool layer during downstream tasks. For CMAF, we insert the cross-attention into the top 6 blocks of the vision and text encoders. Moreover, for direct comparison with existing uni-modal baselines, we re-train VoLTA with ResNet5050 (He et al., 2016) and Swin-Tiny image encoders.

3 Implementation Details

4 Results on Vision-only tasks

We first experiment on three uni-modal tasks - classification, object detection, and instance segmentation. For a direct comparison with existing ResNet5050 and Swin-T baselines, we re-train identical encoders with VoLTA pipeline. Furthermore, since the uni-modal tasks do not utilize cross-attention parameters, we perform an ablation by dropping the MLM and ITM objectives from VoLTA.

Image Classification: Table 1 presents the linear probing results of uni-label classification on ImageNet and multi-label classification on VOC0707 and COCO. For all uni-modal tasks, we report results with COCO pre-training for a fair comparison with existing baselines. For ImageNet, we adopt all COCO baselines from Yuan et al. (2021). Even without the MLM and ITM objectives, VoLTA achieves better performance than all baselines across three datasets with ResNet5050 backbone. The Swin backbones and cross-attention module further improve the performance. For VOC0707, we report the results for both SVM and MLP-based linear classifiers. VoLTA with ResNet5050 backbone achieves state-of-the-art results on VOC0707 SVM evaluation, beating the nearest baseline, SwAV, by 0.70.7 mAP score. These results indicate the ability of VoLTA to learn effective image-level visual features.

Object Detection & Instance Segmentation: Next, we perform two uni-modal region-level tasks - object detection on VOC07+1207+12 and COCO20172017, and instance segmentation on COCO20172017. As shown in Table 2, VoLTA yields the state-of-the-art performance in both tasks across the majority of metrics. The fine-grained region-level understanding helps VoLTA to perform well on detection and segmentation tasks.

5 Results on Fine-grained Vision-Language tasks

Next, we perform region-level multi-modal downstream tasks - referring expression comprehension (REC) and language-guided object detection.

REC: This task aims to localize target objects in an image described by a referring expression phrased in natural language and, thus, perfectly evaluates the fine-grained feature representation capability of VoLTA. As depicted in Table 3, VoLTA beats larger-sized UNITER-L and VILLA-L models on the challenging testB split of both RefCOCO and RefCOCO++. Moreover, VoLTA performs comparably with MDETR and UniTAB, even without being trained on grounding data. These results indicate our model’s efficacy in learning fine-grained local visual features.

Object Detection: We evaluate VoLTA on two challenging language-conditioned object detection benchmarks - COCO and LVIS. Note that, all existing baselines for this tasks are pre-trained on fine-grained image-text-box data, whereas VoLTA only utilizes image-caption pairs. Table 4 shows that VoLTA performs comparatively with these strong baselines. Note that VoLTA beats Mask R-CNN, MDETR, and GLIP-B on LVIS APr, which denotes average precision on rare objects. Thus, we conclude that VoLTA achieves impressive localization ability and robustness, even without utilizing any grounding annotations.

6 Results on Coarse-grained Vision-Language tasks

Next, we perform image-level multi-modal downstream tasks - visual question answering (VQA), visual reasoning, retrieval, and captioning.

VQA & Visual Reasoning: As reported in Table 5, VoLTA achieves the best performance on VQA and visual reasoning across the baselines pre-trained with a comparable amount of data. Moreover, on VQA, VoLTA beats LXMERT, which is trained with 2×2\times more data. These results demonstrate the efficacy of our method even when utilizing a mid-scale pre-training corpus.

Retrieval: Most existing VLP methods use a fusion encoder for image and text retrieval and feed every image-text pair into the model. Though such fine-tuning often results in higher performance, it introduces quadratic time cost and is not scalable. Following Dou et al. (2022a), we adopt a more efficient strategy. We drop the cross-attention parameters for this task and compute the dot product of image and text features extracted separately in the dual-encoder setting. As shown in Table 5, even with such an approach, VoLTA produces superior performance among the baselines trained with a similar amount of data, beating all three baselines by a significant margin.

Captioning: We perform captioning on the COCO dataset to evaluate if VoLTA can adopt a generation task. We integrate GOLD (Pang & He, 2021) into VoLTA during fine-tuning as it produces significant improvements. As shown in Table 5, our approach maintains superior captioning performance across all baselines pre-trained with comparable data. Using CIDEr optimization further improves performance.

It is worth mentioning that besides achieving a superior result than all baselines using a comparable amount of data on multi-modal coarse-grained tasks, VoLTA also outperforms multiple methods pre-trained using magnitude more data. These results, shown in Table F.1, indicate the effectiveness and generalizability of VoLTA across these tasks.

7 Ablation Study

We perform ablation studies on the pre-training objectives, GOT loss weight, and the dimension of projectors.

We also verify the effectiveness of the multi-modal BT objective by ablating the intra- and inter-modal terms. The first row of Table 7(a) is identical to the original image-only BT objective. Next, we introduce the text branch and add the same BT objective between the two views of the caption. Afterward, we add the inter-modal BT objectives. As shown in Table 7(a), each loss term improved the image classification performance, demonstrating the importance of intra- and inter-modal objectives. Overall, this set of experiments demonstrates that all objectives are necessary for our model to perform well on different fine-grained multi-modal tasks.

Ablation on Projector Dimension: The design of the projector head plays a pivotal role in the downstream performance of the model (Garrido et al., 2022). To investigate the impact of hidden and feature (projector output) dimensions, we have tested 44 different configurations on uni-modal downstream classification tasks. It can be observed (see Table 7(c)) that an increase in the number of parameters in the projector head does not necessarily lead to an increase in performance. For example, a projector configuration of 81928192-81928192-256256 has roughly eight times more parameters than 20482048-20482048-10241024. However, the latter performs better in downstream tasks (Table 7(c)), indicating that the output dimension of the projector plays a crucial role in the final performance of the model.

8 Qualitative Results & Error Analysis

Figure 4 shows the fine-grained alignment of image regions and caption words achieved by the pre-trained VoLTA system. The transport plan from GOT module outputs the similarity across every image patch and caption token. To obtain the visualizations in Figure 4, we choose the similarity scores between the red words in the caption with every image patch. Next, we apply bilinear interpolation to these similarity scores to convert them to the same dimension as the input image. Finally, we superimpose these interpolated similarity maps on the input images to obtain Figure 4 as the outcome. In most cases, the pre-trained model accurately learns to localize various objects using only global image-caption data. However, objects in extremely cluttered scenarios are occasionally not focused. We show such error cases in Section E.

Conclusion

We present VoLTA, a unified VLP paradigm that utilizes image-caption data but achieves fine-grained region-level image understanding, eliminating the use of expensive box annotations. VoLTA adopts graph optimal transport-based weakly supervised patch-token alignment and produces an explicit, self-normalized, and interpretable low-level matching criterion. Extensive experiments demonstrate the effectiveness of VoLTA on a wide range of coarse- and fine-grained tasks.

Acknowledgement

The codebase for this work is built on the Barlow Twins (Zbontar et al., 2021), GOT (Chen et al., 2020a), and FIBER (Dou et al., 2022a) repository. We would like to thank the respective authors for their contribution, and the Meta AI team for discussions and feedback. Shraman Pramanick and Rama Chellappa were partially supported by an ONR MURI Grant N00014-20-1-2787.

References

Appendix A Pesudo Code of VoLTA

The training pseudo code for VoLTA is as follows:

Appendix B Overview of Vision-Language Pre-training Models

Vision-Language Pre-trained (VLP) models have proven extremely beneficial for multi-modal tasks in recent years. Earlier works were predominantly focused on using pre-trained object detectors to extract patch (region) level information from corresponding images (Lu et al., 2019; Li et al., 2020a; Tan & Bansal, 2019; Chen et al., 2020d; Su et al., 2019). In some of these models, such as ViLBERT (Lu et al., 2019), and LXMERT (Tan & Bansal, 2019), multi-modality fusion has been achieved via co-attention using a third transformer which contains fused information independently obtained from respective vision and language encoders. On the contrary, VisualBERT (Li et al., 2020a), VL-BERT (Su et al., 2019), and UNITER (Chen et al., 2020d) employ a merged attention strategy to fuse both image patches and text features together into a unified transformer through corresponding image and text embedders. In addition to these, OSCAR (Li et al., 2020b) uses object tags as inputs. VinVL (Zhang et al., 2021) follows a similar strategy to that of OSCAR, the only difference being their novel 3-way contrastive loss which optimizes the training objectives used for VQA and text-image matching. VL-T5 (Cho et al., 2021) exploits bounding-box coordinate information, image IDs, and region IDs along with ROI features for visual embedding. Here, encoded visual and textual features are fed into a bi-directional multi-modal encoder and an auto-regressive text decoder framework, respectively, for pre-training.

In all the above methods, pre-trained object detectors are kept frozen during the training. Furthermore, extracting region-level features from images can be tedious. To address these shortcomings, end-to-end pre-training methods have been developed. PixelBERT (Huang et al., 2020) uses a CNN-based vision encoder and sentence encoder to obtain image and text representations, respectively. These representations are subsequently fed into a transformer via a cross-modality alignment. SOHO (Huang et al., 2021) uses grid features-based discretization via a learned vision dictionary which is then fed into a cross-modal module. SimVLM (Wang et al., 2021b) uses CNN and text token embedding for image and text feature representation extraction with a unified encoder-decoder transformer trained on a PrefixLM objective. Finally, MDETR (Kamath et al., 2021) uses CNN and RoBERTa (along with corresponding projection layers) for image and text feature extraction. These extracted features are concatenated before passing through a unified transformer trained on 1.3M Image-Text-Box (I-T-B) annotated data.

In recent years, the rise of Vision Transformers (ViT) (Dosovitskiy et al., 2021) has motivated the research community to have an all-transformer framework by incorporating ViTs (instead of CNN backbones) in VLP models. Image patch features and text token embeddings are fed directly into a ViT model for pre-training in ViLT (Kim et al., 2021). Visual Parsing (Xue et al., 2021), ALBEF (Li et al., 2021a), and METER (Dou et al., 2022b) use ViTs as vision encoders for image feature generation. ALBEF and METER use co-attention in their pre-training frameworks for multimodality fusion.

Another class of VLP models in the form of CLIP (Radford et al., 2021), DeCLIP (Li et al., 2021b), and ALIGN (Jia et al., 2021) has been introduced lately. Although known for their impressive zero-shot recognition ability and excellent transferability to downstream tasks, these models typically rely on huge amounts of image-text pairs for pre-training. Contrastive loss forms the core component of the pre-training objectives in these VLP models. In such models (e.g., CLIP (Radford et al., 2021), DeCLIP (Li et al., 2021b)), separate encoders have been used for each modality. On the contrary, modality-shared contrastive language-image pre-training (MS-CLIP) (You et al., 2022) leverages knowledge distribution across multiple modalities (image and text) through parameter sharing. In their unified framework, the parameters which are being shared between two modalities include the attention and feedforward modules and the layerNorm layers.

GLIP (Li et al., 2022c) and GLIPv2 (Zhang et al., 2022) use a localization loss along with a word-region alignment loss for pre-training corresponding encoders using image-text-box annotations. BLIP (Li et al., 2022b) employs image and text encoders connected through a cross-modality multi-head attention which are pre-trained on image-text pairs using contrastive and language modeling objectives. OmniVL (Wang et al., 2022a) utilizes a unified image (and video) encoder and a text encoder pre-trained on image-text, image-label, video-text, and video-label pairs using unified vision-language contrastive, vision-language matching and language modeling losses. Furthermore, a visual-grounded alignment decoder is also present for facilitating better learning and alignment between various modalities. X-VLM (Zeng et al., 2022) employs a vision transformer to extract features from the subset of patches representing images/regions/objects. These patch features are then paired with associated text features for contrastive learning, matching, and masked language modeling. Additionally, image and text pairings are also done for bounding-box prediction which is used to locate visual concepts in the image. CMAL (Ma et al., 2022) proposes interactions between features (obtained from respective image and text encoders) via cross-modal associative mappings which help in fine-grained semantic alignment between the learned representations. LOUPE (Li et al., 2022a) implements token-level and semantics-level Shapley interaction modeling with global image-text contrastive loss (in a dual-encoder setting) for explicit learning of fine-grained semantic alignment between visual regions and textual phrases without using expensive bounding-box annotations. FILIP (Yao et al., 2022) removes the need for cross-modality attention fusion by modeling the fine-grained semantic alignment between visual and textual tokens via a novel cross-modal late interaction mechanism in contrastive loss. TCL (Yang et al., 2022b) uses global cross-modal alignment, intra-modal alignment, and local mutual information maximization losses along with masked language modeling and image-text matching to learn robust image-text representations during pre-training. UniCL (Yang et al., 2022a) utilizes a unified learning method with a two-way contrastive loss (image-to-text and text-to-image) in the image-text-label space which can learn representations from either of the image-label and image-text data or both. UniTAB (Yang et al., 2022c) employs a transformer-based encoder-decoder framework that can jointly output open-ended text and box, encouraging alignment between words and boxes.

In order to accelerate the convergence of VL pretraining, Wang et al. (2023c) proposed free language modeling (FLM) which addresses the issues inherent to masked language modeling (MLM) and autoregressive modeling. Using FLM as a pre-training objective, the authors have achieved impressive performance on several downstream tasks. Fame-ViL (Han et al., 2023) introduces a parameter-efficient VLP approach employing a task-versatile architecture with cross-attention and task-specific adapters. Fame-ViL applies a single model for various heterogeneous fashion tasks achieving performance gains over previous SOTA benchmarks. Wang et al. (2023b) have introduced a simple yet effective position-guided text prompt (PTP) paradigm to improve the visual grounding capability of existing cross-modal VL architectures and help them better handle various downstream tasks. Park & Han (2023) have proposed a VL framework based on the explainable soft feature masking and regularization via diversification strategies for improving the performance of VL models in several downstream tasks. Li et al. (2023a) have devised BLIP-2, where a lightweight querying transformer is pre-trained using a two-stage strategy to bridge the modality gap. A frozen encoder is used in the first stage to bootstrap VL representation learning, and vision-to-language generative learning is bootstrapped in the second stage employing a frozen LLM, allowing zero-shot generation capabilities. Jang et al. (2023) have developed a simple and unified VL model (as a single tower) in a modality-agnostic manner. BEiT-3 (Wang et al., 2023d) introduces a general-purpose multimodal foundation model to pre-train a multiway transformer by performing masked data modeling on inputs irrespective of modalities (i.e., images, texts, and image-text pairs). FLIP (Li et al., 2023b) extends CLIP (Radford et al., 2021) by performing contrastive learning on pairs of masked image patches and corresponding texts without reconstructing the masked image content. GIT (Wang et al., 2023a) unifies the VL architecture (an image encoder and a text decoder) under a single language modeling task while also scaling up the pre-training data and the model size to gain superior performance on captioning and question-answering downstream tasks.

FIBER (Dou et al., 2022a) fuses vision and language encoder backbones through merged co-attention which are then pre-trained on 4M data with two-stage pre-training (coarse- and fine-grained). Image-text pairs are used in the coarse-grained pre-training stage which is then followed by a fine-grained pre-training stage with image-text-box annotations. However, these bounding box annotations come with extra overheads. Therefore, in our model, VoLTA, we propose an alternate solution for optimal-transport based local feature-level alignment using global image-caption annotations which performs well not only on coarse-grained tasks (such as VQA and Image Captioning), but also on fine-grained tasks (such as Referring Expression Comprehension and Object Detection). Table B.1 encapsulates an overview of the details of all these aforementioned methods.

Appendix C Downstream Datasets

Our downstream tasks can be categorized into three groups: uni-modal, multi-modal coarse-grained, and multi-modal fine-grained.

Uni-modal: For uni-modal tasks, we fine-tune (and validate) our pre-trained model on ImageNet-1k (Deng et al., 2009) for image classification, VOC07+12 (Everingham et al., 2010) for image classification and object detection, and COCO (Lin et al., 2014) for image classification, object detection, and instance segmentation.

Multi-modal Coarse-grained: Here, we fine-tune (and validate) our pre-trained model on VQAv2 (Antol et al., 2015) for visual question answering, NLVR2 (Suhr et al., 2019) for visual reasoning, Flickr30k (Plummer et al., 2015) for image and text retrieval, and COCO (Lin et al., 2014) for image captioning.

Multi-modal Fine-grained: For these tasks, we fine-tune (and validate) our pre-trained model on RefCOCO, RefCOCO+, and RefCOCOg (Kazemzadeh et al., 2014; Yu et al., 2016) for referring expression comprehension, and COCO (Lin et al., 2014) and LVIS Mini (Gupta et al., 2019) for language-conditioned object detection.

Several multi-modal downstream tasks are built based on the COCO dataset, where the validation and test splits of these downstream tasks are scattered across the raw COCO splits. Therefore, during pre-training, we carefully selected the portion of the COCO dataset which does not overlap with the validation/test splits of these multi-modal downstream tasks.

Appendix D Implementation Details &\& Hyper-parameter Values

We use ResNet5050/Swin-T/Swin-B (He et al., 2016; Liu et al., 2021) as image encoder and RoBERTa (Liu et al., 2019) as text encoder. Each encoder is followed by a projector network which is a 33-layer MLP with the configuration [dd-20482048-20482048-10241024]. Here, dd represents the embedding dimension of the encoder’s output.

Image Augmentations: Two sets of random transformations sampled from an augmentation pool are applied on each input image to generate two disparate distorted views. The augmentation policy is composed of RandomResizedCrop, RandomHorizontalFlip, ColorJitter, RandomGrayscale, GaussianBlur, and Solarization augmentations. RandomResizedCrop is applied with a probability of 1.0, whereas the remaining ones are applied randomly with varying probabilities following Zbontar et al. (2021) (see Table D.1).

Text Augmentations: Two sets of random transformations are applied on input text using EDA (Wei & Zou, 2019) including synonym replacement, random insertion, random swap, and random deletion with different probabilities as outlined in Table D.1.

D.2 Pre-training Setup

Table D.2 shows the details of hyper-parameters used during training.

VoLTA comprises a vision encoder and a language encoder with a merged co-attention for cross-modality fusion. In our experiments, we have considered two types of vision encoder backbones - ResNet-5050 (He et al., 2016) and Swin Transformer (Liu et al., 2021). For fair comparisons with related works (Dou et al., 2022b; a), the input image resolution for ResNet-50 encoder backbone is kept as 224 ×\times 224, whereas for Swin-B, it is 384 ×\times 384. The output embedding dimension of the image encoder in both cases is 1024. Similarly, to be consistent with Dou et al. (2022b; a), we have selected RoBERTa as the language encoder with a vocabulary size of 50265, a tokenizer as ‘roberta-base’, a maximum input text length of 30, and an output embedding dimension of 768 (please refer to Table D.2 for more details).

Separate projector heads follow vision and language encoders. A projector head consists of 3 linear layers, each with 2048 output units (except for the last one, which has 1024 output units), followed by a Batch Normalization layer and ReLU activation (for exact configuration, please refer to Table D.2). The final projected output denotes the input (image/text) feature representation used in downstream tasks. The embeddings (i.e., output from respective encoders) are fed into the loss function of VoLTA to learn these representations.

The loss function of VoLTA includes four different loss components, namely, multi-modal Barlow Twins for intra- and inter-modality redundancy reduction, GOT for alignment of local features, and MLM and ITM together for encouraging cross-modal attention fusion. For MLM, we randomly mask 15%15\%Following BERT, we decompose this 15%15\% into 10%10\% random words, 10%10\% unchanged, and 80%80\% with a special token [MASK]. (MLM probability in Table D.2) of the input tokens, and the model is trained to reconstruct the original tokens. For ITM, the model predicts whether a given image-text pair is matched.

Pre-training Cost: Our Swin-B backbone takes 6 hours per epoch to train on 64 V100100 GPUs, with per GPU batch-size of 44.

D.3 Downstream Setup

For ImageNet, the linear classifier has been trained for 100 epochs with a batch size of 256, an LR of 0.3, and a cosine LR schedule. Cross-entropy loss is minimized with SGDM optimizer (momentum of 0.9), and a weight decay of 1e-6. For both COCO and VOC, the linear classifier has been trained for 100 epochs with AdamW optimizer with batch size of 256, an LR of 5e-2, and a weight decay of 1e-6.

Object Detection: For training the detection model, the detectron2 library (Wu et al., 2019) has been used. The backbone networks for Faster R-CNN (Ren et al., 2015) and Mask R-CNN (He et al., 2017) has been initialized using our pre-trained model.

For VOC07+12, we used the trainval set comprising 16K images for training a Faster R-CNN (Ren et al., 2015) C-4 backbone for 24K iterations using a batch size of 16 across 8 GPUs (using SyncBatchNorm). The initial learning rate for the model is 0.15, which is reduced by a factor of 10 after 18K and 22K iterations. Linear warmup (Goyal et al., 2017) is used with a slope of 0.333 for 1000 iterations.

For COCO, Mask R-CNN (He et al., 2017) with a C-4 backbone on the COCO 2017 train split is used for training, and the results are reported on the val split. A learning rate of 0.03 is used, and other parameters are kept the same as in the 1×\times schedule in detectron2 (Wu et al., 2019).

D.3.2 Coarse-grained multi-modal downstream tasks

Vision-Language Classification (VQAv2 and NLVR2): Vision-Language Classification task encompasses VQAv2 and NLVR2, whose hyper-parameter setup has been taken from METER (Dou et al., 2022b) and FIBER (Dou et al., 2022a). Model finetuning is done with peak learning rates of 2e-5 for the backbones, 1e-4 for the cross-modal parameters, and 1e-3 for the head layer for 10 epochs with a batch size of 512. The image resolutions are set to 576 for VQAv2 and 384 for NLVR2 and the models are evaluated with the VQA-Scores for VQAv2 and accuracy for NLVR2 (Table C.1).

Image-Text Retrieval (IRTR): We follow Dou et al. (2022a) for IR-TR setup for the Flickr30k dataset, where the cross-attention layers in the backbones are removed during IR-TR fine-tuning and evaluation. The peak learning rates are set to 2e-5 for the backbones, and 1e-4 for the head layer. Furthermore, a batch size of 1024 is considered, and each image resolution is set to 576. We evaluate the Recall@1 metric for both the text and image retrieval tasks as outlined in Table C.1.

Image Captioning: For image captioning, only the image-to-text attentions are kept for the cross-modality attention fusion, and the model is converted into a standard seq2seq model (Dou et al., 2022a). We used a causal mask on the decoding side, and the outputs are predicted auto-regressively (Dou et al., 2022a). Models are trained with the cross-entropy loss for 5 epochs with the peak learning rates of 5e-5 for the backbones, and 2.5e-4 for the rest of the parameters, followed by a two-stage finetuning. In the first stage, finetuning with GOLD (Pang & He, 2021) is done for 5 epochs with a peak learning rate of 1e-5 for the backbones, since it is efficient and has been proven to be effective when the model input can correspond to different outputs. The second stage of fine-tuning involves CIDEr optimization where the learning rate is further reduced to 1e-6, and the model is trained for 3 epochs. A batch size of 512 is considered in both these cases, and a beam size of 5 is used during inference. Evaluation metrics include BLEU (Papineni et al., 2002), METEOR (Banerjee & Lavie, 2005), CIDEr (Vedantam et al., 2015), and SPICE (Anderson et al., 2016) scores (shown in Table C.1).

D.3.3 Fine-grained multi-modal downstream tasks

Referring Expression Comprehension (REC): We follow Dou et al. (2022a) for training and evaluation on 3 different datasets (RefCOCO, RefCOCO+, and RefCOCO) where the models are finetuned with a batch size of 16 for 20 epochs. A warmup of 2000 steps with a peak LR of 1e-5 is used for the OD head as well as the rest of the model’s parameters. LR drops twice, once at 67% and the other at 89% of the total number of steps. Horizontal flip augmentation has been turned off during REC training because it was observed by Dou et al. (2022a) that horizontal flip adversely affected the performance, particularly on the RefCOCO dataset. Accuracy is used as the evaluation metric in this case (Table C.1).

Object Detection: We follow the training and evaluation setup of Dou et al. (2022a) for text-conditioned (multi-modal) object detection. For both COCO and LVIS datasets, the model has been finetuned for 24 epochs with a batch size of 32, an LR of 1e-5, and two learning rate drops, once at 67% and the other at 89% of the total number of steps. AP scores are used in this case for model evaluation (Table C.1).

Appendix E Error Analysis

Although VoLTA learns impressive fine-grained region-level understanding during pre-training, there are still some cases where the model fails to identify tiny and hindered objects, especially in cluttered environments. We show four such examples in Figure E.1. In the first image, the object ‘red peppers’ is barely visible even in human eyes, and thus, VoLTA can not precisely identify these objects. However, it can identify the coarse region (the fruit basket) where ‘red peppers’ can be present. In the second image, VoLTA confuses a ‘dishwasher’ with a ‘microwave oven,’ probably because the ‘microwave oven’ is present in a cluttered environment, and both objects have similar appearances in low-resolution frames. In the third image, VoLTA can correctly identify ‘red apples’, but fails to spot ‘green apples’, probably because VoLTA has not seen enough such samples. In the last image, the face of the ‘person’ is hindered by the camera, and VoLTA fails to locate it. Since we pre-train VoLTA with 224×224224\times 224 images, such tiny objects are often hard to be distinguished. However, higher-resolution images will be more helpful in addressing such intricate scenarios, which we plan to explore in future works.

Table F.1 presents a comparison of VoLTA on the multi-modal coarse-grained tasks with state-of-the-art methods pre-trained using magnitude more data. On VQA, VoLTA beats ViLBERT, UNITER-B, VILLA-B, UNIMO-B, and ViLT-B, each pre-trained on 3−43-4M datasets. Please note that VoLTA is trained only on COCO and VG, whereas the other methods use a combination of COCO, VG, CC, and SBU datasets. Such strong performance proves the generalizability of VoLTA. On captioning, VoLTA beats Unified VLP, OSCAR, UFO-B, ViTCAP, VinVL-B, METER-CLIP-B, and XGPT. However, for IRTR and NLVR VoLTA can not yield better performance over these baselines. We assume that the large domain difference between pre-training and downstream datasets is the reason behind the limited performance on IRTR and NLVR.

Appendix G Additional Qualitative Results

Visual Question Answering and Visual Reasoning: Visual question answering (VQA) is a widely recognized multi-modal task that infers an answer in response to a text-based question about an image. In Figure G.1, we demonstrated several examples image-question pairs and corresponding answers predicted by VoLTA on the VQAv2 validation set. The primary aim of the visual reasoning task is to ascertain the veracity of a natural language statement against an associated image pair. Figure G.2 displays examples of responses (True/False) predicted by VoLTA on the NLVR2 validation set.

Language-conditioned Object Detection: Object detection forms an indispensable constituent of several multi-modal understanding systems. However, the conventional object detection pipeline is employed as a black-box tool and predicts all possible objects in the image. On the other hand, for better apprehension of combinations of these objects in free-form texts, a language-conditioned object detection task is considered (Kamath et al., 2021; Dou et al., 2022a). We use pre-trained VoLTA and fine-tuned and evaluated COCO and LVIS datasets for the text-conditioned object detection task. As illustrated in Figure G.3, VoLTA predicts bounding boxes relevant to the text prompts (captions) and labels them with the corresponding spans from the text. For example, the top-middle image has 4 objects. However, based on the text prompt, our model predicts boxes only for person and cup.

Referring Expression Comprehension (REC): The objective of REC is to align the entire referring expression (text) with the corresponding box by disambiguating among the several occurrences of an object belonging to the same category and therefore, one box per expression is to be predicted. For example, the bottom-left image in Figure G.4 depicts VoLTA’s box prediction for the corresponding referring expression: the slice of cake on the left.

Comparison of CLIP vs. VoLTA on Referring Expression Comprehension (REC): Figure G.5 shows a comparative qualitative evaluation between frozen CLIP + dynamic head (Dai et al., 2021), and VoLTA. We concatenate the vision and text features from the CLIP encoders and train a dynamic head on top of frozen features for the REC task. Since CLIP is a dual-encoder system pre-trained with image-level features, it can not learn superior fine-grained features. Hence, CLIP fails on harder REC samples. For example, if multiple similar-looking objects (benches or bowls) exist in an image, CLIP fails to distinguish between them. However, VoLTA succeeds on such complex samples, which can be attributed to the fine-grained alignment achieved by the GOT objective.