Learning Customized Visual Models with Retrieval-Augmented Knowledge
Haotian Liu, Kilho Son, Jianwei Yang, Ce Liu, Jianfeng Gao, Yong Jae Lee, Chunyuan Li
Introduction
It has been a fundamental research problem in computer vision (CV) to build a transferable visual system that can easily adapt to a wide range of downstream tasks. With remarkable advances in deep learning, a de facto solution to achieve this is to train deep neural networks on a large amount of data to pursue the so-called generic visual representations. This dates back to the standard supervised training on ImageNet , whose superb representation power is further demonstrated in BiT /ViT by scaling up the training to JFT300M . Along the way, recent efforts have been applied to the popular image self-supervised learning to reduce the demand for labeled data. The third approach is image-text contrastive learning trained on billion-scale web-crawled image-text pairs. Such models, like CLIP and ALIGN , are able to achieve great performance on different downstream domains, without the need of any human labels.
Excellent empirical performance has been achieved with the above three pre-training methods, by following the well established two-stage pre-training then adaptation pipeline: model pre-training from scratch on large data, then model adaptation directly on downstream tasks. Specifically, the pre-trained models are adapted to downstream tasks by considering the available task-specific samples only: either evaluated in a zero-shot task transfer manner, or updated using linear probing (LP) , finetuning (FT) , or prompt tuning . Following this two-stage pipeline, most research has reverted to the faith that building transferable visual systems is equivalent to developing more generic visual models by feeding all knowledge in the model pre-training stage. Therefore, the community has been witnessing a trend in exploring scaling success of pre-training model and data size with less care on the target domain, hoping that the model can adapt to any downstream scenario.
In this paper, we argue that the conventional two-stage pipeline above is over-simplified and less efficient, in achieving the goal of building a transferable visual system in real-world settings. Instead, we propose a customization stage in between the pre-training and adaptation, where customization is implemented by systematically leveraging retrieved external knowledge. The inspiration comes from how humans are specialized in society for better generalization: instead of trying to memorize all concepts, humans are trained/prepared in a relevant subject to master a certain skill, while maintaining the basic skills in pre-training.
To this end, we explore a systematic approach to acquire and learn with external knowledge sources from a large image-text corpus for model customization. The process of collecting external image-text knowledge is fully automatic without extra human annotation. The acquired knowledge typically contains richer information about the concept: relevant images that never appear in the downstream training and evaluation set, and richer text descriptions about concept semantics. Such multi-modal knowledge sources are generally available on the web, and further open-sourced like LAION . They cover a variety of domains, making it possible to develop customized visual models for task-level transfer. Similar retrieval-augmented intuitions have been exploited in computer vision for class-level transfer , but not yet for task-level transfer (similar to that of CLIP). Our main findings/contributions can be summarized as follows.
We propose to explore the potential of the web-scale image-text corpus as external knowledge to significantly improve task-level transfer performance on the target domain at an affordable cost. A simple and effective strategy is proposed. To begin with, we build a large-scale multi-modal indexing system to retrieve the relevant image-text pairs using CLIP features and approximate nearest neighbor search. For a CV problem, the task instruction is often sufficiently specified with text such as class names, which allows us to utilize them as queries to retrieve the relevant image-text pair knowledge from the indexing system. No images from the CV problem are needed. To efficiently build the customized visual model, we propose a novel modularized learning strategy: only updating the additional trainable weights on the retrieved knowledge, and freezing the original model weights. Hence, the model masters the new skill without forgetting basic skills.
The generality and effectiveness of the proposed customization strategy is demonstrated on four CV problems. We instantiate it with CLIP, and develop the customized visual models for image classification on ImageNet and 20 datasets in Elevater , image-text retrieval on COCO /Flickr , as well as object detection and semantic segmentation on COCO . The knowledge bases are considered as LAION and larger web-crawled multi-modal data. The retrieval-augmented knowledge (3% image-text pairs compared with the original training data) significantly improves the model’s zero-shot performance without the need of accessing any images on downstream tasks. See Figure 1 for highlighted results. For example, our ViT-L/14 checkpoint achieves 78.5% zero-shot accuracy on ImageNet , surpassing all public checkpoints from CLIP and OpenCLIP , including those with larger model size and trained on a much larger LAION-2B . The new customized visual models also demonstrate higher few/full-shot performance than the original generic model counterparts.
Our retrieval system, codebase, and pre-trained models will be publicly available. To make this line of research more accessible, our retrieved subsets for both Elevater and ImageNet will also be made available, with an easy-to-use toolkit to download the subsets without storing the whole dataset locally. It poses a feasible direction for leveraging the ever-increasing data from the Internet for customized visual recognition, especially for the low-resource regimes.
Related Work
Learning transferable visual representations from natural language supervision is an emerging research area. The pioneering works of CLIP and ALIGN make use of contrastive learning to pretrain models on billion-scale web-crawled image-text pairs. There are an increasing number of studies to improve their generality from various modeling perspectives, including training objectives , scaling techniques , data efficiency , and leveraging multilingual correlations . In academia, several works demonstrate techniques to improve the learned semantic representations on datasets at a smaller scale (e.g. CC3M , CC12M , YFCC15M ), by exploring pretraining on a unified image-text-label space , token-level contrastive loss , and auxiliary within-modality contrastive loss . Complementary to the above works, we build on top of existing pre-trained generic models, and aim to improve the model’s performance by customizing them using retrieved relevant image-text pairs.
In natural language processing, several works augment large language models with external data encoded with structured language and relation representations . Motivated by retrieval-augmented models in NLP, several recent works leverage visual and / or textual knowledge to improve classification , question answering , image generation , and multi-modal tasks simultaneously . RAC improves long-tail classification by retrieving from a non-parametric memory consisting of pre-encoded images and text. K-LITE enhances the text prompts with the retrieved external knowledge that is encoded in natural language. Our paper leverages the paired knowledge of image-text and aims to improve task transfer performance for core vision problems such as classification, retrieval, detection and segmentation.
CLIP demonstrates impressive zero-shot and linear probing performance on different downstream domains. Several works explore improving the domain adaptation performance on CLIP models. Elevater leverages the text encoder outputs to initialize the task-specific linear head to improve the linear probe and finetuning performance of CLIP. Inspired by prompting techniques in NLP, recent works make use of learnable prompts that are trained on a few samples on downstream tasks. Similar to these works, this paper aims to improve CLIP’s performance on downstream tasks, while making use of relevant image-text pairs data to improve the model’s performance, without access to the downstream images. Furthermore, when downstream samples are available, they are complimentary to our method.
Retrieval-Augmented Customization
Computer vision models have achieved strong transfer performance, when learning with large-scale image data only , image-label data and/or image-caption data . Without loss of generality, we follow and define a unified triplet-wise format for image-text-label data, where is an image, is its language description, and is a label indicating the index of the unique language description in the dataset. In a general form, the language description is a text sequence . It ranges from simple category names representing visual concepts when is small, to more free-form and semantic-rich sentences such as captions when is relatively large.
Zero-shot. In a customized setting, the simplest task definition can be provided as a set of category names for visual recognition, leading to the task instruction . No training image is available, not to mention the corresponding label .
Few/Full-shot. The users may spend annotation cost to curate image-label pairs as the training instances, making the task instruction more specific, , which allows updating the image encoder model for better adaptation performance.
In this paper, we assume there exists a web-scale image-text corpus as the external knowledge source , where is the database size, e.g. 400M for LAION . One may use the task instruction as a query to seek additional relevant knowledge to build a more transferable visual system. Given the downstream task instruction and an external knowledge source , our goal is to learn customized visual-semantic representations, which are readily transferable to the downstream task of interest, whose training and evaluation images are not observed during the customization process. To this end, we propose React. We illustrate the high-level idea in Figure 2, and describe the process as follows.
2 Multi-modal External Knowledge
We explore web-scale image-text data as the multi-modal knowledge base in this paper. Ideally, one may consider the entire web as the knowledge base, and use Google or Bing search to retrieve the relevant knowledge. We consider two large static datasets with image-text pairs. To control the experiment complexity and ensure reproducibility, we use LAION-400M , a publicly available database with 400M pairs, for most of the experiments. To further study the scaling influence of the retrieval base, we conduct comparisons on Web-800M, a privately collected web database with 800M pairs.
To facilitate an efficient knowledge acquisition process, we use pre-trained contrastive models (e.g. CLIP) as the feature extractor, and build a cross-modal retrieval system using FAISS . We use its Hierarchical Navigable Small World (HNSW) approximate -NN lookup to balance performance and efficiency. For more details, please refer to the supplementary materials. After the retrieval system is built on the designated retrieval pool, it can be efficiently used for retrieving relevant image-text pairs for various downstream domains.
To facilitate the same interface for various customized visual tasks in the wild, it is desirable to have the same uniform task instruction schema. In NLP, all task instructions can follow the same uniform schema, composed of task definition and positive/negative examples . Here, the task definition defines a given task in natural language, completely specifying how an input is expected to be mapped to an output text. We note a coherence connection between this NLP task schema and the customized zero/few/full-shot CV settings in Section 3.1. Following a similar schema, the minimum requirement to specify a visual task is the task definition , where category names illustrate the target visual concepts in natural language. Though adding human-annotated examples is a natural way to clarify the task and yield the complete schema , extra cost is introduced.
It is of high interest to clarify the task using relevant examples, without human curating cost. Therefore, we propose to augment the task instruction with the retrieved examples from the external multi-modal knowledge base . For each concept in a given task, we first represent it in natural language using the language prompt as in , through inserting the concept into a set of task-specific templates . The task definition is expanded in its natural language form:
Next, we perform our knowledge retrieval process to acquire the relevant image-text pair from the source . Two types of retrieval processes are considered to acquire the top- pairs:
Text-to-Text (T2T) retrieval allows us to retrieve more relevant examples as they have a better match with our target concept. The T2T-retrieved set for is:
Text-to-image (T2I) retrieval allows us to have more diversity in the text descriptions in our retrieved examples. The T2I-retrieved set for is:
Both and are retrieved examples to augment the task definition , without accessing the images in the training or validation set of the task. Compared to , they are “free” external knowledge to clarify the task and can be used to build a more transferable system.
3 Model Customization
After retrieving the relevant multi-modal examples, one may employ the naive customization solution by fine-tuning the full-model initialized from pre-trained weights, as in Figure 3(a). Alternatively, we propose an affordable solution to endow pre-trained models with a new capability to leverage this external knowledge. The pre-trained generic visual models have gained strong transfer abilities and access to a large amount of internal knowledge stored in the model weights. We freeze the weights of these models so that their initial capacity remains unchanged. To bridge these pretrained models harmoniously to the customized domain, we consider locked-text gated-image tuning with the following two techniques, illustrated in Figure 3(d).
In order to provide sufficient expressivity to the model and make it able to adapt well on retrieved knowledge, we insert gated self-attention dense blocks in between the original layers of the image encoder, and train the new blocks from scratch. Those blocks are made of a self-attention layer, that attends the early layer inputs, followed by an extra dense feed-forward layer. Please see a visual illustration of this gated block in the rightmost of Figure 3(d). We denote the parameters of all new modules as . This design is inspired by the gated cross-attention-dense blocks in Flamingo and frozen multi-modal model . The difference is that the trainable module is introduced in Flamingo to enable cross-modal conditioning, while we adapt it for model growing in new customized domains.
The text encoder in language-image contrast models represents the task semantic space. To maintain it, we propose locked-text tuning, which freezes the text model weights so that the generic task encoding knowledge remains locked; see Figure 3(c). This is in contrast with locked-image tuning (LiT) in Figure 3(b), where the image encoder is frozen and the text encoder is fine-tuned, which teaches a text model to read out good representations from a pre-trained image model for new tasks.
We extract the normalized feature vectors in a hyper-sphere using and . To customize the model wrt task definition , we update using a bidirectional learning objective between images and language on the retrieved knowledge pool :
where is a temperature hyper-parameter controlling the strength of penalties on hard negative samples, and , . We set for classification tasks to force image-text pairs sharing the similar text to be positive. Note (4) is a general form; it reduces to UniCL when ; it further reduces to the training objective of CLIP or ALIGN when there is a one-to-one mapping between an image and its paired caption in a batch, i.e. and .
In our empirical study we find that locked pre-trained image and text encoders with trainable gated modules in image encoder work best. Once the customized visual models are trained with the retrieved knowledge, we transfer it to the downstream domain for zero/few/full-shot evaluation.
4 Discussions with Data-Centric Methods
It is recommended in that the task learning capabilities of machine learning (ML) systems can be measured by task-level zero-shot transfer. This recommended evaluation setting is further generalized in by showing that few/full-shot transfer consistently yields higher performance than zero-shot transfer. We argue that the task learning capabilities of ML systems can be improved from both the model and data perspectives. Most existing efforts devote to model-centric methods such as efficient network architectures , smarter training objectives , and scaling up model size . Data-centric methods are less explored, where our retrieval-augmented approach attempts to fill this data gap. We discuss the unique properties of React and build the connections with existing data-centric paradigms.
To build transferable visual systems, K-LITE enriches entities in language supervision with structural knowledge in WordNet and Wiktionary , in both model training and evaluation stages. It provides the first strong evidence that structural knowledge is effective in task-level transfer for CLIP/UniCL. Our paper is different in two aspects: Knowledge sources. K-LITE considers textual common sense knowledge bases, while ours considers the web-scale image-text corpus. Motivation. K-LITE aims to improve the generality of visual models via structural human knowledge, while ours improves the customization of visual models using a plug-and-play task instruction augmentation process.
As a semi-supervised learning algorithm, self-training provides pseudo labels to the unlabelled images using a pre-trained neural (teacher) model. Though sharing the similarity in expanding the task-relevant data, the two methods are different in the augmented knowledge: For an image, the supervision signal in self-training is based on the teacher model’s internal “dark knowledge” , which is limited in a fixed prediction space. The supervision signal in our method is the paired text, which is collected from web as the external knowledge, which may contain richer semantics to describe the image. We build a retrieval process to acquire task-relevant images, which is lacking in self-training. The two methods can mutually benefit: self-training can start from our retrieval-augmented pool, while we could use pseudo labels from self-training to get additional supervision.
Experiments
In this section, we conduct experiments to answer three research questions: (1) What are the unique advantages of retrieval-augmented image-text knowledge for task transfer? (2) How does our design choice of locked-text gated-image tuning compare to existing methods for model customization? (3) Is customization still beneficial in settings where the training data in downstream tasks are observed, i.e., in few-shot or full-shot settings?
We evaluate our models on two CV problems: image classification and image-text retrieval. We first consider ImageNet for zero-shot task transfer. We then further evaluate our model on Elevater , which is an open-set image classification benchmark that contains 20 datasets. We also conduct experiments on image-text retrieval with MSCOCO and Flickr datasets.
One of the most intriguing benefits of React is that it does not need access to any images from the downstream task. Therefore, we first evaluate on task-level zero-shot transfer, which requires no images in the target to be observed . This setting is different from traditional class-level zero-shot , where both the category and images in evaluation should not be observed in training. We argue that ImageNet concepts have been observed in CLIP (Sec. 2.2 of ) and other web-scale trained models , as WordNet synsets and common words in English Wikipedia are explicitly added in the query list when searching for (image, text) pairs in their training data construction process.
As shown in Table 1, by customizing the generic model CLIP/OpenCLIP on 10M retrieved image-text pairs from LAION-400M, React achieves a significant and consistent gain (up to 5.4%) on zero-shot image classification on ImageNet-1K, with different backbones and original pretraining datasets. There are three interesting findings.
F1: React can benefit from model’s own pre-training data. Compared to OpenCLIP (ViT-B/32) trained on LAION-400M, by training on 10M relevant pairs from the same LAION-400M dataset, React improves over OpenCLIP by 3.5%. Note that the model purely uses the image-text pairs that it has seen during its pre-training, and does not see any extra data. This shows that React can more adequately adapt to the target domain during the model customization stage, suggesting a favorable property that no new data is required for customization.
F2: React efficiently explores new image-text sources, even for large models. We costomize CLIP ViT-L/14 on 10M retrieved relevant image-text pairs, and the model achieves a 2.8% improvement to 78.1%. This surpasses all publicly available checkpoints from CLIP and OpenCLIP, including the checkpoint with a much larger ViT-H/14 backbone and trained on a much larger LAION-2B dataset. This suggests that React is a more sample-efficient approach to improve the model performance on the domain-of-interest.
F3: Scaling up the retrieval pool increases performance. We perform React in a privately collected dataset with over 800M pairs, and train a customized model on 6M retrieved pairs. The performance is increased to 78.5%, yielding 0.9% gain compared with 6M pairs retrieved from LAION-400M. This suggests that React scales well with the larger retrieval pool. It showcases React as a cost-efficient approach to leveraging the ever-increasing web image-text corpus.
We extend the scope of ImageNet-1K experiments to low-shot settings: 1% and 10% labelled data settings, and provide the first strong baselines using CLIP checkpoints. We present results in Table 2. First, we find that when the linear head is randomly initialized, it often results in sub-optimal low-shot performance, as the knowledge from the CLIP’s language encoder is completely discarded. We advocate using language-augmented initialization of the linear head , which improves the 1% label adaptation performance of CLIP ViT-B/16 from 70.9% to 74.3% (+3.4%). With React customization stage, it further improves by 3.1% to 77.4%, outperforming prior arts with similar model sizes. Furthermore, when we further scale up the model size to ViT-L/14, CLIP achieves 80.5% accuracy, which is on par with the previous SoTA. React further improves the accuracy by 1.1%, setting a new state-of-the-art of 81.6% accuracy on 1% label settings. Similar trend is observed in 10% label setting: CLIP is on-par with the prior art, and our React customization pushes the new SoTA towards 85.1% accuracy.
1.2 Zero-, Few-, and Full-Shot on Elevater
As a proxy for performing vision tasks for many customized scenarios in the wild, we consider the image classification in the wild (ICinW) benchmark in Elevater . It consists of 20 datasets from a diverse selection of domains and covers a wide range of concepts, totaling 1151 classes.
We perform multi-modal knowledge retrieval for 20 datasets together – the retrieved samples are around 10M image-text pairs in total, on which one single customized visual model is trained. After the process, we feed the customized model to different downstream tasks separately. For each downstream dataset, we use the official Elevater toolkit to obtain the train/val/test splits, and perform zero-shot, few-shot, and full-shot evaluation.
We report the average scores in Table 3. It achieves 3.8% improvement in the zero-shot setting, even when we do not perform a separate customization for different datasets. This demonstrates the robustness of our customization process. Further, we see the consistent improvement in few-shot and full-shot settings, including both linear probe (LP) and fine-tuning (FT). This result is encouraging, as it demonstrates that when we have access to some or all data from the downstream task, the proposed model customization stage remains beneficial. Therefore, we advocate model customization process in both data-limited and data-rich settings.
Next, we ask why does the retrieved image-text knowledge improve the zero-shot task transfer performance on a broad range of datasets? We compare the breakdown performance on all 20 datasets in Figure 4 for the zero-shot settings. Out of 20 datasets, the retrieval-augmented knowledge shows superior/comparable/inferior performance to the baseline on 15/1/4 datasets for CLIP and 14/0/6 datasets for OpenCLIP, respectively. Most of the improved and failure datasets are consistent for both checkpoints. For the top two datasets that gains the most, i.e. StanfordCars and FGVC Aircraft, relevant image-text knowledge is retrieved from the web-crawled data LAION-400M to describe the concepts; see Fig. 5(a). Interestingly, this observation is complementary to K-LITE , which failed on these two datasets, because no knowledge was extracted from Wiktionary for them, as it often requires domain-specific knowledge and even visual knowledge to best define a car brand (e.g. BMW X6 SUV or Audi R8) or an aircraft model type (e.g. DC-10 or A321).
As shown in Fig. 4, React struggles on the PatchCamelyon dataset, a cancer cell recognition benchmark. We visualize the retrieved samples and the samples from the original training set in Fig. 5(b). The retrieved images are either instruction photos and from another sensing method, which exhibits a different visual distribution from PatchCamelyon. This suggests the importance of ensuring the retrieval quality for the domain-of-interest.
2 Image-Text Retrieval
To demonstrate the generality of React, we consider Flickr30K and MSCOCO image-text retrieval tasks, in both zero-shot and full-shot settings. We use the standard image-text contrastive objective . For image-text retrieval task, following , we use the CLIP-L/14 with 336x336 input resolution in both zero-shot, customization, and fine-tuning stage. We use the captions from MSCOCO as queries to retrieve 6M image-text pairs and perform customization. Note that none of the caption queries are used in the model training stage.
As shown in Table 4, React improves the generic CLIP counterparts on both zero-shot and full-shot retrieval for Flickr30K and MSCOCO datasets. The gain on zero-shot task transfer is large. On Flickr30K, it achieves 3.4%/10.0% recall improvement for I2T and T2I retrieval, respectively. Afer fine-tuning on full training data, React still improves over the baseline slightly. It provides another piece of evidence for React in data-rich settings. Furthermore, we conduct the same customization procedure of React on a large checkpoint Bletchley with 864M parameters, and observe consistent gains over both datasets. It demonstrates that React scales well with model size on retrieval tasks.
3 Dense Prediction Tasks
Although React is optimized with the image-level contrastive loss during the customization stage, we find it beneficial for dense prediction tasks as well. We showcase its application to dense prediction tasks on object detection and semantic segmentation.
For object detection, we choose the state-of-the-art RegionCLIP as our framework. We conduct experiments in two settings: zero-shot inference with ground-truth (GT)/ Region Proposal Networks (RPN) boxes and open-vocabulary object detection on MSCOCO dataset. We perform the model customization following the same setting as Sec. 4.2. Following RegionCLIP, we conduct experiments on ResNet50 backbone. Additionally, we present results on ViT-B/16 backbone for zero-shot inference.
RegionCLIP alters the architecture of an object detector. It is able to perform zero-shot object detection, by (1) initializing its backbone and prediction head from the CLIP-pretrained checkpoints, and (2) employing a pretrained RPN network. The results are shown in Table 5. React consistently improves over CLIP checkpoint under all settings. When ground-truth region proposal is used, React improves over CLIP by +1.0 on overall AP50; when the pretrained RPN is used, React demonstrates +1.5/+1.4/+1.9 AP50 improvements on novel, base, and all classes, respectively. These results are encouraging, as it shows that the customized knowledge from React transfers well to dense prediction tasks like object detection, under the RegionCLIP framework.
We further conduct experiments on the open-vocabulary settings, where the model finetunes on a set of selected categories (base classes), and evaluate on both seen (base) and unseen (novel) classes. We report the results in Table 5 (OVD).
We can see that with the React customization, the detector yields improved performance on base with +2.3 AP50, and importantly, it significantly improves novel categories with +6.4 AP50. This suggests that the injected knowledge during the model customization stage improves the learned fine-grained visual feature that is beneficial to both seen and unseen categories for object detection, when the downstream coarse-grained data is available. This is favored, because (1) the weakly-supervised data such as the coarse-grained image-text pairs requires much less human annotation cost than fine-grained bounding box annotation, (2) the paired data in React is free, as it is retrieved from the web, where COCO image-text pairs are not used in customized training.
3.2 Semantic Segmentation
For semantic segmentation, we choose the state-of-the-art MaskCLIP as the framework. It investigates three evaluation settings for segmentation. First, it makes use of the pretrained CLIP checkpoint to discover the alignment between grid visual features and the text prompt features, so as to perform zero-shot semantic segmentation. Second, to further improve the performance, MaskCLIP proposes two techniques for refining its zero-shot predictions: key smoothing and prompt denoising. Lastly, when the training images are available, without the need to access the training labels, it further proposes MaskCLIP+ to perform full-shot finetuning on the target training set using pseudo-labels. Following MaskCLIP , we use ViT-B/16 checkpoints, and use their official code base to train and evaluate models. We report results in Table 6.
On all of the three settings, React demonstrates improvements over the MaskCLIP, with both locked-text and lock-text-gated-image tuning. This indicates that React-based model customization is beneficial for dense prediction tasks as well. Notably, when key smoothing and prompt denoising is used, React with locked-text tuning demonstrates a significant 3.6% gain in mIoU. Besides, the locked-text tuning benefit more than gated-image tuning, when prediction refinement techniques are used. We hypothesize that locked-text tuning may generate more noisy predictions comparing with gated-image predictions, and it can thus benefit more with the proposed refinement techniques.
Surprisingly, without seeing the downstream COCO images, the zero-shot evaluation of React (18.2) even slightly outperforms MaskCLIP+ (18.1), which is finetuned on the downstream training COCO images with self-training.
4 Ablation Studies
We ablate React on ImageNet with CLIP ViT-B/32 checkpoint, with 10M retrieved image-text pairs from LAION-400M.
We conduct an experiment by adding a single GSA block before each Transformer block and continual pretraining the model on the retrieved image-text pairs. We then visualize the learned gated values in Fig. 6 (top): as the network goes deeper, the gate values become larger, and compared with other blocks, the learned alpha gates in the first six layers have a much smaller value. We hypothesize that earlier blocks have a smaller modulation to the base network, and removing them has minimal effect on the model’s performance. Therefore, we vary the number of last layers which we add gated blocks to, and show empirical results in Fig. 6 (bottom). With more layers added to the base network, the model’s zero-shot performance gradually increases, and saturates at 6 blocks. Therefore, we add gated blocks to last 6 layers as the default to balance the accuracy and the efficiency.
We ablate the design choices in the model customization stage: (1) direct tuning the pre-trained weights vs. training gated blocks from scratch; (2) updating visual vs. text encoder. We report results in Table 7. First, a frozen text encoder consistently outperforms a frozen visual encoder. We conjecture the phenomenon is due to that the retrieved texts have a much more limited space, comparing to text space in the original pre-training set (e.g. LAION-400M), as the concepts are limited to the query classes from the target domain. Therefore, fine-tuning the text encoder may tend to collapse the pre-trained semantic space.
We advocate two tuning methods for model customization. Locked-text gated-image tuning has a strong adaptation power, and is efficient in the model customization stage, with fewer trainable parameters. Locked-text tuning is also an effective way of customizing the models to downstream tasks, without the need of adding extra parameters. By default, we use gated blocks for its superior performance and efficiency.
We observe that training with a small retrieval size is more likely to overfit. We find that a larger retrieval size generally yields better performance, and saturates at around 6-10M image-text pairs.
Three search methods are compared: T2I, T2T, and T2I/T2T-combined. We find all modes consistently improve over the baseline CLIP. T2I retrieval alone yields a slightly worse performance, which may be partly due to its retrieved samples being more noisy and less relevant than T2T retrieval. T2T alone or T2I/T2T-combined has a similar performance. We use T2I/T2T-combined as our default strategy.
We study and compare with the pseudo labeling strategy on retrieval-augmented model customization. After the relevant image-text pairs are retrieved, we use the pretrained CLIP checkpoint to assign pseudo labels to each retrieved image, and finetune the model as a classification task using the retrieved samples and pseudo labels. We optimize the network with UniCL loss . With frozen text, pseudo-labeling can improve the model’s performance with T2I/T2T data, while having a decreased performance with T2I data. This can be due to that the relevance of the T2I data is lower than other splits, and pseudo label assigned can be incorrect. Besides, gated self-attention is complimentary to pseudo labeling with 2% improvements. In contrast to pseudo labeling, training directly on the retrieved image-text pairs does not use heuristics to create pseudo labels, and the model can receive additional supervision signal, which we empirically find helpful for model adaptation and robustness.
Conclusion
We presented React, a plug-and-play framework for leveraging large-scale image-text corpus as external knowledge to efficiently customize models on downstream tasks. Extensive experiments demonstrate its generality and effectiveness in image classification, image-text retrieval, object detection, and semantic segmentation, on more than 20 different downstream datasets. We highly advocate the model customization stage for building more transferable visual system for different downstream tasks.
We thank Min Gao, Ping Jin, Houdong Hu from Microsoft Azure AI for LAION data preprocessing and evaluation scripts, Qiang Liu and Subhojit Som from Microsoft Turing Team for the Bletchley model checkpoints. This work was supported in part by NSF CAREER IIS2150012, NASA 80NSSC21K0295, and Institute of Information & communications Technology Planning & Evaluation(IITP) grants funded by the Korea government(MSIT) (No. 2022- 0-00871, Development of AI Autonomy and Knowledge Enhancement for AI Agent Collaboration) and (No. RS-2022-00187238, Development of Large Korean Language Model Technology for Efficient Pre-training).
References
Appendix
In Section A, we present more results on ImageNet (Sec. A.1) and ELEVATER (Sec. A.2), with additional studies on a broader selection of checkpoints and more visualizations for a better standing of our approach.
In Section B, we provide implementation details (Sec. B.1,B.2) and cost analysis (Sec. B.3) of our retrieval system and model customization pipelines.
Appendix A More Results on Image-level Tasks
To further study the improvement of React over the pretrained vision-language models, we present more results with different pretraining data (WIT-400M, LAION-400M, LAION-2B) and different vision Transformer model sizes (B/32, B/16, L/14). For a fair study, we use the same set of 10M retrieved image-text pairs from LAION-400M dataset for all configurations. Results are presented in Table 8 (first column).
From the table, we can see that React with both tuning strategies consistently improves over the base checkpoints across different pretrained data and different vision backbones, and locked-text-gated-image tuning consistently performs better than the locked-text tuning only. Though both benefiting from React customization, CLIP checkpoints that are trained on WIT-400M data benefit slightly more than OpenCLIP checkpoints. This suggests that during the model customization stage, leveraging unseen data can potentially give the model a larger gain compared with the seen data during pretraining.
We further study the case when all retrieval data is already observed by models during their pretraining stage. Specifically, we study OpenCLIP checkpoints pretrained on LAION-400M , and LAION-2B (a super-set of LAION-400M). From the results, we see that by revisiting the already observed LAION-400M data, React (locked-text) shows +2.8/+3.0 improvements on B32/B16 checkpoints, respectively, which purely comes from the model customization stage, with neither additional model parameters, nor additional training data. Interestingly, even on OpenCLIP checkpoints that is pretrained with a much larger LAION-2B, React can still improves over OpenCLIP by +0.9/+2.9 with B32 backbone with locked-text and locked-text-gated-image tuning strategy, respectively. These findings suggest that leveraging the original pretraining dataset only at the pre-training stage is sub-optimal, it is of much larger potential to explore the web-scale data using the proposed model customization stage.
We also conduct zero-shot evaluation on other ImageNet variants: ImageNet-V2 , ImageNet-R , ImageNet-A , ImageNet-Sketch in Table 8. With the React customization, model robustness towards different ImageNet variants consistently improves on ImageNet-V2, ImageNet-R, and ImageNet-Sketch. We notice that for some checkpoints, the accuracy drops after model customization on ImageNet-R dataset: an adversarial dataset with a collection of selected images from the web that can “fool” common classifiers. We find that classifiers trained on the LAION dataset are more prone to such adversarial attacks, while React customization helps it recover from such attacks to some extent: accuracy improves for OpenCLIP checkpoints that are trained on LAION, especially when locked-text gated-image strategy is used.
We further study the full-shot performance on ImageNet-1K of React using the linear probing protocol. ImageNet-1K contains around 1.28M training images, and it represents one of the most standard data-rich settings. We use the DINO code base for the linear probe experiments. As shown in Table 9, React improves over CLIP by +0.6/+1.9 with the locked-text and locked-text-gated-image tuning, respectively. This suggests that the React customization adequately adapts the visual encoder to the ImageNet domain, resulting in better feature representations.
There is a chance that the LAION-400M dataset contains some of downstream ImageNet images, and our retrieval system may retrieve these image-text pairs. One may question that if the performance gain of React model customization actually comes from these samples.
We carefully study the de-duplication experiments. We compute the pairwise distance of the visual features between the images from the retrieved set and the ImageNet train/val set, and set the cutoff threshold to 0.95 (Fig. 7, Bottom). Note that 0.95 is a high threshold, as 85K (1% of 10M total retrieved images) images are removed, among which, only a few of them overlap with ImageNet train/val. This suggests that the LAION data contain ImageNet images, making the publicly available OpenCLIP checkpoints less rigorous when reporting the zero-shot task transfer performance. As for CLIP, as its pre-training data is not publicly available, it remains unknown if any ImageNet images are observed in its pre-training. We set it to 0.95 mainly to ensure that the overlapping images are removed from the retrieved sets so as to carefully study its effect.
As shown in Table 10, even after aggressively removing 85K images, the final model’s performance is similar (-0.2%) to the checkpoint trained on the unfiltered retrieved set. We further visualize the validation accuracy curve as training proceeds in Fig. 7 (Top). The model behaviors during training are very similar between filtered (dashed curve) and unfiltered (solid curve) retrieved sets, for both tuning strategies. This ensures that the gains are not due to the overlapping samples, and React effectively learns and adapts to the ImageNet domain during the model customization stage.
The quality of retrieved data matters for vision-language pre-training. As an initial attempt of the quality control, we consider to use the CLIP score to select the high relevant retrieved image-text pairs. The distribution of the CLIP score is visualized at Fig. 8. We choose CLIP score 0.3 and 0.32 as two thresholds (): a threshold of 0.3 filters the low-quality samples while keeping the total number of retrieved samples roughly the same (93.5% retrieved samples are kept), a threshold of 0.32 performs a more aggressive filtering and keeps around 6M samples (sufficient for React customization according to Sec. 4.4).
As shown in Table 12, React is robust towards noise in the pretraining data. When filtering using a CLIP score threshold of 0.3, the model customization performance roughly remains the same. When filtering with a threshold of 0.32, the customization performance drops by around 1%, which suggests that the filtered samples contain useful information for model customization. In conclusion, our React model customization is robust against the noises in the retrieval dataset. Therefore, we leave a more sophisticated quality control approach to future work.
A.2 ELEVATER
We present the full-spectrum breakdown results for zero-shot, few-shot, and full-shot experiments on the ELEVATER benchmark in Table 11. We mark the experiment runs that React yields gains compared with baseline CLIP in green and with bold font.
First, across zero-shot, few-shot, and full-shot settings, React consistently improves over the baseline CLIP. However, on the ELEVATER benchmark, locked-text and locked-text-gated-image strategy works better in different cases: for zero-shot, locked-text-gated-image works much better than locked-text with 1.4% improvement; while for other cases, they perform similarly well, and locked-text is slightly better in 3/4 cases. This may be partly due to that we are training a unified checkpoint across different domains, and random noises during the final adaptation stage can cause small variations on different datasets.
Second, we find that across different data regimes and different tuning strategies, the gains and losses are mostly consistent on a fixed set of datasets. This provides another clue for that the gains and losses are highly correlated with the retrieval quality, and the gain/loss conclusion can generally transfer to different tuning configurations.
We present more retrieved samples from the ELEVATER datasets to illustrate properties of React.
First, we show one example of the benefit of using retrieved image-text pairs for training instead of using pseudo labels. As shown in Fig. 9, the query is Volvo sedan, while one of the retrieved sample is an Audi. The retrieved text contains the correct brand Audi, and the alignment between the retrieved image-text pairs can help correct the retrieval mistake and aid the model training. If the retrieved images are annotated as the same label with query, and used the pseudo-labelled pairs for training, the aforementioned finding suggests that the approach would perform worse than leverage the “true” image-text pair knowledge crawled from the web.
Second, we visualize examples retrieved by text-to-text and text-to-image retrieval, using the same query: “a painting of a flamingo”. As shown in Fig. 10, text-to-text retrieval can retrieve samples that is more accurately matching the query (“a painting”). On the other hand, text-to-image retrieval gives a more diverse text description, while it may not have a perfect match between the query and the retrieved text sample (“sketch vector”).
Appendix B Implementation details
We mainly conduct our experiments on the vision Transformer backbones. For the ViT architecture, we mainly follow the implementation from CLIP . The feature from the CLS token from the last visual encoder layer is used as the visual feature. For the gated-image experiments, we only add gated blocks to the last 6 layers. The hidden/embedding dimensions for gated blocks are set the same as the layer that it is added to. Following , the gate values are initialized as zero, modulated by tanh operator.
We mainly follow CLIP and UniCL to set up our training hyperparameters. For optimization, we use AdamW with a weight decay of 0.05 for all models. We set the learning rate to 0.0005 for locked-text-gated-image experiments, and 0.00005 for locked-text experiments. We use the same set of data augmentation and regularization as in . For experiments with 10M retrieved samples, the models are trained for 32 epochs with a batch size of 4096. For experiments with fewer retrieved samples, the training epochs are adjusted accordingly so that they have a similar number of optimization steps. For all training, we used a cosine learning rate schedule, with 5000 iterations warmup.
B.2 Our Retrieval System
We implement our retrieval system using FAISS . We use its Hierarchical Navigable Small World (HNSW) approximate -NN lookup to balance performance and efficiency. Product quantization is used to reduce the index size. We use Autofaiss to select the optimal hyperparameters for the index, and build the index using FAISS index factory. For LAION-400M, the selected configuration is: OPQ256_768,IVF131072_HNSW32,PQ256x8. We build two separate indexing systems for T2I and T2T retrieval. For T2I retrieval, CLIP image features are used for building the indexing system. For T2T retrieval, CLIP text features are used.
We benchmark below the latency and the recall of the HNSW -NN lookup on a server with 64 CPU cores. As shown below, the indexing system is able to retrieve the relevant vectors accurately and efficiently.
B.3 Cost Estimation
We provide the cost estimation for the React pipeline. It includes feature extraction of the retrieval pool, indexing for the retrieval system, querying the indexing system to retrieve relevant image-text pairs, and finally model customization.
For feature extraction using CLIP ViT-B/32 checkpoint, it takes around 250 T4 GPU hours, which equates to roughly a day on a desktop with 4x RTX 3090 GPUs.
We build our index system on a cloud VM with 24 CPU cores. Using the selected configuration in Sec. B.2, we can build the index system using FAISS within 20 hours. Note that we use the CPU version of FAISS, and do not leverage GPU acceleration for building the index.
As shown in Sec. B.2, generating the retrieval set for model customization using the indexing system is very efficient. Typically, it takes less than 10 minutes to generate 10M indices for our retrieval pool.
We train most of our models on a compute node with 16V100 GPUs. For ViT-L checkpoints, we use 2-node distributed training, each node with 16V100. It takes around 16/28/42 hours to train the B32/B16/L14 checkpoint, respectively, with either locked-text or locked-text-gated-image tuning strategy.
Note that the feature extraction and building index system only needs to be done once. They are readily available for any queries from any domains. To make this line of research more accessible, our retrieved subsets for both ELEVATER and ImageNet will also be made available, with an easy-touse toolkit to download the subsets without storing the whole dataset locally. Therefore, researchers can directly run experiments on model customization.
In conclusion, React framework provides an accessible way to explore the large-scale web-crawled image-text dataset, and effectively and efficiently customize the models to the domain-of-interest.