Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners

Renrui Zhang, Xiangfei Hu, Bohao Li, Siyuan Huang, Hanqiu Deng, Hongsheng Li, Yu Qiao, Peng Gao

Introduction

Convolutional neural networks and transformers have attained great success on a wide range of vision tasks with abundant datasets . Instead, for some data-deficient and resource-finite scenarios, few-shot learning also becomes a research hotspot, where the networks are constrained to learn from limited images with annotations. Many previous works have been proposed in this field to enhance model’s generalization capability by meta learning , metric learning , and data augmentation . Recently, CLIP pre-trained by large-scale language-image pairs shows favorable zero-shot transfer ability for open-vocabulary visual recognition. The follow-up CoOp , CLIP-Adapter and Tip-Adapter further extend it for few-shot classification and achieve superior performance on various downstream datasets. This indicates that, even if the few-shot training data is insufficient, the large-scale pre-training has endowed the network with strong representation ability, which highly benefits the few-shot learning on downstream domains. Now that there exist various self-supervisory paradigms besides CLIP, could we adaptively integrate their pre-learned knowledge and collaborate them to be a better few-shot learner?

To tackle this issue, we propose CaFo, a Cascade of Fooundation models blending the knowledge from multiple pre-training paradigms with a ‘Prompt, Generate, then Cache’ pipeline. As shown in Figure 1, we integrate CLIP , DINO , DALL-E , and GPT-3 to provide four types of prior knowledge for CaFo. Therein, CLIP is pre-trained to produce paired features in the embedding space for every image and its descriptive text. Guided by texts with different categorical semantics, CLIP can well classify the images aided by language-contrastive knowledge. DINO follows contrastive self-supervised learning to match the representations between two transformations of one same image, which is expert at distinguishing different images with vision-contrastive knowledge. Similar to CLIP , DALL-E is also pre-trained by image-text pairs but learns to predict the encoded image tokens based on the given text tokens. Conditioned on the input text, DALL-E could leverage the vision-generative knowledge to create high-quality synthetic images in a zero-shot manner. Pre-trained by large-scale language corpus, GPT-3 takes a few hand-written templates as input, and autoregressively generates human-like texts, which contain rich language-generative knowledge. Therefore, the four models have distinctive pre-training goals and can provide complementary knowledge to assist the few-shot visual recognition.

In detail, we cascade them by three steps.: 1) Prompt. We adopt GPT-3 to produce textual prompts for CLIP based on a few hand-written templates. These prompts with richer language knowledge are fed into CLIP’s textual encoder. 2) Generate. We adopt DALL-E to generate additional training images for different categories based on the domain-specific texts, which enlarges the few-shot training data, but costs no extra manpower for collection and annotation. 3) Cache. We utilize a cache model to adaptively incorporate the predictions from both CLIP and DINO . Referring to Tip-Adapter , we build the cache model with two kinds of keys respectively for the two pre-trained models. Regarding zero-shot CLIP as the distribution baseline, we adaptively ensemble the predictions of two cached keys as the final output. By only fine-tuning the lightweight cache model via expanded training data, CaFo can learn to fuse diverse prior knowledge and leverage their complementary characteristics for better few-shot visual recognition.

Our main contributions are summarized as follows:

We propose CaFo to incorporate the prior knowledge learned from various pre-training paradigms for better few-shot learning.

By collaborating CLIP, DINO, GPT-3 and DALL-E, CaFo utilizes more semantic prompts, enriches the limited few-shot training data, and adaptively ensembles diverse predictions via the cache model.

We conduct thorough experiments on 11 datasets for few-shot classification, where CaFo achieves state-of-the-art without using extra annotated data.

Related Work

With the breakthroughs in deep learning models , most modern vision models are based on the paradigm of pre-training on ImageNet and fine-tuning on downstream tasks . Pre-trained models have shown promising adaptability for various downstream tasks, such as object detection , semantic segmentation , and 3D recognition . To improve the representation capability by overcoming the constraints of annotation, self-supervised pre-training has attracted wide attention using large-scale unlabeled datasets . Self-supervised learning is initialized by pretext tasks, such as image restoration from corruption , pseudo labels and clustering . Recently, contrast learning, which learns representations by contrasting positive pairs against negative pairs, has gotten well studied for diverse visual representation learning . Besides, language-supervised visual pre-training emerges as a novel paradigm closer to natural visual understanding , among which CLIP obtains powerful zero-shot transferability by contrastive pre-training on image-text pairs from the Internet. In addition, vision-language pre-training can also promote the zero-shot image generation from text. Open generative models, such as DALL-E and CogView pre-trained on large-scale image-text pairs are able to generate images with diverse contents by given texts. In this paper, CaFo cascade three visual pre-training models, CLIP, DINO, and DALL-E, which contributes to better few-shot learning capacity.

Language-assisted Vision Models.

As different form of data, linguistic knowledge normally contains complementary knowledge to images. For vision-language models, several works have showed the format of prompts would highly affect the accuracy on vision tasks. Thus, prompt engineering is worth putting in great effort. Some efforts utilize learnable textual inputs and optimize them during training. Other works propose to leverage linguistic knowledge pre-trained from large language models to generate prompts for each visual category, which enhances vision-language models without any additional training or labeling. Our CaFo refers to CuPL to produce semantic-rich texts to prompt CLIP for better text-image alignment.

Few-shot Learning.

Few-shot learning highly relies on the transferability of the trained neural networks. From the perspective of distance measurement, some metric learning methods learn a metric space by computing the distances from the instances to novel categories . Also, meta-learning is proposed to improve the few-shot adaptation ability of the models by finding a set of initialized parameters that can rapidly adapt to novel domains . More recently, with the vision-language pre-training model CLIP exhibiting strong zero-shot adaptation performance, several efforts have started to find efficient strategies to adapt it to downstream few-shot datasets. CoOp is proposed as a prompt tuning adaptation method by optimizing a set of learnable prompt tokens. Subsequently, to inject textual branch with visual signals, CoCoOp and VT-CLIP propose to train a intermediate network to generate image tokens as conditional inputs for the textual vectors. Referring to adapters in natural language processing, CLIP-Adapter is introduced to fine-tune CLIP by applying lightweight residual-style adapters. Tip-Adapter is then proposed as a training-free adaption method with a constructed key-value cache model. It can also be regarded as a better initialization of CLIP-Adapter with much faster convergence when fine-tuning. CALIP proposes a parameter-free attention to enhance CLIP in a zero-shot manner, and its parametric solution further attains higher few-shot accuracy. SuS-X constructs a dynamic support set and extends Tip-Adapter by leveraging image-text distances. Besides, many follow-up works have also been proposed for further adapting CLIP to various vision tasks. Different from all existing methods, we integrate other powerful pre-training paradigms with CLIP and collaborate them with customized pipelines.

Cascade of Foundation Models

In this section, we first briefly revisit four types of pre-training paradigms in CaFo. Then, we specifically introduce how we cascade them by ‘Prompt, Generate, then Cache’.

The series of contrastive learning between vision and language learn to map the two modalities into the same embedding space via a contrastive loss. Driven by web-scale datasets, e.g., 400 million for CLIP and 1.8 billion for ALIGN , the basic pre-training target is to minimize the embedding distances of images and their textual descriptions, while maximize those unpaired ones. By the cross-modal alignment, we can discriminate images of different categorizes by the texts with different semantics. We denote such learned prior as language-contrastive knowledge and adopt CLIP as the representative model for such pre-training method.

Contrastive Vision Pre-training.

As the traditional self-supervised learning methods, vision-contrastive models focus on the discrimination between different images. Normally, the positive pairs to be drawn close are two transformations of the same image, while the optimization of negative pairs is optional, which can be replaced by a momentum encoder or cluster assignments . Recent works reveal that we can learn self-supervised features without negative pairs between images . Given the strong linear classification capacity, the pre-trained DINO is adopted here to provide vision-contrastive knowledge for collaboration.

Generative Language Pre-training.

With 175 billion parameters, the large-scale pre-trained GPT-3 is powerful to produce human-like texts with diverse contents and incredible quality. Taking as input a few designed language commands, GPT-3 is able to output prompts with rich linguistic semantics for vision-language models. CLIP utilizes handcrafted templates as prompts, e.g., “a photo of a [CLASS]”, which however lacks sufficient textual semantics to align with input images. We thus leverage GPT-3 to produce CLIP’s prompts to better align with visual information from images.

Generative Vision-Language Pre-training.

Learned from millions of image-caption pairs, the DALL-E series can generate language-conditioned images in a zero-shot manner. They are pre-trained to autoregressively predict the encoded image tokens from the textual tokens of the captions. With such language-generative knowledge, the pre-trained DALL-E can be viewed as a free lunch to enlarge the training data without any manpower. Considering publicity, we select DALL-E-mini as the representative among DALL-E models.

2 Prompt, Generate, then Cache

To cascade different pre-training paradigms, we introduce CaFo with a pipeline of ‘Prompt, Generate, then Cache’, which respectively unleashes the powers of different self-supervised knowledge.

Under the NN-way KK-shot settings, we have the few-shot training images IN,KI_{N,K} with labels LN,KL_{N,K} that contain KK samples for each NN categories. As shown in Figure 2, for NN categories, we adopt a unified series of templates as the language command for GPT-3 , e.g., “What a [CLASS] looks like?”, “How can you identify a [CLASS]?”, and “A caption of an image of a [CLASS]:”. We denote the created prompts for NN categories as PNP_{N}, formulated as

Then, we adopt PNP_{N} as the input of CLIP’s textual encoder. Further, for some downstream data with specialized categories, we can customize the language commands for producing prompts with more domain-specific semantics. For example, in OxfordPets dataset of pet images, we adopt the input of GPT-3 as “This is a pet bulldog, it has thin neck, short face, floppy ears. It’s coat is short, straight, and in brindle color. This is a pet [CLASS],”. Based on that, GPT-3 continues to describe more details of the [CLASS] pet.

Generate via DALL-E

Via the zero-shot DALL-E , we generate synthesis images to enrich our limited training images IN,KI_{N,K}, as shown in Figure 3 (1). For different categories, we adopt a simple template, e.g., “a photo of a [CLASS].”. After the generation, we utilize CLIP to filter the top-K′K^{\prime} best-quality images as the newly-expanded training samples for each category. Then, we obtain the NN-category (K+K′K+K^{\prime})-sample training images, formulated as

where TNT_{N} denotes the NN-category textual inputs. We keep K′K^{\prime} comparable with KK to ensure the synthesis quality and also preserve the low-data regimes. By the pre-trained language-generative knowledge, the data expansion is totally zero-shot, which requires no manpower to collect or annotate the data, and alleviates the data deficiency issue inherently for few-shot learning.

Cache by CLIP and DINO.

We construct a key-value cache model for adaptive knowledge ensemble. Different from Tip-Adapter only adapting CLIP, our cache model contains the pre-learned knowledge from both CLIP and DINO by caching two kinds of keys. Specifically in Figure 4 (2), we first utilize CLIP and DINO to independently extract visual features of the few-shot training images, formulated as

3 Adaptive Inference

where CLIPtex represents CLIP’s textual encoder, PNP_{N} denotes GPT-3’s created prompts, and fCLIPFCLIPTf_{\text{CLIP}}F_{\text{CLIP}}^{T} denotes the query-key affinity matrix of the CLIP’s keys, analogous to DINO’s. φ(x)=exp⁡(−β⋅(1−x))\varphi(x)=\exp(-\beta\cdot(1-x)) serves as a non-linear modulator to control the sharpness of affinity matrix.

As the language-contrastive pZSp_{\text{ZS}} is pre-trained by 400 million data and can perform strong zero-shot transfer ability, we regard pZSp_{\text{ZS}} as the prediction baseline and calculate the weights of pCLIP,pDINOp_{\text{CLIP}},p_{\text{DINO}} for ensemble based on their distribution similarity with pZSp_{\text{ZS}}. By this, we can suppress some obviously false category possibilities in pCLIP,pDINOp_{\text{CLIP}},p_{\text{DINO}} and also amplify the moderately correct ones during ensemble. Firstly, we respectively normalize the scales of three classification logits into -1∼\sim1 by their each mean and standard deviation. We then calculate the distribution similarities as the ensemble weights for the two logits of the cache as

Finally, we adopt the softmax function to normalize the weights and obtain the final ensemble logits as

where i∈{CLIP,DINO}i\in\{\text{CLIP},\text{DINO}\}. By such similarity-based ensemble, penp_{en} can adaptively fuse the prior knowledge learned by CLIP and DINO’s pre-training and achieve stronger few-shot image classification.

Experiments

We conduct few-shot experiments on 11 publicly available datasets: ImageNet , StandfordCars , UCF101 , Caltech101 , Flowers102 , SUN397 , DTD , EuroSAT , FGVCAircraft , OxfordPets , and Food101 . We follow Tip-Adapter to train CaFo with 1, 2, 4, 8, 16 shots and test on the full test set. As we adopt DALL-E to generate training images in a zero-shot manner, we can train CaFo only by the generated images and report its zero-shot performance without few-shot training set.

Implementation.

Our CaFo integrates the knowledge from pre-trained CLIP , DINO , DALL-E , and GPT-3 . For CLIP, we utilize ResNet-50 as the visual encoder and its aligned transformer as the textual encoder. To align with the visual representation from CLIP, we also adopt DINO pre-trained upon ResNet-50. For DALL-E, we adopt different domain-specific textual templates as the input for different datasets, which correspond to the original textual prompts for CLIP’s textual encoder. For GPT-3, we adopt five simple templates as the language commands shared by different categories. Each command outputs ten prompts, which obtains fifty prompts in total. For each category, we simply ensemble the features of different prompts following CuPL . During training, we only set the two kinds of keys in cache model to be learnable and utilize the data augmentation following Tip-Adapter-F. We train CaFo using batch size 64 only for 20 epochs, and adopt AdamW optimizer with the initial learning rate 0.0001 with a cosine scheduler. Note that, we tune the hyperparameters in CaFo by the official validation sets.

2 Performance

We compare CaFo with other CLIP-based adaption methods on the most representative ImageNet : CALIP , Linear-probe CLIP , CoOp , CLIP-Adapter , Tip-Adapter-F , and CALIP-FS . All these methods are based on the pre-trained CLIP with ResNet-50 visual encoders. As reported in Figure 5 and Table 2, CaFo surpasses all existing methods for different shot settings. Remarkably, CaFo with 1 shot even outperforms the 8-shot Linear-probe CLIP and CoOp, and CaFo with 8 shots is better than all methods with 16 shots. For zero-shot learning, CaFo sigiificantly surpasses CLIP and CALIP, demonstrating the importance of DALL-E’s generation. In Table 1, we present the efficiency of CaFo concerning training epochs and time. Our CaFo achieves the best performance-efficiency trade-off with 68.79% accuracy and only 10 minutes training.

On Other Datasets.

To further assess the robustness in different scenarios, we test CaFo on extra 10 datasets in Figure 6. For different semantic domains including real-world scenes, detailed textures, and satellite-captured landscapes, CaFo consistently shows leading performance and indicates excellent robustness via the collaboration of diverse knowledge. Notably, on some datasets, e.g., Caltech101 and OxfordPets, the zero-shot CaFo perform even comparably to other methods with 4 shots, demonstrating the effectiveness of zero-shot DALL-E for few-shot data expansion.

Distribution Shift.

We further evaluate the robustness of CaFo to distribution shift by training on “Source” dataset and testing on “Target” datasets. In Table 3, we select the “Source” as ImageNet and the “Target” as ImageNet-V2 and ImageNet-Sketch . As we can utilize some prior knowledge of the target domain for GPT-3 and DALL-E for prompting and generation, CaFo achieves the best out-of-distribution performance on the two “Target” datasets, surpassing the second-best Tip-Adapter-F by +3.28%, +0.88%, and +3.43%, respectively.

3 Ablation study

In Table 4, we explore how each pre-trained model contributes to the collaboration on different shots of ImageNet. Therein, “CLIP” denotes the zero-shot CLIP with cache model containing only CLIP’s keys, and “DINO” denotes only the cache model with DINO’s keys. As shown in the first three rows, the CLIP’s language-contrastive knowledge performs stronger than DINO’s vision-contrastive knowledge, which might benefit from millions of pre-training data. Their adaptive ensemble by cache model can bring larger improvement when the shot number increases. For the next two rows, DALL-E and GPT-3 can independently boost both CLIP and DINO for nearly all shots with the prompts and generated synthetic images. The last row represents our final solution, CaFo that incorporates all three pre-trained models with the best performance for all shots.

Generated Number via DALL-E.

We utilize DALL-E to generate synthetic images as the expanded few-shot training data. In Table 6, we explore the best synthetic number K′K^{\prime} for each category of different shots on ImageNet. We observe that the larger K′K^{\prime} does not lead to better few-shot performance. As we adopt pre-trained CLIP to select the top-K′K^{\prime} generated images, which are scored by the similarities between CLIP-encoded images and category texts, the larger K′K^{\prime} would contain more low-quality images and adversely affect the cache model. Furthermore, the amount of expanded data is comparable to the original KK shots and thus preserves the characteristic of few-shot learning.

Adaptive Inference.

In Table 5, we ablate different ensemble methods of CLIP and DINO’s predictions during inference on ImageNet. The first two rows represent the cache model with one type of keys respectively for two pre-trained models without ensemble. Then, we adopt average and maximum pooling between the two predictions and ensemble the result with pZSp_{\text{ZS}}. However, such naive integration without adaptive weights causes accuracy degradation. In the last three rows, we calculate the distribution similarities for adaptive ensemble and respectively select the three logits as the baseline. As shown, using pZSp_{\text{ZS}} as the distribution baseline performs the best, since pZSp_{\text{ZS}} itself shows strong transfer ability and can effectively suppress the wrong predictions of other logits.

CLIP’s Visual Encoders.

We conduct CaFo with different CLIP’s visual encoders for comparison with other methods. As shown in Table 7, CaFo consistently achieves leading performance with different visual backbones, indicating our generalizability to network architectures.

4 Visualization

In Figure 7, we visualize the synthetic images generated by DALL-E on ImageNet , OxfordPets and Caltech101 . As shown, benefited from the vision-generative knowledge, the generated images can well highlight the downstream semantics of target category and effectively expand the few-shot training set in low-data regimes.

GPT-3’s Prompts for CLIP.

In Figure 8, We present a rectified example in ImageNet aided by GPT-3’s prompts in CaFo. As shown, prompting by GPT-3 (Left) produces more semantic texts compared to CLIP’s handcrafted templates(Right), and better depicts the visual appearances in the image, which predicts the correct category of goldfish.

Conclusion

We propose CaFo, a cascade of foundation models that comprehends diverse knowledge from different pre-training and follows the ‘Prompt, Generate, then Cache’ pipeline. We first incorporate the generative language model, GPT-3, for prompting CLIP with more semantic texts, and adopt DALL-E to expand the few-shot training data. Then, we adaptively fuse the vision-contrastive DINO with CLIP via a unified cache model. By collaboration, CaFo achieves state-of-the-art performance for few-shot learning on 11 datasets. Although CaFo has unified four types of pre-training, our future direction will focus on integrating more existing pre-trained knowledge, such as the masked-generative MAE , the 3D-contrastive CrossPoint , and 3D-generative I2P-MAE .

References

Appendix A Additional Performance Comparison

In Figure 9, we compare the performance of CaFo without DALL-E ’s generated images or GPT-3 ’s created prompts on 10 datasets, which still consistently outperform the second-best Tip-Adapter-F.

Appendix B Additional Ablation Study

For the cache model, we investigate other pre-trained foundation models besides CLIP and DINO , including SimCLR , MAE , and SLIP . We preserve the prompting and generation by GPT-3 and DALL-E , along with We the pZSp_{\text{ZS}} as the ensemble baseline during adaptive inference. As shown in Table 8, ‘CLIP+DINO’, as our final solution, performs the best. Also, as an enhanced version of CLIP, SLIP can achieve higher accuracy in CaFo.

Zero-shot CaFo.

As we leverage the pre-trained DALL-E to generate the supplementary few-shot training set in a zero-shot manner, our CaFo can be evaluated under zero-shot settings the same as CLIP, for which none of the human-annotated training images is given. In Table 10, we report the best generated image number K′K^{\prime} of DALL-E for zero-shot CaFo. The number “0” denotes Zero-shot CLIP. For different datasets, the best number varies ranging from 1∼\sim16, and the larger number normally cannot get the better result, probably due to the low-quality synthetic images. On Caltech101 and EuroSAT , zero-shot CaFo largely surpasses CLIP by +4.62% and +7.54%, indicating our superiority under zero-shot settings.

Hyperparameter β𝛽\beta.

In Formula 5 and 6, we utilize a non-linear modulator φ(x)=exp⁡(−β⋅(1−x))\varphi(x)=\exp(-\beta\cdot(1-x)) for the affinity matrix of CLIP and DINO in the cache model, where β\beta controls the matrix sharpness. In Table 9, we experiment CaFo with different β\beta on 16-shot ImageNet and observe 0.6 performs the best.

Appendix C Additional Visualization

In Figure 12 and 13, we show more visualization of the prompts produced by GPT-3 and how they assist our CaFo to rectify false predictions of the original CLIP’s templates.

DALL-E’s Generated Images.

In Figure 14, we visualize more synthetic images generated by DALL-E on different datasets. Benefited from the pre-trained DALL-E, the generated images can well highlight the semantics of target category and effectively expand the few-shot training set in low-data regimes.

t-SNE.

We present the t-SNE visualization of our CaFo and the second-best Tip-Adapter-F in Figure 10. CaFo shows more contrastive distribution of category clusters and well mitigates some aliasing between similar classes.

Learning Curves.

In Figure 11, we visualize the 20-epoch learning curves of test accuracy on 16-shot ImageNet. Compared to the single CLIP, collaborating with DALL-E, DINO and GPT-3 significantly improves the convergence speed and classification accuracy on test set.