Neural Prompt Search

Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu

Introduction

The size of vision models has grown exponentially from tens of millions a few years ago (e.g., ResNet ) to today’s hundreds of millions , or even billions , for Transformers . Such an increase can cause a number of problems to transfer learning , and the first and foremost is that fine-tuning becomes more difficult as large model size can easily lead to overfitting in a typical-sized dataset, let alone the increase of compute and storage costs.

Recently, there is a growing interest in developing parameter-efficient tuning methods . The key idea is to insert a tiny trainable module to a large pre-trained model and only adjust its parameters by optimizing some task-specific losses like the cross-entropy for classification problems. The most representative methods are Adapter , Low-Rank Adaptation (LoRA) , and Visual Prompt Tuning (VPT) . As exemplified in Fig. 1(a), Adapter is a bottleneck-shaped neural network appended to a network block’s output; LoRA is a “residual” layer consisting of rank decomposition matrices; VPT prepends additional tokens to the input of a Transformer block, which can be seen as adding learnable “pixels.”

By evaluating the three parameter-efficient tuning methods on a commonly-used transfer learning benchmark, i.e., VTAB-1k , we identify a couple of critical issues. First, none of the three methods performs consistently well on all datasets, as illustrated in Fig. 1(b). For instance, when it comes to scene structure understanding tasks, VPT outperforms Adapter and LoRA on SmallNORB/ azimuth , but its performance plunges on SmallNORB/elevation and Clevr/count , which is largely behind the two competitors. The results suggest that, for a specific dataset, one needs to perform an extensive evaluation on different tuning methods in order to identify the most suitable one. Second, performance is found to be sensitive to the selection of model parameters, such as Adapter’s feature dimension or the token length in VPT—this is also observed by Jia et al. that the optimal token length in VPT varies from 1 to 200 on different datasets.

In this work, we view the existing parameter-efficient tuning methods as prompt modules and propose to automatically search for the optimal prompt design from data via a neural architecture search (NAS) algorithm. Specifically, we introduce the concept of Neural prOmpt seArcH (NOAH) for large vision models, particularly those equipped with the Transformer block . The search space is constructed by subsuming Adapter , LoRA and VPT into each Transformer block, as depicted in Fig. 1(a). The specific model parameters, including the feature dimension for Adapter and LoRA and the token length for VPT, are determined by a one-shot NAS algorithm.

We conduct extensive experiments on VTAB-1k , which is composed of 19 diverse vision datasets and covers a wide spectrum of visual domains like objects, scenes, textures and satellite imagery. The results show that NOAH significantly outperforms the individual prompt modules on 10 out of 19 datasets while the performance on the remaining is highly competitive (see Fig. 1(b) for an overview of the results). We also evaluate on few-shot learning and domain generalization where the results also confirm the superiority of NOAH to the hand-crafted prompt modules.

Our contributions are summarized as follows. (i) We present a systematic study of three representative prompt modules and expose some critical issues associated with performance and efficiency. (ii) A novel concept, neural prompt search, is proposed to address the challenge of hand-engineering prompt modules. (iii) An efficient NAS-based implementation of NOAH is provided. (iv) We demonstrate that NOAH is better than individual prompt modules in downstream transfer learning, few-shot learning, and domain generalization. The models and code will be released.

Neural Prompt Search

where 1d\frac{1}{\sqrt{d}} is a scaling factor.

Below we briefly review the three representative—and top-performing—parameter-efficient tuning methods, i.e., Adapter , LoRA and VPT , which will be incorporated into our search space. An illustration of these three methods can be found in Fig. 1(a). Note that VPT has been studied for vision models while Adapter and LoRA have only been studied for language models.

Adapter

LoRA

where ss is a fixed scaling parameter for modulating the updates.

Visual Prompt Tuning (VPT)

2 Prompt Search Algorithm

As discussed, none of the individual parameter-efficient tuning methods, or prompt modules called in this paper, shows dominance in the transfer learning benchmark. Our approach, neural prompt search (NOAH), incorporates Adapter , LoRA and VPT into each Transformer block and learns the design that best suits a dataset through neural architecture search (NAS). Specifically, we employ a one-shot NAS algorithm, AutoFormer , for prompt module search. Our supernet is a ViT-like model composed of 12 Transformer blocks (layers). Below we detail the search space and how the search is done.

As shown in Fig. 1(a), we embed the three prompt modules into each Transformer block following the guidelines proposed in the original work . Concretely, we install VPT in the input position, add LoRA alongside the two projection matrices as residuals, and insert Adapter after the normalized output of the MLP. The search space mainly contains the model parameters associated with the three prompt modules. Specifically, each prompt module has two sets of parameters to search from: (i) the embedding dimension ∈{5,10,50,100}\in\{5,10,50,100\} or {1,5,10}\{1,5,10\}; (ii) the depth ∈{3,6,9,12}\in\{3,6,9,12\}. A depth means up to which layer a module is applied, e.g., depth=3\text{depth}=3 for VPT means layers 0, 1 and 2 have VPT installed while the remaining layers, 3 to 11, do not have VPT.Zero-based indexing is adopted. For VPT, the embedding dimension means the token length whereas for Adapter and LoRA, the embedding dimension means the down-sampled dimension, i.e., rr.

Supernet Training

The supernet, as mentioned, has 12 Transformer layers, each containing the three prompt modules with full embedding dimension, i.e., 100. During each forward pass, a subnet is randomly sampled from the supernet for training. Specifically, for each prompt module, a depth is first sampled from {3,6,9,12}\{3,6,9,12\} to determine which layers should have the module. Then, for each layer within the depth range, an embedding dimension is chosen from {5,10,50,100}\{5,10,50,100\} or {1,5,10}\{1,5,10\}, all with a uniform probability. Note that only the prompt modules’ parameters are learned while the pre-trained model is kept fixed. AutoFormer allows the weights in each prompt module to be entangled during training, meaning that different weights are maximally shared, e.g., in a VPT module, if 100 tokens are selected for training, the previously trained tokens, such as 50, will be reused and trained together with other 50 tokens. This way, as suggested in AutoFormer , leads to faster convergence and low memory cost.

Evolutionary Search

After the supernet is trained, evolutionary search is conducted to obtain the optimal subnet architecture under a parameter size limit . Specifically, we first select KK random architectures, from which the top kk architectures (with the best performance) are used as parents to produce the next generation through crossover and mutation. For crossover, two candidates are randomly chosen and crossed to produce a “child” architecture. For mutation, a candidate mutates its prompt module design with a probability. See Fig. 2 for an illustration.

Experiments

In this section, we mainly address the following questions: (i) Is NOAH better than the individual prompt modules? (ii) Can NOAH work in a few-shot setting? (iii) Are models learned by NOAH robust to domain shift? The answers are discussed in Sec. 3.1, 3.2 and 3.3, respectively. We also conduct some analyses in Sec. 3.4 to have a deeper understanding of NOAH, such as what a subnet looks like and whether it is transferable beyond the dataset in which the architecture was found.

The main competitors are the three representative prompt modules subsumed by NOAH, which are Adapter , LoRA and VPT . Among them, only VPT is specifically designed for vision models while the other two are originally developed for language models. We also compare two common fine-tuning methods on the VTAB-1k benchmark: full tuning (Full) and linear probing (Linear). Full simply tunes the entire model parameters whereas Linear freezes the pre-trained part and only adjusts the newly added linear classification layer.Note that all methods have a new linear classification layer to learn. It is worth mentioning that Full has been considered as a strong baseline in existing studies .

Implementation Details

We keep the training parameters identical across all experiments throughout this paper. ViT-B/16 pre-trained on ImageNet-22K is used as the base model, which is strong enough so the results are fair and convincing. The supernet for NOAH is trained for 500 epochs and the ultimate subnet is trained for 100 epochs— note that “subnet” means the prompt modules/architectures. Since the AutoFormer algorithm allows a subnet to be used without retraining, we demonstrate later that the subnet found by NOAH without retraining is also comparable to the retrained one. The evolutionary search in NOAH takes 5 epochs in total and each step of random pick/crossover/mutation produces 50 new subnets. The probability for crossover and mutation is set to 0.2, which follows AutoFormer . The individual prompt modules, i.e., Adapter , LoRA and VPT , are constructed using the best recipes suggested by the original papers (also trained for 100 epochs; see the Supplementary for more details). The parameter sizes for Adapter, LoRA and VPT are 0.16M, 0.29M and 0.64M, respectively. For fair comparison, we set the upper-limit of parameter size of the final subnet in NOAH to 0.64M so the resulting size would be comparable to the baselines. More implementation details including image augmentation and other hyper-parameters are provided in the Supplementary.

1 Experiments on VTAB-1k

We choose the VTAB-1k benchmark to evaluate the transfer learning performance of our approach. VTAB-1k consists of 19 vision datasets, which are clustered into three groups: Natural, Specialized and Structured. The Natural group contains natural images that are captured by standard cameras and cover a broad spectrum of concepts including generic, fine-grained and abstract objects. The Specialized group contains images captured by specialist equipment for remote sensing (like aerial images) and medical purposes. The Structured group is designed specifically for scene structure understanding, such as object counting, depth prediction and orientation prediction. Each dataset in VTAB-1k contains 1,000 labeled examples, which are split into a train (80%) and a val (20%) set (the latter is used for hyper-parameter tuning), while the test data comes from the original test set. The final model used for evaluation is trained using the full 1,000 examples in each dataset. Top-1 classification accuracy is used as the performance measure.

Results

Table 1 presents the full results on the VTAB-1k benchmark. A high-level summary is shown earlier in Fig. 1(b). The average performance within each group is summarized in Fig. 3. We have the following observations.

Observation 1: Overall, NOAH is the best parameter-efficient tuning method. First and foremost, we demonstrate that searching for the optimal combination of the individual prompt modules works the best. This is evidenced by the 1% average gain over the strongest prompt module, i.e., LoRA. Given the diversity of the benchmark, the 1% average gain can be considered to be significant. It is also worth mentioning that Adapter was previously proved to be the best-performing prompt module in NLP , but in our study for computer vision tasks, LoRA takes over the seat. This further confirms that search is a better option than hand-engineering in practice.

Observation 2: NOAH slightly dims in the Specialized group. The results suggest that NOAH’s weakness seems to be in the Specialized tasks where the individual modules achieve the on-par performance: NOAH’s results are not too far from those of the competitors, e.g., NOAH’s 84.9 vs LoRA’s 84.6 on average. And while NOAH is superior on a Retinopathy, it lags on the other datasets, especially on the EuroSAT. Since the individual modules require a manual search over architecture and hyper-parameters, NOAH is more compelling.

2 Experiments on Few-Shot Learning

We choose five fine-grained visual recognition datasets, which include Food101 , OxfordFlowers102 , StandfordCars , OxfordPets , and FGVCAircraft . The categories in these datasets cover a wide range of visual concepts closely related to our daily life: food, plant, vehicle and animal. We follow existing studies to evaluate on 1, 2, 4, 8 and 16 shots, which are sufficient for observing the trend.

Results

The results are summarized in Fig. 4. In terms of the average performance, we can observe that: (i) In the low-data regime like 1 or 2 shots, NOAH, LoRA and Adapter perform similarly but VPT largely lags behind; (ii) NOAH shows clear dominance when more shots become available, e.g., with 16 shots the gap between NOAH and the runner-up is around 2%. By looking at the individual graphs, we can see that none of the individual prompt modules performs consistently well on all datasets, which, again, justifies that search is better than hand-engineering.

3 Experiments on Domain Generalization

Since domain shift is ubiquitous in real-world applications , we are interested to know how our search-based approach compares with the individual prompt modules in terms of domain generalization ability. Following prior studies , we first train a model on ImageNet (using 16 shots per category) and then directly test it on four other variants of ImageNet that undergo different types of domain shift. Specifically, the test datasets include (i) ImageNetV2 , which is collected from different sources than ImageNet but following the same collection protocol, (ii) ImageNet-Sketch , which is composed of sketch images of the same 1,000 classes in ImageNet, (iii) ImageNet-A , which contains adversarially-filtered images, (iv) ImageNet-R , which is a rendition of ImageNet. Both ImageNet-A and -R have 200 classes derived from a subset of ImageNet’s 1000 classes. All results are averaged over three random seeds.

Results

Table 2 compares NOAH with the three individual prompt modules. On ImageNet, which is the source dataset, the gap between NOAH and the individual modules is small, which is about 1%. However, on the four test datasets, NOAH demonstrates significantly stronger robustness than the baselines: over 6.8%, 4.8%, 5% and 5.2% improvements on -V2, -Sketch, -A and -R, respectively. The results, together with those from previous subsections, justify that our search-based approach is superior to the individual prompt modules.

4 Further Analysis

A key question to answer is: how does NOAH’s subnet, i.e., the ultimate architecture, look like. To make the results convincing, we visualize the average architecture—instead of individual ones—found within each group of VTAB-1k, as well as the global average over all datasets, in Fig. 5. The x-axis represents the network depth while the y-axis represents the embedding dimension. An intriguing observation is that Adapter and LoRA, across all groups (Fig. 5(a)), mainly appear in deep layers with the embedding dimension larger than four and reduced to zero in some shallow layers. In contrast, VPT can be found nearly in all depths (layers), but the dimensions vary significantly in different groups, which indicates different demands for VPT. For instance, in the Natural group (Fig. 5(b)), shallow layers need more VPT modules; but in the Structured group (Fig. 5(d)), deep layers need more VPT modules. Moreover, the co-existence of the three modules, especially in deep layers, suggests that they are complementary to each other—such a synergy is difficult to obtain by manual design. In summary, the observed high variances in the module designs strongly indicate that search is much more efficient than hand-engineering when it comes to developing parameter-efficient tuning methods.

Transferability of Subnet

As discussed previously, the subnet (i.e., architecture) found for different datasets differs dramatically. Here we study whether, or in what circumstances, the subnet found from one dataset can be transferred to another. To this end, we train NOAH on ImageNet and apply the ultimate subnet to the VTAB-1k benchmark where the model is retrained and evaluated. To measure transferability, we compare the ImageNet subnet with the dataset-specific subnets on VTAB-1k. Fig. 6 shows the comparisons. Overall, the gap between the ImageNet subnet and the 19 dataset-specific subnets on VTAB-1k is below 3%, meaning that NOAH has fair transferability. By digging deeper into the results, we find that the transfer gap is smaller when the source (i.e., ImageNet) and target datasets are closer, and vice versa. For instance, the gaps in the Natural group are less than 1%, which make sense because the ImageNet images and those from the Natural group share similar visual concepts, such as generic objects, flowers and animals.

With vs Without Retraining

Thanks to the weight entanglement strategy in AutoFormer , the subnet extracted from the supernet can be directly deployed for use without retraining. To verify if such a rule also applies to NOAH, we compare the subnets with and without retraining on VTAB-1k. The results averaged over each group are shown in Table 3 where we observe that NOAH without retraining (denoted as inherited) is still competitive: the inherited version still outperforms the VPT and Adapter. The results suggest that the retraining cost can be safely removed without incurring any significant loss.

Related Work

A recent trend in transfer learning is to develop parameter-efficient tuning methods , which is spurred by the rapid increase in model size. Existing methods can be generally divided into two groups. The first group fine-tunes a small portion of the internal parameters, such as biases . The second group adds tiny learnable modules like Adapter or LoRA , which is more relevant to our research and thus the focus here. Adapter and LoRA essentially share similar architectures—both look like a bottleneck—but are installed at different places: Adapter is often installed at the output of a block while LoRA is treated as residuals to the projection matrices in a Transformer block. It is worth noting that these methods are first studied in natural language processing (NLP) since pre-trained language models typically have an enormous parameter size that reaches the billion level. Another popular design in NLP is prompt learning , which turns some text prompt tokens into learnable vectors. Such an idea has recently been applied to vision-language models and is also the source of inspiration for the recently proposed VPT , which adds learnable “pixels” to the input of ViT .

More relevant to our work are those trying to unify different parameter-efficient tuning methods . He et al. build a connection between Adapter and prompt learning and cast the problem into the learning of a modification vector, which leads to a unified view and a stronger baseline. UNIPELT is another unified framework, which subsumes several prompt modules in a block and learns a set of gating functions to selectively activate them. Our work differs from these studies in two crucial ways: (i) we target computer vision problems whereas the previous studies focus on NLP; (ii) we unify prompt modules from the NAS perspective with a much more fine-grained control over the model hyper-parameters (e.g., token length and embedding dimension). This allows our model to be deployed in a resource-constrained environment. In the future, we plan to apply our approach to NLP.

2 Neural Architecture Search

Neural architecture search (NAS) consists of two crucial components: search space and search algorithm. A search space can subsume various designs of how neurons are connected , diverse combinations of model hyper-parameters , or different arrangements of specific modules like normalization layers . When it comes to the search algorithm part, the community has witnessed significant advances: from costly methods like reinforcement learning or evolutionary search to more efficient ones based on weight-sharing or differentiable optimization . The most relevant work to ours is AutoFormer , which is a one-shot NAS method focusing on Transformer models. AutoFormer features a weight entanglement strategy, which allows different subnets sampled from a big supernet to share weights among each other. Our work leverages AutoFormer to solve the problem of engineering parameter-efficient tuning methods, which we hope can inspire future work to address efficient transfer learning.

Discussion, Limitation and Future Work

With the proliferation of large-scale pre-training data , model size in neural networks has also been increased correspondingly in order to reach a certain learning capacity. On the other hand, the rapid increase in model size has also spurred interests in developing efficient transfer learning methods .

Our research presents timely studies on how some recently-proposed parameter-efficient tuning methods, or prompt modules, fare in computer vision problems. Crucially, our studies expose a critical issue that, for any specific downstream dataset, hand-designing an optimal prompt module is extremely challenging. More importantly, we for the first time solve the problem from a NAS perspective and demonstrate the potential of our search-based approach in terms of downstream transfer learning performance, the ability to work in low-data regimes, and robustness to domain shift, which is ubiquitous in real-world data .

Our studies also unveil some intriguing phenomena. In particular, we find that the ultimate subnet exhibits different architectural patterns for the three prompt modules across datasets of different natures. Since neural networks’ features, as often suggested , progress from low-level visual primitives in bottom layers to high-level abstractions in top layers, the aforementioned findings entail that different prompt modules work best for features at different levels. We hope such findings and insights can inspire future work on designing more advanced prompt modules.

In terms of limitations, NOAH requires additional training for the supernet, which inevitably increases the development cost. Moreover, as suggested by the few-shot learning results, NOAH’s advantages become clearer when more labeled images are available. In other words, NOAH would require more labels to unleash its full power in practice. For future work, we plan to dig deeper into the mechanisms behind NOAH for better interpretation of the intriguing results and apply NOAH to broader application domains beyond computer vision, such as NLP.

Appendix A Implementation Details

Table 4 briefly introduces dataset that we used.

A.2 Training Details

For the VTAB-1k , we follows its default augmentation settings, implementing the resizing and normalization for input images. Specifically, we resize a input image to 224×224224\times 224, followed by normalizing it with ImageNet means and standard deviation. For few-shot learning and domain generalization experiments, we implement color-jitters with the factor as 0.4, and RandAugmentation with magnitude equals 9, magnitude standard deviation equals 0.5.

Hyperparameters

We consistently set the embedding dimension of Adapter and LoRA as 8. We set prompt length of VPT following the paper. For few-shot learning (FS) and domain generalization (DG) experiments, we consistently set the VPT prompt length as 8. We use AdamW optimizer with the cosine scheduler. The weight decay equals 11e−3-3, warm-up epochs equals 10, and the batch size equals 64. Other hyperparameters are shown in below.

References