LION: Implicit Vision Prompt Tuning

Haixin Wang, Jianlong Chang, Xiao Luo, Jinan Sun, Zhouchen Lin, Qi Tian

Introduction

With the development of deep learning in the field of computer vision, models with stronger representations and larger sizes have been developed. Despite this, training these models with a large number of parameters is becoming increasingly challenging.

One common approach to addressing this issue is pre-training on a large dataset, such as ImageNet , for general vision tasks and then fine-tuning the model on downstream tasks to improve performance. While this method has been widely used, it has several drawbacks that should be considered. Firstly, fine-tuning requires a large amount of computational resources, especially for large models such as ViT-B (85.84M parameters) and Swin-B (86.87M parameters). Secondly, the model may become overfitted to the small target dataset and cannot be used for other tasks after fine-tuning. This means that separate sets of model parameters are needed for each task, leading to a high storage requirement.

In recent years, prompt-based learning has generated considerable interest in the natural language processing (NLP) community because of its amazing performance on a variety of downstream problems . Prompt tuning aims to design a trainable lightweight block as a supplementary input, which can guide or direct the generation of powerful vision representations to achieve desirable performances, rather than fine-tuning pre-trained models to adapt to downstream tasks.

To design a prompt block that combines both a lightweight architecture and strong representation capability, we conducted a comprehensive study and analysis of the limitations of current vision prompt tuning methods. First, existing approaches insert trainable networks as prompt blocks between each layer of the network , assuming that the feature representations from different levels contribute to the network’s generalization performance, especially for low- and mid-level representations. This, however, goes against the lightweight design philosophy of prompt tuning. Additionally, the architecture design is complex and heavily reliant on tuning skills, making it difficult to apply to various vision backbone models with different architectures. Second, finding the right prompt is a challenging task that often takes a significant amount of time. Even small changes in the activation input can have a significant impact on performance. This can be attributed to the depth of big model architectures, making the trainable parameters of shallow network layers more difficult to train and converge.

Based on the challenges above, we naturally raise a question, can we design a single-layer network with favorable convergence to iterate continuously? We hope it can achieve the effect of multi-layer network training and thus greatly reduce the training parameters. Therefore, we propose the impLicit vIsion prOmpt tuNing (LION), which is motivated by deep implicit models with stable memory costs for various complex tasks. In particular, we merely insect two equilibrium implicit layers in two ends of the pre-trained main backbone with parameters in the backbone frozen. LION enables the tuning of various vision models, including convolutional neural networks (CNNs) and vision transformers.

Specifically, LION constructs lightweight prompt blocks to generate task-specific discriminative prompts for each downstream input image. LION can generate a compact and robust downstream model that adapts tuning demands across a wide range of vision tasks while only training lightweight additional parameters per downstream task, by blending the input and the representations with the learned prompts. Besides, since the hyper-parameters like learning rate can severely effects the robustness of training the vision prompts, we prune the parameters in these two layers according to lottery hypothesis. Only the important parameters are kept in training to avoid the over-fitting. More surprising is that LION essentially compresses the parameters of the existing vision prompt network, which allows it to be generalized to any subsequent vision prompt tuning method with trainable parameters.

Our proposed LION can be used to tune CNN-based and Transformer-based vision models, surpassing fine-tuning for various recognition tasks of image classification, and can perform well under a variety of practical scenarios, including generic objects, class imbalance, and few-shot learning. Learning vision-specific information while maintaining the pretrained knowledge, our LION deliver an average improvement of up to 1.3% compared with VPT, but with much fewer trainable parameters. In summary, the main contributions of our work are three-fold:

We propose LION, a significantly lightweight yet effective vision tuning method which leverages a single trainable implicit layer to adapt pretrained model to downstream visual tasks.

We construct a more robust optimization strategy by lottery hypothesis, while proving the good convergence of our method through theoretical analysis and validation.

Experimental results show that LION outperforms previous vision prompt tuning methods, evidencing its feasibility and usefulness for vision models especially on few-shot and long-tail scenarios.

Related Works

With the development of deep neural networks, more and more works fine-tune the pre-trained models as the backbone for downstream tasks , which are trained on large-scale image dataset for common tasks like image classification. The use of fine-tuning is highly flexible: they can be applied to new input domains or to new tasks with different output semantics.

Early work in fine-tuning focused on adapting pre-trained models through layer-wise fine-tuning , where only a subset of layers in the network were fine-tuned, while the rest of the network remained frozen. This technique was shown to be effective in reducing the amount of computation required for fine-tuning and improving the stability of the optimization process. More recent work has focused on developing fine-tuning methods that allow for the adaptation of the entire network , rather than just a subset of layers. These methods often use regularization techniques, such as early stopping, to prevent overfitting and ensure the stability of the optimization process. Additionally, some of these methods have also explored the use of adversarial training to improve the robustness of the fine-tuned models. Another line of related work has focused on developing fine-tuning methods that can be used in a transfer learning setting , where the pre-trained model is adapted to a new task in a different domain. This has proven to be a challenging problem, as the pre-trained model may not be well-suited to the new task due to domain shift. To address this, researchers have proposed fine-tuning methods that incorporate domain adaptation techniques, such as domain-adversarial training.

2 Prompt-based Learning

Prompt-based learning is a technique that utilizes task-specific descriptions to enhance the understanding of downstream tasks by pre-trained models. This approach was popularized by the GPT series in the field of NLP and specifically, GPT-3 by treating each downstream task as a masked language modeling problem, where the model generates a textual response in the prompt. This has led to many studies focused on developing effective prompt strategies for extracting knowledge from pre-trained language models. Similarly, recent vision-language models , such as CLIP and ALIGN , generate visual concepts from natural language by training a large contrastive learning model over a large number of image-text pairs, using visual categories as prompts. This method has achieved impressive performance on various vision tasks without the need for fine-tuning. However, these prompt-based tuning methods are not suitable for pre-trained vision models. Our work aims to bridge this gap by developing a parameter-efficient prompt tuning approach specifically for vision models, to adapt frozen pre-trained vision models to downstream tasks across a broad distribution.

3 Deep Implicit Models

In recent years, there has been a growing interest in the use of implicit layers in deep learning. Researchers have explored different approaches to utilizing numerical analysis methods to replace the representation mechanism in existing deep networks. Some notable examples include SparseMAP , OptNet , and SATNet . These approaches have shown promising results in improving the efficiency and performance of deep learning models. One particular type of implicit model that has gained attention is Deep Equilibrium Models (DEQ) . DEQ is an implicit model with infinite depth, yet it is interesting as a single-layer network because it allows for analytical backpropagation through the equilibrium point. Regardless of the depth of the network, training and predicting with DEQ only require constant memory. Moreover, DEQ has been shown to achieve comparable performance with efficient memory cost, as illustrated in . Another advantage of DEQ is its interpretability. The use of implicit models can make it difficult to interpret the behavior of the network, but DEQ provides a transparent mechanism for understanding the network’s inner workings. These advantages make DEQ a suitable candidate for constructing light-weight vision prompt layers.

Methodology

Here, MLPMLP represents the fully-connected layer for representation projection. αi\alpha_{i} and βi\beta_{i} represent balance coefficients that are determined based on their importance in the attention mechanism . To generate these balance coefficients, we initialize two adjustable parameters, gαig_{{\alpha}_{i}} and gβig_{{\beta}_{i}}, subject to the following constraints: αi+βi=1{\alpha}_{i}+{\beta}_{i}=1 and 0<αi,βi<10<{\alpha}_{i},{\beta}_{i}<1.

In this way, the task-specific knowledge of the downstream data is distilled and effectively incorporated into the trainable parameter set: Θt=concat(Θp1∣∣Θp2∣∣Θh)\Theta_{t}=concat(\Theta_{p1}||\Theta_{p2}||\Theta_{h}), activating the pre-trained model. The prompt-based blending representation promotes a balance between the original representations and the learned prompt, resulting in adaptive control. Our prompt blocks perform a generic architectural modification that enables the pre-trained vision model to be adapted to a downstream task with only a few additional parameters Θt\Theta_{t}.

Intuitively, our method is illustrated in Figure 2. The whole process can be listed as follows: i) Blend the downstream input image with the vision prompt using an adaptive coefficient by feeding the downstream input image into the equilibrium layer. ii) Feed the combined image to the frozen pre-trained model to get the feature representations. iii) Input the downstream image into the equilibrium layer to produce vision prompts, and map the prompts to the representation size to create a combination. iv) Add a fully-connected layer to the top layers of the pre-trained model for the final prediction.

2 Implicit Prompt Design

The Deep Equilibrium (DEQ) Model, as described in , employs a single layer that finds the fixed point of an iterative procedure. This layer, being capable of expressing the entire deep network as an equilibrium computation, is just as powerful as multiple stacked explicit layers. Our proposed method, LION, leverages this capability by using a single DEQ layer to adapt a pre-trained vision model to downstream tasks through implicit layer training. The downstream inputs are equipped with a single implicit layer, implemented as a ResNet layer , which is trained to learn task-specific prompts, while the pretrained model remains frozen. Drawing inspiration from previous studies , we aim to improve the feature representation ability by utilizing both high-frequency and low-frequency information as the vision prompt. To attain this goal while reducing parameter storage burden, we propose the use of two-level prompt blocks, located prior to the input layer and the head layer. Our design of this lightweight architecture is supported by the demonstration of its convergence ability in Section 3.5.

Unlike a conventional network where the output is the activations from the LL th layer, the output of a equilibrium layer is the equilibrium point itself. Therefore, the forward evaluation could be any procedure that solves for this equilibrium point. Conventional deep neural networks, if they converge to an equilibrium, can be considered a form of fixed-point iterations for the forward process:

Our goal will be to compute the vector-Jacobian product ∂z∗(⋅)∂(⋅)Ty\frac{\partial z^{*}(\cdot)}{\partial(\cdot)}^{T}y for some vector yy, where (⋅)(\cdot) here is a stand-in for any quantity we want to differentiate the fixed point with respect to (i.e, the input xx, or any parameters of the function P\mathcal{P}, both of which of course will affect the final fixed point z∗z^{*}). Since this vector-Jacobian product is the key aspect to integrating these DEQ layers within backpropagation, such a routine allows us to integrate the DEQ layer within standard automatic differentiation tools.

The derivation of the vector-Jacobian product largely mirrors that in previous chapters, but we include the full derivation again here for completeness. Differentiating both sides of the fixed point solution, we have:

In order to calculate the vector-Jacobian product, we need the following information:

The key term of interest here is the solution in linear system, we can utilize the proxy o=(I−∂P(z∗,x)∂(z∗))−Tyo=(I-\frac{\partial\mathcal{P}(z^{*},x)}{\partial(z^{*})})^{-T}y to solve the fixed point equation and compute the final Jacobian vector product.

3 Robust Training

To overcome the instability during prompt tuning, we utilize the robust training mechanism for trainable parameters Θt\Theta_{t} motivated by . The Lottery Ticket Hypothesis posits that only a subset of a model’s parameters are crucial for achieving good generalization. The rest are likely to overfit. To address this, we employ a criterion to separate crucial and non-crucial parameters and optimize them differently. Crucial parameters are updated with stochastic gradient descent while non-crucial parameters are constrained to reduce their ability to overfit.

Recent pruning methods suggest that crucial parameters should have substantial magnitudes, as they play a key role in network propagation. Furthermore, early optimization shows that parameters with large gradients tend to contribute to generalized patterns , which are crucial for learning from clean samples. Therefore, both the values and gradients of the parameters should be considered when determining their importance. To capture this, we use the product of their values and gradients as the criterion for determining criticality. In mathematical terms, the criticality of parameter θt\theta_{t} from Θt={θti}i=1M\Theta_{t}=\{\theta_{t}^{i}\}_{i=1}^{M} is represented as:

where L\mathcal{L} represents the loss function.

The equation shows that when the gradient or value of a parameter is close to zero, its criticality is low, making it a non-crucial parameter that is prone to overfit. On the other hand, when the value of z(θt)z(\theta_{t}) is large, θt\theta_{t} is considered a crucial parameter for learning basic and generalized patterns. To control the number of crucial parameters, we introduce a threshold τ\tau, such that crucial parameters are selected and represented as:

Unlike pruning methods, we do not eliminate the non-crucial parameters Θtn\Theta_{t}^{n}, instead we adopt a distinct optimization strategy. Here, the crucial parameters are updated in the usual way, while the non-crucial parameters are restricted to converge to zero for better generalization. Mathematically, the update rule for θt∈Θtc\theta_{t}\in\Theta_{t}^{c} is represented as:

The symbol η\eta denotes the learning rate in the above equation. For the non-crucial parameters, we shrink them by means of strict regularization, instead of minimizing the loss. In other words, the update rule for θt∈Θtn\theta_{t}\in\Theta_{t}^{n} is:

4 Optimization

For our training process, we keep the pre-trained parameters Θf\Theta_{f} intact and only modify a limited set of parameters Θt\Theta_{t}. This selective update of parameters makes our LION modular and efficient - it allows us to utilize an existing pre-trained vision model without having to modify or re-train it. Instead, we add a small number of additional parameters specific to each task, which can be formulated as:

With the pre-trained model frozen, we minimize the prediction error using cross-entropy (CE) loss. The ability of our LION to adapt a pre-trained vision model to a wide range of tasks while maintaining a high level of accuracy makes it a highly attractive solution for deployment in cloud services. The potential benefits of reduced computational and storage overhead, as well as the ability to offer real-time adaptation to new tasks, make our method an ideal choice for cloud service providers seeking to improve their offerings.

5 Theoretical Analysis

Note that we have only two prompt blocks in different position and architecture. We show that we can directly add the vision prompts to the first layer of the pre-trained model, while it cannot be implemented on the last layer. We should utilize another MLP block to ensure the optimal solution. Here we assume the σ\sigma to be ReLU: σ(xi)=max(xi,0)\sigma(x_{i})=max(x_{i},0). Given that the model is pre-trained with parameters θ^:(v^pre,W^pre)\hat{\theta}:(\hat{v}_{pre},\hat{W}_{pre}) which reach the optimal solution L(v^,W^)=0\mathcal{L}(\hat{v},\hat{W})=0. With the single-layer DEQ as the vision prompts network, we can derive the vision prompts with the suppose that: xpro=Axx_{pro}=Ax and zpro=Bzz_{pro}=Bz for some invertible matrix A,BA,B, where the corresponding label is unchanged: ypro=yy_{pro}=y.

There exists the vision prompt xpro=Axx_{pro}=Ax for invertible AA and ypro=yy_{pro}=y that can minimize the population loss: minWL(v^,W)=0min_{W}\mathcal{L}(\hat{v},W)=0. However, the vision prompt zpro=Bzz_{pro}=Bz may not be sufficient: there exists such BB such that the population loss is non-zero for any choice of the parameter vv: minvL(v,W^)>0min_{v}\mathcal{L}(v,\hat{W})>0.

Let B^,v^\hat{B},\hat{v} can reach the optimal solutions so that y=v^σ(W^x)y=\hat{v}\sigma(\hat{W}x) for all x,yx,y. Let W=W^A−1W=\hat{W}A^{-1}, we have for all xprox_{pro}

Therefore, the parameters v^,W\hat{v},W achieves L(v^,W)=0\mathcal{L}(\hat{v},W)=0.

Following is a counterexample showing that last-layer prompts are impossible. Since σ\sigma is the element-wise ReLU function, W^z\hat{W}z has only positive entries for all zz. Let B=−IB=-I, which is an invertible diagonal matrix full of -1. Then for any vv, we have vσ(W^zpro)=vσ(−W^x)=0v\sigma(\hat{W}z_{pro})=v\sigma(-\hat{W}x)=0, so the expected loss is positive. Therefore, minvL(v,W^)>0min_{v}\mathcal{L}(v,\hat{W})>0. ∎

6 Complexity Analysis

Experiments

To fully evaluate the performance of our proposed method, we implement detailed experiments on image classification from three categories of settings, including regular setting, long-tailed distribution setting, and out-of-distribution setting.

Dataset. CIFAR10 is a dataset of 60,000 32x32 color images in 10 classes, with 6,000 images per class. There are 50,000 training images and 10,000 test images. CIFAR100 is a dataset of the same size as CIFAR10, but it has 100 classes containing 600 images each. There are 500 training images and 100 testing images per class. ImageNet100 is a subset of the ImageNet dataset, which contains 100 classes of natural images. Each class has between 500 and 1000 images for training and 50 to 100 images for testing. Flower is a dataset of images of flowers from 5 different species. It contains 4242 images with 80-90 images per class. Stanford Dogs is a dataset of images of 120 breeds of dogs, with a total of 20,580 images. Stanford Cars is a dataset of images of cars, with a total of 16,185 images of 196 classes of cars. The dataset is organized by make, model, and year. Clothing is a dataset which contains images of various types of clothing.

2 Baselines

We compare our proposed LION to several commonly used protocols in order to provide perspective on our results. These protocols include: 1) Retraining, which trains the entire vision model from scratch; 2) Head Fine-tuning, which fine-tunes the last layers of a pre-trained model while freezing the remaining layers and retraining the head classifier; 3) Fine-tuning, which adjusts the weights of a pre-trained model and retrains the head classifier; 4) Adapter , which adds a new adapter structure to the transformer and updates only its parameters; 5) Bias , which updates only the bias terms of the parameters; 6) VPT , which fine-tunes the model by incorporating prompts as input tokens.

3 Implementation Details

For a fair comparison, we use the model pretrained on ImageNet as the initialization for the following tuning. Additionally, we extend our method to include CNN-based (ResNet-50, ResNet-101 ) and Transformer-based (ViT , Swin Transformer ) backbones. The implementation of baselines for additional backbones involves leveraging the core idea presented in the original paper, and adapting it to suit the specific capabilities of each new backbone architecture. In experiments on the aforementioned datasets, we utilize the Adam optimizer with a momentum of 0.9, batch size of 64, and learning rate of 1e-5. The whole experiments are implemented on the NVIDIA V100 GPU with PyTorch.

4 Performances

Image Classification. Quantitative results can be seen in Table 1 and Table 2. It can be observed that the proposed LION performs the best overall on all six tasks when using all the four models as the backbones. The best performance for each task is highlighted in bold. For example, on CIFAR100, LION achieves an accuracy of 54.84 %\% when using ResNet-50, and 58.98 %\% when using ResNet-101, both of which are the highest among all the methods. For the transformer-based (i.e., ViT-B and Swin-B) methods, LION achieves accuracy improvement of 1.36 %\%, 0.68 %\%, 0.51 %\%, 3.31 %\%, 0.38 %\% and 0.19 %\% over other tuning methods, respectively on CIFAR10, CIFAR100, ImageNet100, Flower, Dogs, and Cars datasets.In terms of the maximum number of trainable parameters, the proposed method has the lowest value among all the methods, with only 0.097 M trainable parameters. It is also worth noting that the performance of the other methods varies depending on the task and the backbone model used. For example, while the VPT method performs well on CIFAR100, it is not as effective on other tasks. On the other hand, the adapter method performs relatively well on the Flower and Dogs tasks, but not as well on the other tasks. In summary, the proposed LION is the most effective tuning method across all tasks and backbone models, with the added advantage of having the lowest number of trainable parameters.

Long-tail Class Distribution. We perform experiments on benchmark datasets that have a long-tail class distribution, such as CIFAR10-LongTail and CIFAR100-LongTail. The results of the imbalance ratio 50 and 100 are shown in Table 3. Our proposed LION algorithm outperforms the best baseline, VPT under all the settings. Specifically, we observe a gain of approximately 4% in validation accuracy, while reducing the trainable parameters by 88%. This trend holds across other settings as well. Additionally, when compared with VPT on long-tailed CIFAR-10 with an imbalanced ratio of 100 under ViT-B, LION achieves superior validation accuracy using only 4.2x fewer trainable parameters. When evaluated under Swin-B, LION outperforms VPT on long-tailed CIFAR-10 with imbalanced ratios of 50 and 100, while also reducing the trainable parameters by 64.7%.

Few-shot Learning. We perform experiments on benchmark datasets, such as Pets , Food-10 , Cars , and Flower with 8 examples per class, which are widely used for evaluating few-shot learning algorithms. Our experimental results, shown in Table 3, demonstrate that LION achieves state-of-the-art results on average, while using the fewest trainable parameters. It is worth noting that although VPT outperforms our method on some datasets for the image classification task, we still achieve the best performance on all datasets in this scenario, which highlights our method’s advantage. These observations confirm the capability and efficiency of our method in the low-data regime and further verify the effectiveness of the lightweight implicit vision prompt design.

5 Ablation Study

With the ImageNet pre-trained ResNet-50 and Vit-B, we perform extensive ablation studies to systematically analyze the developed LION. We introduce three model variants as follows: (1) LION-P1 removes the equilibrium layer in the front of the pre-trained backbone to activate low-level features, i.e., P1\mathcal{P}_{1} in the Eq. 1. (2) LION-P2 removes the equilibrium layer behiend the pre-trained backbone to activate high-level features, i.e., P2\mathcal{P}_{2} in the Eq. 2. (3) LION-R removes the robust training mechanism with the normal optimization for all the parameters, i.e., optimization in Eq. 8 and Eq. 9. The results of these model variants are summarized in Table 4. We have the following observations. First, our LION outperforms LION-P1 and LION-P2, which indicates that implicit vision prompt blocks work for vision semantic information activations. Second, LION-P2 obtains much worse than LION, which shows that the high-level activation is significantly important for the vision prompt tuning. Third, the robust training mechanism allows better network optimization, as shown by LION-R’s lower performance of 2% compared to LION, validating the superiority of robust training for better optimization.

6 Sensitivity Analysis

In Figure 3, we investigate the sensitivity of two hyper-parameters: τ\tau in the robust training mechanism, which separates the crucial and non-crucial parameters in Eq. 8 and Eq. 9. And we also investigate the number of the layer in DEQ by stacking the single layer attracted by its striking performances. We first vary τ\tau in {0.2, 0.4, 0.6, 0.8} with other parameters fixed. As τ\tau rises, the performance first increases and then decreases a little. The potential reason is that too large τ\tau could filter the essential parameters for optimization. We can observe that the performance of ours is not sensitive to τ\tau in the range of [0.4,0.6][0.4,0.6] and we can set it to any values in that interval. Further, we the number of layers from 1 to 4 with other parameters fixed. Obviously, our method can achieve a small performance gain with the layer ranging from 11 to 44 but at the cost of several times the computation time. Therefore, τ\tau and the layer number are set to 0.40.4 and 11 as default respectively.

Conclusion

In conclusion, this paper proposes an efficient vision model named LION that addresses the heavy computational costs associated with ViT. By drawing inspiration from deep implicit models with stable memory costs, LION only requires two equilibrium implicit layers in two ends of the pre-trained main backbone with parameters in the backbone frozen. Additionally, pruning the parameters in these two layers according to the lottery hypothesis reduces the number of training parameters. LION can obtain higher performance with smaller parameter size compared to the state-of-the-art baseline VPT, especially under challenging scenes. Our experiments demonstrate that LION has a good generalization performance, making it an easy way to boost applications in the future. Overall, LION provides an economic solution for vision tasks and is promising for a wide range of datasets.

References