Visual Tuning

Bruce X. B. Yu, Jianlong Chang, Haixin Wang, Lingbo Liu, Shijie Wang, Zhiyu Wang, Junfan Lin, Lingxi Xie, Haojie Li, Zhouchen Lin, Qi Tian, Chang Wen Chen

Introduction

Since the wide usage of Transformer models and the emergence of large foundation models , the paradigm of deep learning in vision intelligence has been experiencing the hype of adapting downstream tasks to the foundation models. The astonishing performance of the recent Visual ChatGPT is enabled with myriad computation resources in the pre-training process , and human feedback during the tuning process. The pre-trained foundation model (i.e., GPT-3) shows strong capability but entails large storage space, around 800GB, to store 175B parameters , which makes it expensive to retrain independent model copies for different downstream tasks. Foundation models are expected to continue to scale up, and how to reuse the foundation model via parameter-efficient transfer learning (PETL) methods (e.g., prompt, prefix, adapter, etc.) quickly becomes a research hype. In the past two years, taking the inspiration of PETL methods in natural language processing (NLP) , numerous visual tuning techniques have been proposed for adapting downstream tasks to pre-trained vision or visual-language models.

In the era of increasingly large models, vision models have been scaled up from EfficientNet-based (480480M parameters) to Transformer-based (2,1002,100M parameters) ones that recently reach increasing scales such as 2222B parameters and 562562B . For such large models, PETL methods aim to make good reuse of the shared parameter weights (usually interpreted as the knowledge of large models) deployed on the cloud to save storage overhead and to empower edge devices such as autonomous vehicles, drones, and robots that are intensive in computing and battery resources . This practice is different from the modus operandi of transfer learning that either fully fine-tunes the whole model or just fine-tunes the task head (e.g., the last fully connected layer) .

Given the emergence of increasingly large models (i.e., foundation models), we are in a new paradigm of visual tuning that is beyond tuning the entire model or the task head. How to effectively reuse the knowledge thereof with PETL methods, leading to less memory usage and higher inference speed is a hot topic in various vision tasks . Starting with a detailed background in Section 2, this paper provides an in-depth review of recent tuning advances in the vision domain, categorizing them into five common types and elaborating their current technical state with discussions in Section 3. Last but not the least, we provide insights into future research directions that hold significant promise in Section 4, followed by a conclusion. To the best of our knowledge, this is the first comprehensive survey on visual tuning, which bears great importance for researchers to understand the mechanisms and trends of this practice.

Background

In the early days, machine learning methods relied on feature engineering such as SIFT , BRIEF , and ORB to handle specific tasks, which is later on dominated by the deep learning paradigm since the introduction of ImageNet . Deep learning models pre-trained on ImageNet are able to benefit various downstream vision tasks such as image recognition, object detection, and image segmentation via fine-tuning. Fine-tuning is the second step of typical transfer learning, which makes use of the knowledge acquired from the source domain to facilitate the learning process of the target domain .

According to the principles of transfer learning , downstream tasks can greatly benefit from the knowledge learned via pre-trained models. For instance, most state-of-the-art baselines on large-scale vision benchmark datasets such as ImageNet and Kinetics rely on extra data. Early pre-trained models such as CNN-based and Transformer-based can be pre-trained via supervised or self-supervised learning. Recently, towards unifying natural language and visual understanding tasks, advanced pre-training paradigms such as CLIP , ALIGN , GATO , CoCa , FLAVA , Flamingo , and SWAG have been proposed, which are known as foundation models as a whole. Given the promise of large-scale pre-trained models, visual tuning techniques beyond fine-tuning have attracted increasing research interest, leading to the visual tuning paradigm as illustrated in Fig. 1. This section will elaborate on the background of visual tuning from five perspectives: theory, definition, model architecture, model pre-training, and model tuning.

In the 1990s, the machine learning community largely ignored neural networks and backpropagation due to concerns about overfitting and the potential for poor local minima. However, in the recent era of deep learning, these concerns have been greatly alleviated via advancements in theories and empirical experiences . In this section, we present the fundamental theories that underpin the current state of visual tuning, exploring these theories from three distinct perspectives as follows.

Like the origin convolutional neural network (CNN) architecture was inspired by the receptive fields in the visual cortex , learning models can be inspired or motivated by biological and neuroscience discoveries. Meanwhile, many computer vision problem definitions draw inspiration in some ways from human perception capabilities. The developments of neuroscience and computer vision bring mutual insights from each other. For example, recent empirical neuroscience findings suggest that backpropagation is related to the synaptic updating mechanism of the human brain, bringing possible insights for understanding the exact learning mechanism in the cortex . However, it remains far from fully understanding the human brain’s learning mechanism when researchers are working toward more explainable and interpretable AI. For instance, Ullman et al. conducted experiments to investigate the atom of recognition in human and computer vision, suggesting that existing deep learning models failed to learn minimal image changes (e.g., reducing image resolution and cropping part of the image) at the human level.

Particularly, researchers are working on enabling computer vision to have capabilities that are similar to human vision. First, human vision can efficiently process huge amounts of continuous visual streams. Regarding the intrinsic mechanism of this ability, classical biological findings suggest that humans perceive real-world scenes by contextualizing information from local parts (such as small edges) as a whole (i.e., subjective contours), which are respectively handled by cortical areas V1 and V2 . It is also suggested that human vision is embodied and developed in interactive ecological environments . This motivates researchers to work on effective solutions concerning the aspects such as accuracy and efficiency. Second, humans are good at generalizing visual understanding to unseen brand-new scenes by reasoning their physical and geometric properties . This motivates emerged foundation models to be tested via increasingly challenging setups such as zero-shot learning, continual learning, multi-task learning, etc.

1.2 Model Perspective

Taking inspiration from human vision, the recently emerged paradigm of tuning large vision models aims to effectively reuse the knowledge in the large pre-trained model in an efficient way regarding computation and data. Generative pre-trained large language models such as GPT-3 show significant continual performance improvements when the model size is scaling up from 0.1B to 175B parameters . This observation is known as the scaling up law: larger pre-trained model will benefit downstream tasks, which shows insights that adapting from a larger knowledge base can lead to better performance for downstream tasks. This scaling up law has also been proved in recent literature . Sung et al. elaborated on the reason for the reduced training memory of PETL techniques from the perspective of backpropagation and further reduced their training memory by skipping the gradient traversal through the frozen backbone, which steps further on the analysis of existing PETL technique regarding training and inference memories.

1.3 Statistical Perspective

Machine learning models are restricted by some statistical assumptions such as independent and identically distributed, the law of large numbers, central limit theorem, etc. , making machine learning practitioners conduct regularization techniques, collect large-scale datasets, and normalize the input data, respectively. In the era of large models, breaking these statistical boundaries becomes imaginable with encouraging recent progress (surveyed in Section 3), which intrinsically improves models’ generalization ability to out-of-distribution or long-tail data with less training data (from few-shot to zero-shot learning) and tunable parameters. To guide tuning with statistical rules, there are some works proposed based on measurable domain bound. For instance, Ye et al. proposed the concept of expansion function, quantifying regularization or bound restrictions as the “variation” between the source and target domains, and the “informativeness” of a feature. Liu and Zhang also attempted to measure the domain gap by using the test error. Zhang et al. proposed to use the margin loss to replace the 0-1 loss for domain adaptation. The margin loss is expected to relax the restriction and provide a more informative generalization bound. Nilesh et al. defined task diversity from a statistical perspective, providing generalization upper bounds of sample complexity for multi-task transfer learning.

2 Notation and Definition

In order to understand efficient fine-tuning, let’s start by defining domains, tasks, transfer learning, and other notations. A joint distribution X×Y\mathcal{X}\times\mathcal{Y} can be expressed as P(X,Y)P(X,Y) (i.e., PXYP_{XY}), where X\mathcal{X} and Y\mathcal{Y} represent its corresponding feature space and label space, respectively. ( XX and YY represent the observed instance set and its corresponding label set. ) Given PXYP_{XY}, we refer P(X)P(X) (i.e., PXP_{X}) as the marginal distribution on XX, PY∣XP_{Y|X} the posterior distribution of YY, and PX∣YP_{X|Y} the class-conditional distribution of XX given YY.

Definition 1 (Domain): A domain D={X,P(X)}\mathcal{D}=\{\mathcal{X},P(X)\} is defined by its feature space X\mathcal{X} and a marginal distribution P(X)P(X), where XX denotes an instance set defined as X={x∣xi∈X,i=1,...,n}X=\{x|x_{i}\in\mathcal{X},i=1,...,n\}. A domain can be with or without labeling information.

Definition 2 (Task): A task can be denoted as T={Y,f}\mathcal{T}=\{\mathcal{Y},f\}, where Y\mathcal{Y} and ff represent a label space and a decision function, respectively. For the classification task of a source domain TS\mathcal{T^{S}}, the goal is usually to predict the conditional distribution of instances, which can be denoted as f(xj)={P(yk∣xj)∣yk∈Y,k=1,...,∣Y∣}f(x_{j})=\{P(y_{k}|x_{j})|y_{k}\in\mathcal{Y},k=1,...,|\mathcal{Y}|\}. In this case, the task TS\mathcal{T^{S}} can be regarded as forming a typical source domain DS\mathcal{D^{S}} with labeling information, being denoted as (DS,TS)={(x,y)∣xi∈XS,yi∈YS,i=1,...,nS}(\mathcal{D^{S}},\mathcal{T^{S}})=\{(x,y)|x_{i}\in\mathcal{X}^{S},y_{i}\in\mathcal{Y}^{S},i=1,...,n^{S}\}.

3 Model Architecture

Pre-trained foundation models for vision have been surveyed in , which develops from CNN- and GAN-based models to recent Transformer-based models. We recommend readers refer to for the detailed pre-training strategies. This section briefly introduces these representative models’ basic structures: CNN-based, Transformer-based, and CNN+Transformer.

CNN is one of the most popular deep learning models such as AlexNet , VGGNet , Inception , ResNet , EfficientNet , etc., which has been surveyed time to time . EfficientNet is lightweight yet can achieve comparable performance to Transformer-based models via pre-trained initialization on various visual tasks such as image classification and video understanding . Except for 2D CNN, a couple of 3D CNN models such as C3D , I3D , S3D , and X3D have been introduced for video understanding tasks. In addition, temporal convolutional network has also been proposed for tasks such as segmentation and pose estimation . Notably, the basic layer of the temporal convolutional network is implemented with 2D convolution while Transformer is implemented via 1D convolution in addition to the attention mechanism.

The typical architecture of a Transformer model is structured with several basic Transformer layers. Each layer can be made of a varied number of Transformer blocks composed of a multi-head self-attention (Attention) module and a fully connected feed-forward network (FFN) implemented with a 2-layer multilayer perceptron (MLP). Layer normalization (LN) and residual connection are respectively performed before and after both FFN and Attention modules. One such Transformer block can be represented as:

where Z^l\hat{Z}^{l} and ZlZ^{l} respectively indicate the output of Attention and FNN modules, and Zl−1Z^{l-1} denotes the output of a previous FNN module. Building upon the basic Transformer model, Transformer has been dominating increasing tasks . Early Transformer models for vision are Vision Transformer (ViT) , Data-efficient image transformer (DeiT) , while their representative variations are TNT , T2T , PVT , Swin-Vit , Video Swin Transformer , CPVT .

Transformer models are well known for their ability to capture long-range dependencies of input data. Whereas CNNs might be better at representing local features. Models combining Transformer and CNN can achieve better performance. Twins-SVT used 2D convolutional to calculate the attention module, leading to improved performance on image-based tasks with considerably more model parameters. Representative methods combining CNN and Transformer are Shuffle , CMT , VOLO , etc. Although they can achieve superior performance, how they can be used for visual tuning seems under-explored.

4 Model Pre-training

Pre-training methods can be roughly grouped into supervised and self-supervised ones. Early vision models were pre-trained via supervised learning on large-scale datasets such as ImageNet , JFT-300M , Kinetics , etc. Since fine-tuning models pre-trained with supervised learning, larger-scale pre-training has been conducted recently. For example, Gato uses multi-task learning with the supervision of varied tasks to enable the large model to acquire more knowledge for the adaptation of downstream tasks. Multi-label learning is used to pre-train a pure vision model that reaches 22B parameters , showing fantastic performance on downstream visual tasks.

In the regime of supervised pre-training, the non-trivial annotation cost imposes a practical obstacle to scaling up the benefit of transfer learning. Alternatively, self-supervised learning on unannotated data can also make the models richer and potentially more useful . The paradigm of fine-tuning models pre-trained via self-supervision brings the possibility of learning knowledge from unannotated data at a larger scale, which is enabled by advanced computing power, the Transformer model, and more data. Models pre-trained with self-supervised learning are termed “foundation models” by Bommasani et al. . Recent notable examples include MAE in vision; CLIP , ALIGN , Florence , BEiT , GATO , CoCa , SWAG , etc. for visual-language models. Taking the initial success in NLP, this paradigm has started showing success in vision and various other realms such as climate science , protein design , etc. Bommasani et al. identified the key significance of foundation models as emergence regarding capability and homogenization regarding model, modality, tasks, and domains.

5 Model Tuning

Given the knowledge learned via pre-trained models, downstream tasks can greatly benefit from them. Early modus operandis of fine-tuning includes updating the whole parameters of the pre-trained model and tuning the task head only (e.g., fully connected layer). With the popularity of large language models such as GPT-3 pre-trained via meta-learning in an unsupervised manner, enabling them to handle a broad set of skills (the inner loop termed “in-context learning”). Given the ability of multiple skills, the current leading paradigm in NLP is to adapt downstream tasks to the large language models, entering the learning paradigm of “pre-train, prompt, predict” from “pre-train, fine-tune” .

On the one hand, a couple of recent works achieve promising performances on vision downstream tasks by fine-tuning visual-language models. However, according to the results in and , fine-tuning visual-language models do not lead to results as good as fine-tuning supervised pre-trained vision models. In addition, pure vision models are also increasingly large (reach 22B parameters) and gain great advances recently with varied pre-training strategies . As such, it needs further investigation on proper pre-training techniques and fine-tuning techniques for vision downstream tasks.

Visual Tuning

To the best of our knowledge, there is no survey that systematically summarizes the recent state of visual tuning from the technical perspective. He et al. analyzed different PETL tuning methods such as prompt-tuning, prefix-tuning, and adapters in the NLP domain, showing they are intrinsically similar (i.e., they bring a certain amount of tunable parameters for adaptation). Taking parameter efficient transfer learning methods in NLP into consideration, we group visual tuning methods into five categories: fine-tuning, prompt tuning, adapter tuning, parameter tuning, and remapping tuning (see Table 2) according to their structures and motivations. In the remainder of this section, we introduce the five groups of tuning techniques with discussions of their advantages and disadvantages.

We use fine-tuning to denote the standard practice of transfer learning, which either tunes the whole parameters of pre-trained models or just tunes the task head. Many state-of-the-art methods adopted this practice to achieve impressive performance on vision benchmarks such as ImageNet , Kinetics , COCO , NTU RGB+D 120 , Human3.6M , etc. Tuning the whole pre-trained model parameters intrinsically initiates the learning process of the downstream tasks via the learned model weights. While tuning the task head treats the pre-trained model as a feature extractor.

The full fine-tuning strategy comes with obstacles for adapting large models to downstream tasks. First, it requires one to update and store separate model parameters for different downstream tasks, which can be expensive and infeasible when the foundation models become increasingly large. Second, it relies on high-quality downstream data and can hardly adapt to unseen scenarios that have large distribution shift , which is unlike the learning process of humans who can learn from few samples and generalize well to new circumstances. This issue has been researched in directions such as zero-shot learning, few-shot learning, and continual learning . Alternatively, fine-tuning the downstream task head can avoid updating the entire backbone model, but it usually leads to unsatisfactory experimental performance.

2 Prompt Tuning

Prompt-based learning is first introduced in NLP to efficiently adapt downstream language tasks to foundation models. Unlike the traditional “pre-training, fine-tuning” paradigm which initializes the weight parameters of pre-trained model and optimizes these parameters under the guidance of downstream task-specific loss functions, prompt-based learning leverages textual prompts to reformulate various downstream tasks as the original pre-trained task. Inspired by prompt techniques in NLP, prompt tuning is also introduced into the computer vision field. Specifically, vision prompt tuning could be divided into three groups, i.e., vision-driven prompt, language-driven prompt, and vision-language prompt.

Vision-driven prompt tuning has become a popular parameter-efficient way to transfer the remarkable generalization ability of pre-trained vision models to various downstream tasks. The research efforts of vision-driven prompt strategies can be roughly categorized into two groups, i.e., modifying inputs directly, and designing vision prompt sub-networks to produce vision prompts. Studies of the first family usually tend to directly modify inputs, e.g., adding a set of learnable parameters into input images, which aims to modify the input distribution and further makes downstream tasks close to the solved task during the original pre-training, as shown in Fig. 2(b). Formally, the mathematical formulation can be described as:

where PV\mathcal{P}_{V} indicates the vision-driven prompts, T\mathcal{T} denotes the embeddings of local images or tokens outputted by Transformer, and aia_{i} is the ii-th learnable vector.

Existing extensive works utilize the above principle to design vision prompts to instruct frozen vision pre-trained models to various downstream tasks. Concretely, VPT plugs solely a few learnable parameters and regards these parameters as a part of input tokens of Transformer, which steers pre-trained vision models to perform various downstream tasks. Similar to VPT, DePT also introduces learnable visual prompts into the vision Transformer and only optimizes these source-initialized prompts while keeping the vision Transformer frozen during adaptation. In addition, PViT designs task-specific prompts by introducing a small set of specialized parameters to adopt a shared video transformer backbone to perform synthetic scene tasks and a real video downstream task. Moreover, ZegCLIP injects learnable vectors as vision prompts into each layer of frozen CLIP image encoder to adopt zero-shot semantic segmentation tasks. LPT optimizes shared prompts to explore the general features across the entire long-tailed dataset, and group-specific prompts to endow the fine-grained discrimination ability into frozen pre-trained vision models.

As proven by the above works, prompt learning enables pre-trained visual models to adapt to a variety of visual tasks in natural scenarios. However, prompt learning still has great potential in transferring visual knowledge of pre-trained vision models trained in natural scenarios to downstream tasks that have large domain gaps. A recent study has extended vision prompt from natural scene understanding to diverse vision tasks with huge domain discrepancies, such as point cloud analysis , image generation and even speech understanding . Concretely, PointCLIP converts the raw points into scatter depth maps by projecting them onto predefined image planes, termed as vision prompt, which effectively transfers the remarkable ability of CLIP model. In addition, PointCLIP also narrows the modality discrepancies between unordered point clouds and the visual images, thus producing a unique insight for processing vision tasks with significant domain gaps using prompt technology. P2P proposes the geometry-preserved projection and geometry-aware coloring operations to translate point cloud data into colorful images, which are regarded as vision prompt and further adapt the pre-trained vision model for various point cloud analysis tasks. PromptGen samples images via approximating the energy-based models efficiently to transfer other off-the-shelf generative models and further control the distribution of pre-trained generation models. Kim et. al design three different prompts containing input, feature, and temporal dimension prompts to steer frozen pre-trained vision models to conduct visual speech recognition tasks. These works show that vision-driven prompts can transfer pre-trained vision models from natural scenarios to various downstream tasks even with domain discrepancies.

Excitingly, the above work takes only a simple manner (e.g., adding extra parameters into inputs) to construct visual prompts but makes great progress on transferring the remarkable discrimination and generalization of pre-trained vision models. To further investigate and dig into the effectiveness of vision prompt, the other family of approaches also tend to design a sub-network to construct vision prompts, as shown in Fig. 2(c). Specifically, the vision-driven prompts PV\mathcal{P}_{V} can be denoted as:

where Φ(,)\Phi(,) denotes the designed sub-network to produce vision prompts PV\mathcal{P}_{V}, θ\theta is the learnable parameters in Φ(,)\Phi(,), and XX is the input image. For instance, NOAH implicitly learns the feature dimension for adapter and the token length for vision prompt via a neural architecture search algorithm. PGN learns to produce input-dependent prompts via selectively sampling input images from a commonly learned library of tokens. FPTrans divides the input image into foreground and background using a pre-trained ViT network with a lightweight module, which can be prepended with foreground and background prompts. FRPT explicitly zooms the discriminative regions of input images via designing a lightweight sampling network to obtain the vision prompts. RePro localizes objects from videos as vision prompts utilizing a tracklet detector and further learns the correlation between subjects and objects according to the learned vision prompts. Gan et. al learned domain-specific prompts and domain-agnostic prompts, which capture the knowledge of the source domain and maintain the domain-shared knowledge in the continual adaptation. ViLD generates multiple regions of interest based on the region proposal network, regarded as vision prompts, to align their visual embeddings and textual embeddings for open-vocabulary object detection. These works can produce appropriate prompts according to downstream tasks, thus effectively exploring the remarkable generalization and discrimination ability of pre-trained vision models. More importantly, compared to introducing learnable parameters directly, they can improve the interpretability of these vision prompts such as directly modifying the pixels.

2.2 Language-driven Prompt

Recently, large-scale vision-language models are pre-trained by extensive image-text pairs and focus on open-world visual concepts. Following this ideology of prompt learning in NLP, most existing works tend to transfer large-scale vision-language models into various downstream vision tasks via designing appropriate language-driven prompts . As shown in Fig. 2(a), most works, such as CoOp , firstly extract unified context or class-specific context of visual images as language-driven prompts to adapt frozen pre-trained vision-language models to diverse vision tasks. Formally, the language-driven prompts can be formulated as below:

where PT\mathcal{P}_{T} denotes the language-driven prompts, aia_{i} represents the ii-th learnable vectors, the number of learnable vectors is nn, and <class><class> is the class embeddings. Extensive existing works have focused widely on this line of language-driven prompt and utilize this prompt analogous to the designed in CoOp to adapt various downstream tasks, e.g., domain adaption , semantic segmentation , video understanding , and few-shot learning . In addition, recent methods extend the original language-driven prompt and thus design multiple complementary language-driven prompts to better mine the task-specific knowledge from pre-trained vision-language models. The multiple complementary language-driven prompts PMT\mathcal{P}_{MT} can be represented as

For instance, PLOT learns multiple comprehensive prompts to capture different attributes of classes and aligns visual embeddings and multiple textual embeddings via optimizing the optimal transport distances between multiple prompts.

2.3 Vision-language Prompt

Vision-driven and language-driven prompts have been explored to simultaneously modify the vision and text inputs for pre-trained vision-language models, thus transferring the discrimination and generalization ability of pre-trained vision-language models thanks to effectively aligning visual and textual embeddings . For an instance, UPT designs a shared prompt network to produce the vision prompt and text prompt, thus narrowing the gap between visual representations and textual embeddings. DPT simultaneously optimizes the visual and textual prompts from the vision and text input perspectives, which aims to modify the textual classifier and visual representations of pre-trained vision-language models. MaPLe establishes interactions between vision and text prompt by injecting the vision prompts into the language prompts directly, further improving alignment between the vision and language representations outputted by pre-trained vision-language models. TPT introduces learnable text prompts with random vectors and category names and designs vision prompts generated by cropping input images randomly. These methods can transfer the pre-trained vision-language models to various downstream tasks from the perspective of text and vision inputs.

2.4 Discussion

It is well known that the quantity of labeled data largely determines the upper limit of the vision algorithm. Vision prompt learning usually focuses on solving the problem of few-shot or zero-shot learning, which allows the model to perform relatively well even without labeled data. Besides, visual prompt learning unifies all downstream tasks into pre-training tasks via designing a specific template. In this way, the data derived from the downstream tasks are transformed into novel inputs. In this way, these inputs could exploit the capabilities of the pre-trained models themselves. In other words, the core mechanism of the vision prompts aims to exploit the potential of the upstream pre-trained model, so that the upstream pre-trained model can perform the downstream task as well as possible with a few or less requiring labeled data. As described above, we need to reformulate the different downstream tasks to make them suitable for the pre-trained model.

After examining the work related to the three prompt tuning strategies including language-driven, vision-driven, and vision-language prompts, it is clear that they have their own unique characteristics. For language-driven prompt tuning, although these methods in vision-language models achieve impressive performance on a wide range of vision tasks even without any fine-tuning, the language-driven prompt schemes are tailored for the multi-modality model and thus inapplicable to the pre-trained vision model. In addition, language-driven prompts translate the semantic information of the input image into category-specific descriptions. In this way, the extracted knowledge from the input image tends to solve the original pre-trained tasks instead of the downstream task. Therefore, these works based on language-driven prompt strategies can neglect the task-specific knowledge of inputs, thus reducing the performance of downstream tasks. For vision-driven prompt tuning, they leverage a simple yet effective manner, i.e., introducing learnable parameters, or designing a light-weight sub-network according to the characteristics of downstream tasks, to generate vision prompts for steering pre-trained vision models for various downstream tasks. Combined with the advantage of language-driven and vision-driven prompt tuning, vision-language prompt tuning can transfer the discrimination and generalization of pre-trained vision-language models from the perspective of modifying text and vision inputs.

3 Adapter Tuning

Adapter-based methods are a class of techniques that involves additional trainable parameters into a pre-trained model that has been frozen to facilitate learning for downstream tasks. In the NLP domain, adapters were first introduced by Houlsby et al. as a means of achieving PETL. However, efficient adaptation, particularly in the field of computer vision, has received comparatively little attention. Initial efforts to develop adaptive methods for computer vision have included incremental learning methods and domain adaptation methods . Subsequently, adapters have garnered interest across domains and have been successfully applied in the computer vision field. Adapters provide a lightweight alternative to extensive model fine-tuning.

In this section, we have sorted out the existing vision-related adapter-based tuning methods, which can be roughly divided into three ideas, i.e., sequential adapter, parallel adapter, and mix adapter one by one as follows.

where Z^l\hat{Z}^{l} denotes the optimized features outputted by the sequential adapter.

In sequential adapter strategies, research efforts can be roughly categorized into two groups: inserting residual blocks directly, and using more parameter optimization techniques to minimize the size of the adapter. Studies of the first family emerge in the earlier stage without large-scale models, the initial development of Res-adapt involves a customized deep network structure that utilizes adapter residual modules to dynamically adjust to different visual domains in real-time. Following Res-adapt, DAN once converges to a comparable or even a higher level of performance while requiring a fraction of the parameters (typically 13%) of standard fine-tuning procedures. Their recent efforts develop LST, a technique that trains a separate ladder side network using intermediate activations from backbone networks via shortcut connections to make predictions. This approach improves the model’s accuracy while reducing computational complexity. In addition, Conv-Adapter systematically investigates feasible solutions to learn task-specific knowledge, which adapts intermediate features of each residual block in pre-trained ConvNets using four different variants, including the parallel adapter discussed below.

A second family of approaches focuses on optimizing the network architecture in order to better understand the efficiency of sequential adapters. Alternatively, EPM suggests that universal parametric neural network families are used, but with limited parameters. It discusses a number of methods for parametrization, including parallel and series residual adapters, compression of joint adapters, and parameter allocation. Polyhistor decomposes a hyper-network into a pair of separate hyper-networks and factorizes an adapter weight matrix into two kernels for parameter reduction in multi-tasking architecture. Pro-tuning exploits semantic information from diverse levels to enrich the feature space by building multiple stage-wise prompt blocks, which can be simply plugged into a model for quickly adapting to a new task. By generating attention weights without token-token interactions, AMixer captures long-term and short-term spatial dependencies without self-attention. Shysheya et. al proposed a novel technique called Fit, which scales and shifts activations and leverages a Naive Bayes final layer classifier for image classification, and is automatically configured and achieves promising results. Marouf et. al introduced TINA, an approach that iteratively reduces the size of every adapter using a scoring function compared to neuron importance. This process enables the automatic estimation of every adapter and improves the efficiency of the overall model. Luo et. al proposed RepAdapter, which uses re-parameterization of sparse structure to approach nearby projection weights and further reduces the parameters of the model while maintaining its effectiveness and lightweight nature.

Adapters have become a popular technique for foundation tasks in computer vision, where the pre-training task is often image classification. However, other tasks such as high-level vision tasks , low-level vision tasks , video understanding , and robotic control all require designs that are tailored to their specific architecture, in order to efficiently transfer learned parameters and achieve good performance through PETL. In addition to these task differences, recent research has proposed innovative ways to utilize adapters in different applications. For instance, BDTL and ViTDet adjust a plain backbone with minimal adaptions only during fine-tuning, thus eliminating the hierarchical constraint from the object detection task. In Florence , representations are expanded from coarse to fine, from static to dynamic, and from RGB to a number of modalities (caption, depth). It can be used to perform a wide range of vision tasks by incorporating universal visual-language representations from Web-scale image-text data. SND utilizes a dynamic stacked network as the adapter to activate the pre-trained model in a plug-and-play method for image restoration. MK-Adapter introduces a learnable multi-knowledge adapter to adaptively blend the predictions from CLIP and DINO for few-shot classification. ADA is developed to perform continual learning using pre-trained transformers and adapters, which adds an adapter before layer norm and feed-forward layers (MLP). AIM introduces spatial, temporal, and joint adapter separately and ST-Adapter propose a spatiotemporal adapter, which both aim to equip an image model with spatiotemporal reasoning capability for video understanding. PEA addresses the limitations of classical fine-tuning in robotic manipulation by utilizing the classic bottleneck architecture widely used in image classification as a lossless adaptation. CAOA proposes a content-adaptive optimization framework for image compression that trains adapters inserted into the decoder through optimization in terms of rate-distortion, with adapter parameters being additionally transmitted.

In the field of multi-modal learning, with the development of large-scale cross-modal pre-trained models, i.e., CLIP and ALIGN , adapter technique has been widely adopted, using a design analogous to the one mentioned above, to adapt various downstream tasks for efficient fine-tuning with excellent results. HA firstly analyses empirically and recommends general recipes for efficient multi-modal transfer learning including LayerNorm-tuning. CLIP-Adapter combines residual-style feature blending with an additional bottleneck adapter to learn new features on either a visual or a language branch. On this basis, Tip-Adapter aims to enhance the few-shot capability, but does not require backpropagation during training by using a key-value cache. MAGMA combines visual and textual inputs to produce text using adapter-based fine-tuning, resulting in more complex generative language models. BALLAD augments the representations of tail classes on balanced training samples by employing an additional adapter layer for long-tailed vision language learning. Building on previous work, subsequent methods focus more on parameter efficiency and multi-modal effectiveness. Through a hierarchical adapter module structure, Hierarchical3D integrates multi-modal content into a trained textual summarizer, while tuning only a few model parameters. VL-Adapter adjusts the pre-trained model with sequential adapter layers for the cross-modal domain (i.e., image-text, and video-text) to construct the benchmark. HyperPELT is an approach for fine-tuning small modules in a pre-trained language model using a shared hyper-network that takes trainable hyper-embeddings as input. Then CrossModal-Adapter and MV-Adapter both attempt to make further improvements with a notable difference: allowing cross-modal interactions early in the training process by the weight-sharing mechanism. To perform better on challenging datasets, SVL-Adapter brings together the complementary strengths of vision-language pre-training and self-supervised representation learning for extreme scenes like few-shot learning. LAVISH firstly migrates the adapters for pre-trained ViTs to audio-visual tasks by injecting a small number of trainable parameters into every layer of a frozen ViT. These innovative approaches demonstrate the versatility of adapters and their potential for various applications beyond traditional classification tasks.

3.2 Parallel Adapter

where Z^l\hat{Z}^{l} denotes the optimized features outputted by the parallel adapter.

The simplest application methodology is to insert an adapter module in parallel. For example, ViT-Adapter allows plain ViT to achieve comparable performance to vision-specific transformers with prior information for dense predictions, introducing the image-related inductive biases by a pretraining-free adapter. Then, PESF-KD first mathematically formulates the mismatch as the sharpness gap which can be narrowed with the appropriate smoothness of the soft label. It introduces an adapter module for the teacher, and only updates the adapter to obtain soft labels with appropriate smoothness. AdaptMLP replaces the original MLP block with two branches, including the frozen and the trainable branch in parallel for adapting vision transformers to a large video action recognition, which avoids catastrophic interference with each other. Convpass argues that current adapters have weak inductive bias which limits their performance and utilizes trainable convolutional blocks to bypass this weakness. In the application for multi-modal communities, AMA assumes features from different modalities are already spatially aligned, therefore, it restores the 2D structure for each modality after the channel squeezing to efficiently aggregate local multi-modal features as the adapter. UniAdapter refers to the parallel structure to unify uni-modal and multi-modal adapters. Specifically, adapters are distributed to different modalities, with the total number of tunable parameters reduced by partial weight sharing.

3.3 Mix Adapter

Mix adapter introduces new parameters in different positions with mixed architecture demonstrated in Fig. 3(c), i.e., the multi-head attention blocks in each Transformer layer.

PATT presents a variation of a prefix-tuning module for video-based downstream tasks, which involves bringing trainable parameters to different positions on the backbone in various ways, analyzing different parameter efficient techniques, such as data scales for downstream domains and positions of trainable parameters. Following PATT, ETT uses the latest innovations in attentive prefix tuning (i.e., generating new key-value pairs) and domain residual adapter (i.e., controlling domain gaps) for few-shot learning. PALT initializes several adapters randomly and identifies the influence of each adapter module so that it can prune adapters based on the prestigious lottery ticket hypothesis. VQT aggregates intermediate features of Transformers for effective linear probing, featuring parameter and memory efficient transfer learning on vision tasks. By applying group-wise convolution techniques to features, the Consolidator structures tunable parts for efficient transfer learning in large vision models. TVG compares popular pre-trained models with existing approaches, as well as testing different adapters, so that additional parameters do not have as much impact on the video grounding community.

3.4 Discussion

Adapter-based methods are a type of transfer learning technique that has gained popularity in computer vision research in recent years. The idea behind adapter-based methods is to focus on solving the vision downstream tasks with a frozen backbone and only a small number of parameters. Adapter-based methods have several advantages over traditional transfer learning approaches. First, they are computationally efficient, since they only require training the adapter module and not the entire network from scratch. Second, they are highly modular, allowing for easy adaptation of pre-trained models to new tasks without requiring extensive modifications to the original architecture. Finally, they can improve the generalization performance of pre-trained models by adapting them to new tasks while preserving their learned representations.

After examining the work related to the three adapter architectures (sequential, parallel, and mix), it is clear that each has its own unique advantages. The sequential adapter is advantageous because it is simple and easy to use, making it adaptable to the pre-trained model’s architecture and allowing for better results. This is exemplified by its success in a variety of tasks. On the other hand, parallel adapter architecture excels at activating the knowledge of the pre-trained model, resulting in improved network representation and optimal performance. Lastly, mix adapter architecture offers flexibility and minimal dependence on the number of parameters, which is often achieved through mathematical approximation techniques.

In the future, more relevant work on vision adapter tuning will emerge and be more widely explored. We believe that the following two directions are of some research value: Firstly, more efficient operations can be introduced and more communities can be benefited from adapter-based tuning methods. Secondly, more adapter architecture can be studied. For instance, in the realm of NLP, there exist adapter architectures that exhibit promising performance in adapting to new tasks, which can be leveraged and applied to computer vision. Furthermore, emerging integration techniques will likely enable adapters to achieve improved performance in practical applications.

4 Parameter Tuning

An effective parameter-based tuning involves directly modifying the parameters (either weights or biases) of the pre-trained model in a more aggressive manner. Given a specific layer, it can have its weight-term multiplied to the feature map and a bias-term added to the feature map. As shown in Fig. 4, this section introduces parameter-based methods based on which part of the parameters are tuned: weight part, bias part, and both. Techniques can be grouped into addition and decomposition. Existing works also termed this technique as reparameterization-based methods .

Bitfit is also known as side-tuning, which only tunes the bias part of the pre-trained model (see Fig. 4[a]) and can be represented as:

where the weight parameters WW are frozen, and the bias bb contains the parameters optimized in the tuning process.

Xu et al. introduced side-tuning as two branches: one for predicting mask proposals, and the other for predicting attention bias which is applied in the CLIP model for semantic segmentation. It adds a bias term to the results of the Softmax layer of the attention module. Differentially Private Bias-Term Fine-Tuning (DP-BiTFiT) proposed a differentially private version of bias-tuning. DP-BiTFiT used the optimizer DP-SGD to make the bias term private: first aggregate bias gradient norms across all layers, then use it to compute clipping factor, add Gaussian noise to the sum of clipped gradients, and descend on bias term. DP-BiTFiT basically changed the way for optimizing the bias term, which achieves comparable performance with bias-tuning. DP-BiTFiT’s implementation is worth noting as it does not calculate the gradients for the pre-trained weights, which helps to save over 60%60\% training time.

Namazifar et al. studied the role bias-term of Transformer for NLP tasks. From the mathematical perspective with empirical verification, it concludes that the bias term of the key linear transformation is redundant and can be omitted without any impact on the attention module. Moreover, the bias term of the value linear transformation has a more prominent role than that of the bias term of the query linear transformation.

4.2 Weight Part

The LoRA structure has been applied to an encoder-decoder model called motion style adapters (MoSA) . MoSA uses a lightweight LoRA structure for adapting the motion style (e.g., pedestrians) from a source domain with sufficient labeled data to a target domain (e.g., cyclists).

DyLoRA proposes to truncate the parameters of rank to multiple parts (i.e., ranks) and optimize them separately sequentially without relying on a search mechanism.

DnA remains needs to use SVD to implement the GreBsmo algorithm, bringing additional complexity to the iterative optimization process.

The decomposition method using the Kronecker product is also named as a parameterized hypercomplex multiplication/convolutional (PHM/PHC) layer , being applied for varied tasks such as vision and audio tasks. PHM inspires to form a tunable weight with three terms zi,siz_{i},s_{i}, and AiA_{i}, being added to the pre-trained weight for PETL of NLP tasks. FacT considers two decomposition methods Fact-TT and Fact-TK, using Kronecker product and a multilinear generalization of the SVD (i.e., the Trucker model) , respectively. Fact-TK generally performs better than Fact-TT with slightly more parameters than Fact-TT across 19 image-based tasks, which is far fewer parameters than the basic LoRA method. Dynamic Linear Dimensionality Reduction (DLDR) claims that only optimizing the low-dimensional subspace of a large model can achieve comparable performance. DLDR used SVD to decompose the weight to find the tuned subspace, which achieves comparable performance by training a small number of epochs.

RepAdapter is built based on LoRA structure and introduced a group-wise transformation method to reparameterize the weight term. RepAdapter also interpreted its group-wise divided LoRA layers as a reparameterization process. RepAdapter aims to reduce the inference time and seamlessly integrated the RepAdapter into most giant vision models via structural re-parameterization.

Similarly in NLP, task-adaptive reparameterization (TARP) uses the Kronecker product as a dynamic low-rank decomposition for the MLP module for domain adaptation. Kronecker Adapter (KronA) also introduces the Kronecker product to improve the limited representation power low-rank representation for NLP tasks.

4.3 Weight and Bias

where are respectively interpreted as scale and shift factors. Note that both γ\gamma and β\beta are learnable vectors, which can be much smaller than matrix variables of LoRA or decomposed forms of DnA, Compacter, and KAdaptation.

4.4 Discussion

Differences from the prompt-based and adapter-based methods, parameter-based tuning can use fewer parameters to achieve a similar effect of adaptation. On the tested image-based tasks can even outperform adapter and VPT. SSF introduces the bias-tuning technique to the weight variable via dot product. According to the analysis of SSF, it intrinsically modifies both the weight and bias variables, which is interpreted as follows:

Parameter-based tuning can be less expressive from the perspective of tuned parameters but can sometimes outperform the former two methods. Given the varied techniques available, Mao et al. unified these methods with a gate mechanism. uses pruning techniques to drop the activations during back-propagation, leading to sparse activations. This track of techniques will be further introduced in Section 3.5. In addition to Transformer-based structures, LoRA Winograd convolution aims to use the LoRA mechanism to prune the 3D CNN backbone model (e.g., C3D and R3D-18) for accelerating the Winograd operation with less trainable parameters. By far, most methods are tested on Transformer-based structures, it remains exploration on their effect on CNN-based structures.

5 Remapping Tuning

Instead of directly fine-tuning or processing the pre-existing model, remapping-based tuning is a category of techniques that transfer the knowledge learned by a pre-trained model to a new downstream model. Based on how to utilize the pre-trained model, i.e., output, weight, and network architecture (see Fig. 5), we discuss three forms of knowledge transfer in the following categories: knowledge distillation-based remapping, weight-based remapping, and architecture-based remapping.

Knowledge distillation aims to regularize the downstream model by enforcing it to mimic the output of pre-trained models. Note that, the output typically refers to the final response or intermediate features. In the meantime, knowledge distillation is also an important model compression technique. In this section, we do not involve other model compression techniques such as network pruning , since they are not typically motivated to transfer knowledge from teacher network to student network.

The fundamental idea of knowledge distillation is to transfer the learned knowledge from a large pre-trained teacher model into a small student model by learning the network output or intermediate features of the teacher. Typically, knowledge is distilled from the teacher model to the student model using a soft target distribution for each case. The probability qiq_{i} of each case can be formulated as:

where zz is the output logit of the teacher networks and TT is the temperature of the distillation process.

To our best knowledge, the work of first introduces knowledge distillation to extract knowledge from a pre-existing model. They trained a compressed model with the pseudo data produced by an ensemble of shallow networks while no significant loss occurs in performance. This idea has been expanded to compress the deep and wide networks into shallower ones in . Hinton et al. introduced the teacher-student knowledge distillation framework, where the student network is penalized based on the softened class distribution output of the teacher network.

One of the characteristics of deep neural networks is to obtain increasingly expressive power by learning hierarchical feature representations, as pointed out in . Based on this theory, both the final response and the intermediate feature maps of the teacher network can be employed as the target for training the student model. To substantially exploit the information of intermediate layers, Fitnets introduces intermediate-level hints of the teacher to facilitate training the student. It enforces the intermediate feature alignment between the teacher and student networks via the teacher’s intermediate feature maps as hints. Subsequently, a rich line of work is devoted to aligning the features indirectly . Concretely, Kim et al. developed a factor transfer method that employs paraphrased intermediate features of the teacher as a factor, rendering the knowledge of the teacher network more understandable for the student network. Inspired by neural architecture search (NAS) , Guan et al. developed a two-stage distillation approach that adopts the differentiable search strategy to simultaneously improve the efficiency and the effectiveness of knowledge distillation. Xu et al. developed a feature-normalized distillation method by introducing a sample-specific correction factor for the replacement of the temperature, with the goal of suppressing the impact of noise resulting from the one-hot label. Passalis et al. modeled the information flow of the teacher’s multiple intermediate layers and then train a student model to match this information flow. To realize knowledge transfer for vision transformers, Touvron et al. introduced a token-based distillation strategy termed DeiT, which enforces the student transformer to directly reproduce the label estimated by the pre-trained teacher network using a distillation token. Hao et al. introduced a manifold distillation approach for vision transformers by substantially utilizing patch-level information.

There are also some extensions to further explore knowledge transfer patterns. Park et al. proposed a relational knowledge distillation scheme for mutual relations transfer of outputs instead of individual outputs. As a generalization of vanilla knowledge distillation, they introduced distance-wise and angle-wise distillation losses to sufficiently extract the structural relations in data examples. Liu et al. proposed an architecture-aware knowledge distillation approach termed AKD, with the goal of finding the optimal student networks for distilling a given teacher network. Chen et al. studied the semantics of intermediate layers and employ an attention mechanism to automatically assign the soft layer association between teacher and student networks, which can reduce the impact of over-regularization during the training process. Zhou et al. proposed a holistic knowledge distillation with graph neural networks, where the holistic knowledge contains individual knowledge and relational knowledge . To integrate two knowledge and refine their correlations, graph neural networks are adopted to learn holistic knowledge to provide supervision for the student network by aggregating feature representation from correlated data examples. Chen et al. proposed a residual distillation framework termed Review to effectively learn informative features from multi-level information in the teacher network. Review utilizes multiple layers in the teacher to guide the training for one layer in the student with great performance gains. Zhao et al. modeled the traditional knowledge distillation loss into target class knowledge distillation and non-target class knowledge distillation, then dived into their effects. Based on the observations, they found that the traditional knowledge distillation loss is a highly entangled formulation. To address this issue, they introduced a decoupled method to facilitate the effectiveness and flexibility of knowledge distillation.

For application, most of the above methods focus on image classification. Furthermore, knowledge distillation also demonstrates promising results in more vision tasks, such as object detection , image segmentation , person re-identification , super-resolution , depth estimation , and crowd counting .

5.2 Weight Remapping

Rather than relying on the teacher’s output as supervision to train the student, weight remapping directly transfers the model weights from the teacher network to the student ones. Net2Net is a pioneering effort that rapidly transfers the knowledge stored in a pre-existing network into another network by remapping the weight of a pre-existing teacher network to the student. Its main purpose is to provide substantial speed-up for training the student.

Subsequently, EAS introduces the concept of weight remapping into neural architecture search. EAS proposes an efficient architecture search framework by exploring the search space according to a pre-existing network and reusing its weights. By leveraging reinforcement learning strategies as the meta-controller, EAS can adaptively generate network transformation actions to determine the network depth or width with function-preserving transformations . To tackle variable-length architecture and consider the entire input architecture, EAS employs a bidirectional recurrent network as the encoder network. In this way, the previously trained network can be further exploited to efficiently explore the architecture space and greatly accelerate the training process of the new network. To efficiently compress the teacher network for knowledge transfer, Ashoket et al. proposed a reinforcement learning-based approach termed N2N learning, which models the conversion from the teacher network into a student network as a Markov Decision Process (MDP). N2N learning formulates the process of knowledge transfer as a two-stage action selection. In the first stage, a recurrent policy network selects a sequence of actions including keeping or removing layers of the large teacher network. In the second stage, another policy network performs further reduction in each remaining layer for matching the attenuate configuration. Then, the resulting student network is evaluated to acquire the reward for learning the policy by the process of reinforcement learning. To this end, N2N learning introduces a compression reward term that models the limited computation budget within linear constraints.

Furthermore, some interesting weight remapping methods take the network path topology into consideration instead of merely adding or removing network layers. Elsken et al. introduced a hill climbing-based approach named NASH, which can automatically search the optimal student architecture. By using a series of alternative network morphisms, NASH can train the child networks with a short optimization process by cosine annealing. At each training step, NASH searches for the optimal architectures by a simple hill-climbing strategy . Path-level EAS enforces the meta-controller to change the topology of network connection paths while using function-preserving transformation operations to remap weights. Thus, the designed student network enables complex path topologies. To achieve this, Path-level EAS develops a bidirectional tree-structured meta-controller based on reinforcement learning, in order to enrich the architecture space to the generalization of multi-branch structures. The tree-structured architecture contains edges and nodes to form plentiful paths within each cell, in which each node has an allocation and the merged scheme.

Previous object detection and semantic segmentation approaches use the network weights pre-trained on image classification for performance gains. However, one of the major challenges is that ImageNet pre-training typically requires highly large computation costs. To address this issue, Fang et al. introduced a fast neural network adaptation approach dubbed FNA, making a pre-trained network adapt to a new task by modifying the network such as depth and kernels. In this way, FNA can expand NAS techniques to object detection and semantic segmentation with negligible computation costs. Technically, FNA first designs a seed network by selecting a manually designed network that is pre-trained on ImageNet such as and then enlarges it to a super network. By applying the weight remapping technique, the seed network is used to assign new model parameters. Note that, FNA expands the weight remapping at the level of the kernel, where parameters are allowed to form a smaller network during mapping instead of only the deeper and wider network as in Net2Net. FNA greatly reduces the computation overhead in network adaptation by sufficiently utilizing the pre-trained ImageNet image classification weights. The follow-up work FNA++ extends the weight remapping of FNA to one more task (i.e., human pose estimation) and more network architectures including ResNet and NAS networks with diverse widths, depths, and kernel sizes.

5.3 Architecture Remapping

Architecture remapping refers to the knowledge transfer about network architecture from a pre-existing model. To our best knowledge, this line of work is mainly used in weight-sharing neural network search (NAS). Formally, S\mathcal{S} denotes the architecture and ω\omega denotes the weight of S\mathcal{S}. The goal of NAS is to find the optimal architecture S∗\mathcal{S}^{*} that produces the best performance on the test set:

Specifically, this type of NAS formulates the search space into an over-parameterized super-network, e.g., modeling the search space as multiple repeatable cells . When transferring the searched architecture to downstream tasks, direct architecture transfer, which stacks several searched cells to form a downstream model and then retrain it on the downstream data, is the current mainstream scheme. Canonical examples include DARTS and its variants . Direct architecture transfer has shown impressive results on downstream tasks.

5.4 Discussion

Remapping-based tuning is a class of techniques of knowledge transfer from a pre-trained model to a student downstream model. Different from traditional transfer learning approaches, this line of work focuses on training a new downstream model that is disentangled from the parameters of the pre-existing model. That is, the pre-existing model does not engage with the inference. Thus remapping tuning methods own their exclusive advantages. On the one hand, they can achieve significant parameter-efficient by compressing the downstream model such as knowledge distillation. On the other hand, they have high flexibility in the network structure of the downstream model, since its parameters are decoupled from the pre-existing network.

For three classes of remapping tuning methods, there are also several interesting directions for future work. Knowledge distillation always pays more attention to efficiency, thus greatly reducing computational costs. Despite its efficiency, knowledge distillation typically requires careful balancing of the task-specific loss and distillation loss to ensure fine-tuning performance, even requiring engineering experience sometimes. Weight remapping is a simple and effective solution without manually adding constraints. However, this type of work typically struggles to obtain a lightweight student network compared with knowledge distillation, which restricts its application. Architecture remapping excels at transferring existing network architecture to the downstream model, but its application is typically limited in neural network search.

Visual Tuning Future

To date, the state of visual intelligence forms a transfer learning paradigm of pre-training and tuning, showing great promising performance on numerous benchmarks. Vision contributes a large portion of knowledge acquisition of human intelligence. However, due to the high dimensionality of vision data, the intelligence of machine vision suffers from a relatively small data scale in comparison to that of NLP, and remains lag far behind the general human vision. The future promise of intelligent vision will be expanded beyond the competed benchmark datasets, realizing transformative impacts on more domains via a multidisciplinary coevolutionary process. On one hand, we expect that future pre-training techniques play the role of knowledge acquisition and storage in a “collection-labeling-training-feedback” cycle system. While future tuning is around how to make use of the learned knowledge through more diversified interactions beyond the prompts around conversational systems. Along the way to further understand the mechanisms of deep neural network models and even the human brain, we discuss the future works of vision from perspectives of pre-training and tuning techniques.

Previous works use supervised or self-supervised methods to guide models to learn representations of our visual and visual-text world. The supervised pre-training method is a mainstream practice of the traditional transfer learning paradigm , While self-supervised pre-training scales pre-trained models up to foundation models (introduced in Section 2.4). Although encouraging progress has been made, as data continues to accumulate, we expect that future pre-training techniques will be able to constantly scale up the model size and improve the capabilities of foundation models. Here, we discuss the future directions of model pre-training from three perspectives: data, models, and optimization.

Quality data are the nourishment of foundation models and are regarded as an important property of the data owner. To realize the promise of future foundation models, it is expected to acquire more fundamental knowledge from the open-world multi-modal data with characteristics as follows:

Increasing scales: Concerning the data volume, large vision models that learn knowledge from large-scale datasets are empirically proven effective for adapting to downstream tasks via tuning techniques. However, compared with human vision, existing large-scale vision datasets remain far from the amount of data that humans learn from. On the contrary, the situation in NLP can be different as large language models can be regarded as having a wide knowledge of the Internet, making them more capable in some NLP tasks, i.e., chatGPT. To scale up the data volume, multi-modal data (e.g., image, video, audio, and text), multi-source data (e.g., Internet, generative models such as Nerf and Diffusion ), and multi-sensor data (e.g., different types of cameras, biomarkers, and ambient sensors) can be considered for training large models.

High quality: Before arbitrarily collecting large-scale data, determining what data and how much data are essential concerns. The newly collected data can be redundant or noisy, respectively leading to limited or even negative effects on the model. Chen et al. introduced the diversity rule at the level of feature representation. However, there remains a lack of investigation into the quality at the data level for existing benchmark datasets, giving rise to research on topics such as out-of-distribution generalization and tolerance to noise (see Section 2.1.3). Further investigating measurable factors of data quality (e.g., 16 dimensions summarized in ) and their corresponding consequences on the large model can bring a large impact on the machine intelligence community. Findings will guide evidence-oriented data collection and effectively reduce the expensive labeling cost.

Security and privacy are always the priority throughout the life cycle of the data, especially for domains such as healthcare and finance when interacting with large models on the cloud . Issues around cloud computing can be grouped into four aspects: 1) users’ control over the data, 2) authorized replication, 3) legal requirements, and 4) cloud subcontractors’ processing . Protective actions can be taken at the data level to prevent attacks such as re-identification, dataset reconstruction, and tracing .

1.2 Models

Given the multimodal, multi-source, multi-sensor data, vision large pre-trained models are expected to continuously accumulate knowledge from the new data in an interpretable and secure mechanism.

Theoretical support: Training models with theoretical support from statistical and biological perspectives can make them more interpretable, explainable, and improvable. In the regime of large models, there are a certain number of recent works motivated by some theoretical definitions from the statistical perspective , by which a generalization bound are used to promise efficient knowledge transfer. Except for the statistical aspect, biological and neuroscience discoveries also benefit the development of deep neural networks, which can provide more insights and inspire new ideas for future large vision models. We observe recent works in Section 2.1.1 are mainly delayed from one to another as they are intended to explain each other’s observed domain empirical realities instead of truly inspiring new ideas. Basic neural network connections are inspired by how brain neuron works, but we have not yet known exactly how the human brain learns new knowledge. As such, we are also not clear if the knowledge acquired by existing large models via back-propagation can be effective. On one hand, we humans are sentient beings and acquire knowledge via multiple sensations: vision, sound, haptic, taste, etc. On the other hand, the human brain can be very efficient to activate just a small portion of neurons to complete a task, while existing foundation models do not. By far, we are kind of building something different from the most intelligent and efficient machine in the world (i.e., the human brain). Understanding the brain can be the next turning point (i.e., artificial general intelligence), which brings serious ethical issues.

Continuously updating: As introduced in Section 2.2, a model can be featured with its domains and tasks with their feature space and label space. We expect that foundation models will not only scale up at the parameter aspect but also together with at aspects of domains and tasks. For a single domain, it can have multiple tasks. Models such as GATO and Flamingo are pre-trained with multiple tasks, where the former covers vision tasks while the latter even covers both NLP and vision tasks. For a single task, a classification task can have a novel class, which is defined by an active learning paradigm (i.e., lifelong learning or continual learning). In contrast to batch learning where all training data is available at once, continual learning represents a family of methods that accumulate knowledge and learn continuously with data available in sequential order . The future stronger vision and visual-language models will bring a more profound impact on other domains via the multidisciplinary coevolutionary process.

Security: Aside from privacy issues at the data level, large foundation models (i.e., at the model level) can also be vulnerable to attacks. Foundation models allow users to easily plug and unplug via APIs, which raises security and privacy concerns. such as adversarial attacks and model-inversion. Kaissis et al. introduced the advantages of federated learning and provided an outlook for future works. Although federated learning can mitigate data-level privacy issues, it can be vulnerable to adversarial attack . Around privacy-preserving AI, the adversarial attack will attract more research in the near future.

1.3 Optimization

Current foundation models are generally optimized with back-propagation and reinforcement learning with human feedback. Optimization itself can be relying on hardware devices, hyperparameter configuration, and algorithms as follows:

Hardware: Recent large models are trained with GPUs, which is unaffordable for general researchers or small companies. Fortunately, the newly released NVIDIA Hopper H100 GPU supports FP8 format for accelerating compute-intensive Transformer models (around 9 times faster than previous A100 GPU for training), making trillion-parameter models within the reach of all researchers. While the inference speedup of H100 compared with A100 can be 30 times faster, making tuning a promising direction.

Hyperparameter configuration: In machine learning, hyperparameters such as initial learning rate, batch size, and task-specific parameters often considerably impact performance. To avoid the manual process of trial-and-error, hyperparameter optimization is a sub-field of automated machine learning, which aims to automatically identify a well-performing combination of hyperparameters. Simple techniques are grid or random search. While recent advances in hyperparameter optimization are evolution strategies, Bayesian optimization, Hyperband, etc. .

Algorithm: The combination of back-propagation and stochastic gradient descent remains the mainstream algorithm to make foundation models optimize towards some statistical goals (e.g., the probability that a picture is identified as a cat). Meanwhile, reinforcement learning with human feedback brings more raw human opinions, which can be some kind of human-machine interactions that align the pre-trained large models to more specific human desired tasks.

2 Tuning Techinques

As introduced in Section 3, recent developments in visual tuning techniques can be regarded as originating from the prompt tuning of the NLP domain and working towards the PETL direction. Then a couple of adapter methods are proposed, showing better performance than visual prompt methods but lack of interpretability. The bias-tuning and LoRA methods further reduced the number of parameters, leading to direct parameter tuning methods via addition or decomposition. More recent works are grouped as remapping tuning, among which NAS-based methods show an even more aggressive PETL manner. These techniques provide exciting research foundations for developing future prompts, leading to better use of language and visual knowledge stored in large models via guidance and interaction, respectively. We discuss three core progressive interaction aspects: interpretable prompt, conversational guidance, and diversified interactions, that researchers will discern explosive development as follows.

Prompt engineering will work from intuitive design to more understandable and interpretable directions. Existing text or visual prompts are more like implicit guidance at the high level, describing what is the downstream visual task. As introduced in Section 3.2, many works attempted to learn prompts to facilitate visual downstream tasks. Despite some progress, they suffered from poor interpretability, i.e., it remains difficult to understand what prompts the network has learned. For example, some works (e.g., VPT) learn unordered token-based prompts, which can not be visualized into an understandable prompt. Chen et al. attempted to learn understandable prompts. Regarding other tuning techniques such as adapter, parameter-based, and remapping ones, they are also faced with the interpretability issue, as they intrinsically aim to reduce the number of tuned parameters for adapting the downstream tasks to the large model. Hence, future research should answer questions such as what are good text and vision prompts, and how to evaluate them throughout the learning pipeline (from the input side to the output side); the relationship between vision and text prompts, and in what situations visual and text prompts can be mutually replaced; How to design explicit, consistent, and logical prompts that enable a large model to adapt efficiently.

2.2 Conversational Guidance

We expect the development of visual tuning will lead to new jobs such as prompt engineers that have expertise in providing guidance to large-scale visual-language models. Multi-round conversational systems can provide a natural platform that guides models to adapt toward the desired task goals . It is generally expected that vision models will homogenize with language models . However, due to the fact that “a picture is worth a thousand words”, the development of visual tuning is somehow behind the success of large language models (detailed in Section 2.1.2). Specifically, concerning data complexity and scenario diversity, the industrial applications of the vision domain (not common application scenarios such as autonomous vehicles, transportation recommendation , whether prediction , protein design , etc.) are highly demanding for customization based on the specific task requirements. Given a tumor detection task, a prompt engineer will select or design good segmentation samples in multiple rounds of conversation with the large model to improve some core steps of the task and eventually achieve acceptable results for production.

2.3 Diversified Interactions

In addition to the interaction in a conversational system with text and visual prompts, interactions in vision can be more diversified. Humans can gradually build up the evaluation standard themselves and then practice (i.e., learn or train for a model) towards higher standards. Existing self-learning models have not set up mechanisms with progressively improving goals. We expect, in the long run, universal AI or strong AI in a specific domain will evolve in the form of prompts, guidance, and diversified interactions. In recent works , a segmentation sample is also used as a prompt to tell the model what task will be performed. Currently, interactions in image synthesis use text prompt and sketch images , enabling everyone to become a visual content creator. These visual interactions can be regarded as a kind of visual prompt based on the image. Image can represent a limited part of visual interaction scenarios, which can be regarded as tasks with static viewpoints but provides basic conditions for more plentiful visual interaction. Current image-based interactions via prompts are also known as in-context learning, which aims to mimic the efficient visual understanding of the human brain and intrinsically narrows the searching space of foundation models . In addition to the simple aspect of navigating large models to downstream tasks, there are more diversified interaction scenarios that provide plenty of egocentric visual interactions such as robots, drones, and bionic robot dogs . These video data provide interactive ecological environments that enable the development of human vision mentioned in Section 2.1.1. Although tuning foundation models pre-trained via self-supervised learning indicates a promising future direction, future visual interactions will rely on advanced pre-training techniques (knowledge accumulation techniques) that are beyond currently tested ones that are based on generating masked pixels or contrastive learning. Promising long-term directions that enable diversified interactions can be involving emerging technologies such as brain-computer interface, quantum computing, event camera, etc. This will lead to new generalization capabilities on top of the future “collection-labeling-training-feedback” cycle system.

Conclusion

This survey summarized visual tuning techniques, particularly focusing on the recent state of visual tuning in the coming regime of large models. Starting from fine-tuning, existing states of prompt tuning, adapter tuning, parameter tuning, and remapping tuning are systematically investigated and compared based on a comprehensive understanding of their technical details. Based on the expected emerging large models, future visual tuning directions are discussed from perspectives of prompt, guidance, interaction, and optimization. We hope this first survey on the latest state of visual tuning will offer a new perspective to researchers in the era of large models, facilitating their research based on a better understanding of the current state and grasping the future core research challenges.

Acknowledgments This work is supported by the National Natural Science Foundation of China (NO. 62276004 and NO. 61932020), Huawei Technologies under Grant No.: P0038941. The authors would like to thank the editors and anonymous reviewers who help improve this manuscript.

References