Review of Large Vision Models and Visual Prompt Engineering

Jiaqi Wang, Zhengliang Liu, Lin Zhao, Zihao Wu, Chong Ma, Sigang Yu, Haixing Dai, Qiushi Yang, Yiheng Liu, Songyao Zhang, Enze Shi, Yi Pan, Tuo Zhang, Dajiang Zhu, Xiang Li, Xi Jiang, Bao Ge, Yixuan Yuan, Dinggang Shen, Tianming Liu, Shu Zhang

Introduction

Since the introduction of the Transformer architecture by Vaswani et al. , deep learning models have experienced remarkable advancements in both parameter size and complexity. Over time, the scale of these models has grown exponentially. Early examples of language models include BERT , T5 , GPT-1 , GPT-2 and various BERT variants . In addition, there exists a multitude of domain-specific BERT variants that are tailored to optimize performance in distinct fields of study or industry . More recently, large language models have become building blocks for general-purpose AI systems and are typically trained on extensive datasets through self-supervised learning . This exponential growth in the scale and complexity of these models has significantly enhanced their capacity to comprehend natural language, allowing them to adapt to various downstream tasks . Notable examples include GPT-3 , ChatGPT , GPT-4 , and others , including domain-specific large language models . This ability to generalize across multiple downstream tasks without explicit training, commonly known as zero-shot generalization, represents a groundbreaking advancement in the field .

Inspired by the success of pre-trained language models in natural language processing (NLP), researchers have ventured into exploring pre-trained visual models in the field of computer vision. These visual models are pre-trained on massive image datasets and possess the ability to understand the content of images and extract rich semantic information. Examples of pre-trained visual models include ViT , Swin Transformer , VideoMAE V2 and others . By learning representations and features from a large amount of data, these models enable computers to more effectively comprehend and analyze images for diverse downstream applications . Moreover, multi-modal visual models, such as CLIP and ALIGN , employ contrastive learning to align textual and visual information. This alignment enables pre-trained models to effectively apply learned semantic information to the visual domain, facilitating efficient generalization in downstream tasks. However, despite their remarkable achievements, these models still face limitations in terms of their generalization capabilities.

The rapid advancements in artificial intelligence (AI) have given rise to a plethora of exciting technological breakthroughs, among which the development of AI systems based on foundational models has emerged as a prominent area of research . This conceptual framework has been coined and unified by AI experts, representing an emerging paradigm in the field . The significance of this concept extends to the notion of emergence, which has become increasingly evident with the rise of machine learning techniques. Emergence manifests itself in the execution of tasks such as automatic inference and the progressive emergence of advanced features and functionalities through deep learning, such as contextual learning. The concept of emergence emphasizes that system behavior is intricately induced rather than explicitly constructed, underscoring the dynamic nature of foundational models in the AI landscape.

Recently, the Segment Anything Model (SAM) has brought about a new trend in solving downstream tasks. Models with prompt engineering modules can solve a wide range of downstream tasks through prompts . These models’ remarkable zero-shot generalization capability highlights the significance of prompt engineering in downstream tasks . However, applying large visual models to specific tasks necessitates an effective approach to guide the model’s learning and inference processes . This is where visual prompt engineering comes into play. It is a methodology that involves designing and optimizing visual prompts to steer large models toward generating the desired outputs.

The emergence of foundation models has unleashed tremendous potential for the advancement of artificial intelligence systems, with particular significance in the domain of computer vision. Visual prompt engineering serves as an adaptive interface and a versatile toolkit that seamlessly integrates with large visual models. By synergistically fusing the capabilities of large visual models with the ingenuity of visual prompt engineering, we empower ourselves to harness the full potential of foundation models, resulting in unparalleled flexibility and efficiency in the realm of image analysis and task resolution. This pioneering integration paves the way for exploring vast frontiers in artificial intelligence applications, unveiling many prospects and ushering in unprecedented opportunities.

This review focuses primarily on prompt engineering methods in computer vision. A collection of relevant literature was gathered by crawling arXiv using the keyword ”visual prompt”. As shown in Figure 1, articles unrelated to computer vision were subsequently filtered out using ChatGPT, resulting in a total account of 500 papers. These papers mainly discuss prompt algorithms pertinent to computer vision, ranging from multi-modal visual-language models to visual and general artificial intelligence models. Prompts take on various forms, including text prompts in multi-modal settings, image prompts, and text-image prompts, each requiring distinct characteristics for different tasks. This paper comprehensively reviews prompt engineering in computer vision, providing insights into different aspects such as multi-modal prompt design, image prompt design, and text-image prompt design, taking into account the specific requirements of different tasks. The aim is to shed light on the advancements and current state of prompt engineering in computer vision, thereby facilitating further research in this field.

2 Outline of the Review

This comprehensive review provides a scientific overview of the latest advancements in computer vision prompts and summarizes the existing design methods in the field. The review is structured as follows:

In the introduction, we traced the evolution of foundational AI models from the inception of the Transformer architecture to the development of large-scale vision models, highlighting how the growth and complexity of these models have spurred the innovative use of prompts (including visual prompts).

Section II presents an overview of key models that have contributed significantly to the advancement of visual prompts and AGI, including Transformer, CLIP, Visual Prompt Tuning (VPT), and SAM. These influential models serve as fundamental references for understanding the following discussions on prompt learning and application in AGI.

Section III delves into visual prompt learning, focusing on multi-modal prompts and visual prompt tuning. Different models and their variants specifically designed for multi-modal prompts are explored, highlighting different approaches and applications in this area. In addition, the section also discusses models and their variants that enable the effective tuning of visual prompts to enhance performance in specific tasks or domains, emphasizing the importance of selecting the appropriate prompt modality for different applications, as exemplified by the use of bounding box prompts in medical image segmentation and text prompts in natural image understanding.

Section IV focuses on the application of visual prompts in AGI models, highlighting the integration of prompts within AGI architectures and showcasing their contributions to strong generalization performance. The latest advancements in visual prompts for AGI models are presented, illustrating how proper prompt design can enable improved performance across diverse domains and tasks.

Section V explores future directions and implications of visual prompts research. We discuss potential developments in the field, taking into account advancements in AGI and related areas. Apart from advancements, the section also addresses the challenges and opportunities associated with them in visual prompts technology and provides insights into the broader implications and potential impact of visual prompts.

The conclusion section provides a summary of the main points discussed throughout the review. The critical role of visual prompts in AGI is emphasized, along with their potential for enhancing performance and generalization. The importance of prompt design and different modalities for different applications is reiterated, and future research prospects in this area are highlighted as well. The conclusion presents a concise recap of the significance of visual prompts and their implications for AGI, serving as a closing remark leaving readers with a clear understanding of the importance of visual prompts.

Background Knowledge

In this section, we will introduce some fundamental concepts, beginning with an elucidation of prompt engineering, followed by an overview of foundational models in the field of computer vision. Within this comprehensive review, our primary focus will be on the key techniques and important approaches pertaining to prompt engineering for large visual models.

In the field of NLP, to achieve parameter-efficient tuning for pre-trained models, prompt-based methods are presented by montaging the inputs with additional context. Since prompt-based approaches can bridge the gap between pre-training and downstream tasks and unleash the potential of pre-trained models, they have witnessed remarkable superiority in various NLP tasks. According to the location of prompts within the text, prompts can be grouped into two shapes. The first is the cloze prompt, which exists in the middle of the text, and the other is the prefix prompt usually attaching at the end of the text. Most previous works design the desired prompts by manually defining or automatically learning. The most simple manner to create the prompts is manually defining according to the common knowledge related to downstream tasks. Brown et. al. manually define specific prompts towards multiple downstream NLP tasks including machine translation and question answering. Schick et. al. leverage pre-defined prompts to boost the few-shot text classification and generation tasks. Although manually constructing prompts is simple and intuitive, they rely on complex different strategies for different tasks and need much professional experience, which are expensive and inefficient. Moreover, pre-designed prompts are usually not the optimal ones and they can’t adaptively cope with many difficult tasks well. To yield more efficient prompt templates, many researches propose to automatically learn optimal prompts via sparse supervisions, which can be divided into two categories including the discrete prompts and continuous prompts. As the natural text context, discrete prompts are automatically searched pre-defined discrete space related to corresponding phrases. For example, Jiang et. al. present MINE as a mining-based method to automatically find prompts with both training inputs and outputs. Wallace et al. design a gradient-based search with input tokens to find short texts related to pre-trained models to generate the desired predictions in an iterative manner. Gao et al. regard the searching prompts as a sequence-to-sequence (seq2seq) generation task and leverage a seq2seq pre-trained model into the prompt searching process. Instead of limit the prompts to natural language in the discrete space, other works aim to automatically construct desired prompts in the continuous text embedding space, which relaxes the searching scope and can be optimized via learnable parameters adaptively using downstream datasets. For instance, Li et. al. prepend a sequence of continuous task-specific sequence to the inputs and keep the pre-trained models frozen. Lester et. al. prepend the inputs with special tokens to construct a prompt and explicitly tune the token embeddings.

With prompts, the gap between pre-trained tasks and various downstream tasks can be narrowed, and performance can be efficiently boosted, potentially approaching the level of full parameter fine-tuning. This demonstrates that an appropriate parameter initialization can be considerably beneficial for downstream NLP tasks.

2 Foundation Models

With the gradual development of large models, the Transformer architecture has emerged as a veritable foundational model, heralding a new era in the field. Transformer was first proposed for translation tasks in NLP, which combines Multi-head Self Attention (MSA) with Feed-forward Networks (FFN) to offer a global perceptual field and multi-channel feature extraction capabilities. The subsequent development of the Transformer-based BERT proved to be seminal in NLP, exhibiting exceptional performance across multiple language-related tasks . Leveraging the great flexibility and scalability of the Transformer, researchers have started to train larger Transformer models, including GPT-1 , GPT-2 , GPT-3 , GPT-4 , T5 , PaLM , LLaMA and others. These models have further advanced the performance and generalization capabilities of the Transformer-based architecture, surpassing human-level performance in certain tasks , and there is still potential for further development in terms of improving the effectiveness of training. Meanwhile, the Vision Transformer (ViT) has extended the application of Transformer architecture to the field of computer vision, bridging the gap between Transformer models in textual and image domains, and validating its feasibility as an unified architecture. Subsequent endeavors in computer vision began to improve and extend ViT model, such as DeiT , Swin Transformer , TNT , MAE , MoCo-v3 , BeiT , etc. These works have successfully applied ViT to diverse vision-related tasks and achieved outstanding performance.

In the Transformer architecture , the smallest unit of feature is a token. This inherent characteristic of the Transformer makes it well-suited for handling multi-modal data, as embedding layers can convert any modality into tokens. Consequently, numerous works in the multi-modal domain, such as ViTL , DALL-E , CLIP , VLMO , ALBEF and others, have adopted the Transformer as the primary framework for multi-modal data interaction, including text-to-image and image-to-text retrieval, image captioning and image/text generation, etc. As the era of large models unfolds, researchers have proposed large multi-modal models such as CoCa , Flamingo , BEiT-v3 , PALI , GPT-4 , with the aim of further enhancing performance on a variety of downstream tasks. In summary, Transformer, as a fundamental model, continues to dominated today’s research in the field.

CLIP

OpenAI has unveiled a groundbreaking vision-language model that leverages the association between images and text to perform weakly supervised pre-training, significantly boosting performance by expanding the available data. The developed work involves collecting a massive dataset of 400 million image-text pairs for training. CLIP , short for Contrastive Language-Image Pre-Training, utilizes fixed human-designed prompts that enable zero-shot prediction and demonstrates superior few-shot capabilities surpassing other state-of-the-art models. The success of CLIP highlights the power of combining visual and textual information and underscores the effectiveness of weakly supervised training using large data . CLIP’s achievement signals a potential breakthrough in the understanding and application of multi-modal techniques, showcasing the ability to capture rich feature representations .such as Group VIT , ViLD , Glip , Clipasso , Clip4clip , ActionCLIP and so on. It offers new insights into the advancement of prompt engineering in computer vision, providing a promising avenue for future developments .

VPT

When adapting large vision models to downstream tasks, modifying the input rather than altering the parameters of the pre-trained model itself is often preferred. This approach involves introducing a small number of task-specific learnable parameters into the input space, allowing for the learning of task-specific continuous vectors. This technique, known as prompt tuning, enables efficient fine-tuning of pre-trained models without modifying their underlying parameters. The VPT approach was the first to address and investigate the universality and feasibility of visual prompts. The proposed VPT method includes both deep and shallow versions, which attain impressive results by learning the prompt and class head of the input data while keeping the parameters of the pre-trained transformer model fixed. This work primarily aimed to demonstrate the effectiveness of visual prompting and provided a novel prompt design perspective. By showing that satisfactory results can be obtained by simply modifying the input, VPT showcased the potential of prompt tuning as an efficient strategy for downstream tasks and opened up new avenues for prompt engineering research .

SAM

In 2023, Meta AI released a project aimed at creating a universal image segmentation model capable of addressing a wide range of downstream segmentation tasks on new data through prompt engineering. To achieve this, the SAM is created. SAM leverages prompt engineering to tackle general downstream segmentation tasks by utilizing the prompt segmentation task as a pre-training objective . To enhance the model’s flexibility in adapting to prompts and to improve its robustness against interference, SAM is divided into three components: the image encoder, the prompt encoder, and the mask decoder. This division effectively distributes the computational cost, resulting in a sufficiently adaptable and versatile segmentation model . SAM’s strength lies in its ability to generalize efficiently across different segmentation tasks, thanks to the prompt engineering approach . This methodology of pre-training on prompt segmentation and fine-tuning on specific downstream tasks helps SAM leverage the knowledge learned from the prompt segmentation task to improve performance on a wide range of segmentation problems, including medical image analysis , video object tracking , data annotation , 3D reconstruction , robotics , image editing , and more. Furthermore, SAM’s modular design allows for flexibility and adaptability to different prompt formats, making it a versatile solution for various segmentation challenges.

Visual Prompts Learning

The research field of multi-modal prompt learning has gained significant attention, with several notable works in the area.

One such work is CLIP , an innovative visual language model that incorporates the concept of manually crafted prompts. In CLIP, prompts take the form of ”a photo of a [class],” with [class] denoting the specific data label. This design allows CLIP to comprehend both visual and textual information, establishing meaningful associations between these two modalities. However, fixed manual prompts have been found to be highly sensitive to results and have a significant impact on outcomes, an issue that has been noted in some studies.

CoOP

To address this problem, CoOP introduces the concept of automatic prompts. Automatic prompts represent the downstream task’s prompt as a trainable continuous vector, enhancing prompt flexibility and adjustability. This approach allows prompts to be optimized based on specific task characteristics, rather than being limited to fixed manual settings. By training learnable prompt vectors, the CoOP model can automatically learn the appropriate prompt representation for different tasks. This flexibility enables the model to better adapt to diverse data and task requirements. However, the CoOP method exhibits lower generalization performance compared to CLIP on new data, which may be due to overfitting on downstream tasks. To address this issue, the authors introduce a lightweight Meta-Net that leverages the outputs of the image encoder and combines them with the trainable prompt. This results in a dynamic prompt that not only is a self-adaptive continuous vector learned regarding downstream tasks but also incorporates image features as conditions. The introduction of this dynamic prompt has significant implications for achieving better generalization performance.

DenseCLIP

DenseCLIP is a novel approach aimed at addressing the challenge of transferring large pre-trained models to dense tasks. While contrastive image-text pairing-based pre-training models demonstrate impressive performance on downstream tasks, the transferability to dense tasks remains a challenge for researchers that has yet to be explored. To address this gap, the researchers introduce an innovative method that incorporates CLIP and prompt patterns into dense tasks for the first time, along with a context-aware prompt that can adaptively adjust based on the specific task and input context. The utilization of context-aware prompts enables better capture of the semantic correlation between images and text in dense tasks, transforming the image-text matching problem into a pixel-text matching problem that improves model performance. DenseCLIP leverages large pre-trained models, such as CLIP, to learn the contrastive relationship between images and text, optimizing the model by maximizing the similarity between matched image-text pairs. This transformation and training strategy allows the model to better comprehend the fine-grained relationship between images and text in dense tasks. The introduction of DenseCLIP provides an innovative approach to transferring large pre-trained models to dense tasks. This method combines context-aware prompts and pixel-text matching problems, offering valuable insights and techniques for addressing the image-text correlation challenge in dense tasks.

MaPLe

The aforementioned work has demonstrated a series of significant advancements in the field of natural language processing, starting from manual text prompts to continuous text prompts and further extending to text-image prompts that incorporate image features. However, relying on prompts from a single modality results in suboptimal model performance. To address this issue, a new approach known as Multi-modal Prompt Learning (MaPLe) is proposed, where continuous prompts are used concurrently across multiple modalities. This approach emphasizes the interplay between text and images in prompt construction, enabling models to be enhanced to a certain extent. This dynamic method of prompt construction facilitates interactions and mutual influences between text and image prompts during model training. The introduction of this approach is significant for achieving more accurate and comprehensive multi-modal understanding. It leverages the interrelated information between text and images and incorporates more contextual and semantic information during the model training process, which further improves model performance.

Imagic

Imagic has proposed a groundbreaking framework for text-guided image editing, introducing complex text-based semantic editing for individual real-world images. For the first time, this framework enables the manipulation of object poses and compositions within an image while preserving their original features. By providing both an original image and a target text prompt, Imagic’s framework allows for precise modifications that align with the semantic context of the image.

GALIP

Generative Adversarial CLIPs (GALIP) , a novel framework, has been proposed to enable text-to-image generation. Building upon the intricate scene understanding abilities and image comprehension of CLIP, GALIP introduces CLIP Visual Encoder (CLIPViT) and a learnable mate discriminator (Mate-D) for adversarial training, harnessing the generalization capabilities of CLIP. Ultimately, the framework employs text-conditioned prompts to adapt to downstream tasks, enhancing the synthesis capabilities for complex images.

PTP

To address the limitations of previous visual language model pre-training frameworks in terms of their lack of visual grounding and localization abilities, a new paradigm called Position-Guided Text Prompting (PTP) has been proposed. PTP introduces a novel approach by encouraging the model to predict objects within given blocks or regress the blocks corresponding to specific objects. This reformulates visual grounding tasks as fill-in-the-blank problems using the provided PTP. Research findings have demonstrated that incorporating the PTP module into several state-of-the-art visual language model pre-training frameworks has led to significant improvements in representative cross-modal learning architectures and benchmark performance.

2 Visual Prompts

The utilization of visual prompts in computer vision tasks can be traced back to interactive segmentation, a technique that requires user input, such as clicks , bounding boxes , or scribbles , to guide the algorithm in accurately identifying object boundaries. These visual prompts provide valuable guidance to the segmentation process, enabling more precise and reliable results. In the context of few-shot image segmentation, the annotated support image can also be considered as another form of visual prompt . By leveraging the information in the support image and its corresponding segmentation mask, the algorithm can generalize and adapt its segmentation capabilities to similar target images. Recently, inspired by the success of prompt engineering in Natural Language Processing (NLP), the computer vision domain has witnessed a series of advancements in utilizing prompts. Researchers have started exploring the idea of formulating prompts as continuous vectors tailored to specific visual tasks. This involves designing prompts as guiding signals for computer vision models, helping them to generate or analyze visual content more effectively. By fine-tuning the models with task-specific prompts, notable improvements have been achieved in various visual tasks, including image classification, object detection, and image generation.

In addition to various prompting techniques in different modes, innovative developments in visual image prompting in the field of computer vision have emerged dramatically. VPT draws inspiration from large NLP models and employs prompt engineering to guide the fine-tuning process on a frozen pre-trained backbone model. It achieves this by introducing a small number of trainable parameters in the input space to serve as prompts. By optimizing these prompts, VPT enhances model performance regarding specific visual patterns and task requirements during the inference process, as depicted in the diagram. VPT is efficient, adaptable, and applicable to a wide range of visual tasks. By utilizing well-designed prompts, researchers can guide the model to perform better on specific visual tasks.

AdaptFormer

Similarly, AdaptFormer , an existing ViT-based model, has been optimized to improve its efficiency on action recognition benchmarks by integrating lightweight modules into its architecture. The design philosophy behind AdaptFormer is to utilize trainable modules that are tailored to the task’s specific constraints within the pre-trained ViT model. The experimental results show that AdaptFormer outperforms fully fine-tuned models in action recognition tasks.

Convpass

Convpass is a methodology aimed at rapidly tailoring pre-trained ViT by implementing convolutional bypasses. The primary goal of Convpass is to reduce the computational expenses associated with fine-tuning while increasing the adaptability of pre-trained models for particular computer vision tasks. Convolutional bypasses, integrated in Convpass, permit accelerated training and inference of models, without compromising performance. This technique aims to simplify the learning process and maximize the efficiency of utilizing pre-trained ViT for targeted computer vision applications.

ViPT

ViPT proposes a prompt-tuning approach to address the challenge of limited large data in downstream multi-modal tracking tasks. This technique allows for the direct utilization of existing foundational knowledge for extracting RGB-modal features in RGB+auxiliary modality tracking tasks. The prompt-tuning method involves fine-tuning the parameters of the base model while retaining the knowledge from the pre-training phase. The prompt module allows for flexible adaptation to task-specific data. ViPT tackles the issue of insufficient data in multi-modal tracking tasks by employing the prompt-tuning method. This approach leverages prior knowledge embedded in the pre-trained model to facilitate the extraction of RGB-modal features. By fine-tuning the base model while preserving its existing knowledge, the prompt module enables efficient adaptation to the specific data requirements of the task at hand. ViPT’s prompt-tuning technique is an effective solution to overcome data scarcity in multi-modal tracking tasks, optimizing the base model’s parameters, retaining valuable knowledge from pre-training, and offering adaptability to task-specific data. This approach ensures logical and fluent integration of RGB-modal features, maximizing performance in RGB+auxiliary modality tracking tasks.

DAM-VP

To address the challenge of handling complex distribution shifts from the original pre-training data distribution when using a single dataset-specific prompt, Diversity-Aware Meta Visual Prompting (DAM-VP) introduces the concept of diversity-aware meta visual prompts. This approach employs a diversity-adaptive mechanism to cluster the downstream dataset into smaller, homogeneous subsets, each with its own individually optimized prompt. The aim is to tackle the difficulties arising from transferring knowledge between different data distributions. Research has revealed that leveraging prompt knowledge learned from previous datasets can expedite convergence and improve performance on new datasets. The integration of diversity-aware meta visual prompts in DAM-VP enables the model to adaptively exploit the diversity within the downstream dataset, facilitating more effective transfer learning and improved generalization.

Visual Prompts in AGI

With the impressive generalization capabilities demonstrated by universal models across various domains, significant progress has been made in large computer vision models as well. By training base models on diverse datasets, these models are capable of adapting to downstream tasks through prompt learning. This approach not only alleviates training demands and optimizes resource utilization, but also introduces new avenues for the development of computer vision. For example, the innovative ”segment anything model” achieves powerful zero-shot transfer capabilities for downstream tasks by employing appropriate prompts, hence its versatile applications in various domains. This universal artificial intelligence model can learn general concepts and exhibit zero-shot transfer capabilities on unknown data on account of its transferability, showcasing the immense potential of general artificial intelligence. Nevertheless, prompt engineering is still pivotal to generalizing the model to new tasks and cannot be overwhelmed by other factors. Prompt engineering refers to the process of designing prompts that enable the model to adapt and generalize to different tasks. The question is crucial as a properly designed prompt can lead to effective representation learning, thereby enhancing its performance on various tasks. The prompt should be emphasized as the key to guiding models to understand and accommodate the context and requirements of either a seen or unseen task.

Many other models have emerged, including notable models such as OneFormer , SegGPT , SEEM , Uni-Perceiver v2 , demonstrating powerful capabilities in general artificial intelligence and providing new possibilities for addressing various tasks. These models employ zero-shot transfer methods, making prompt learning a critical manifestation of model generalization.

In this section, we will delve into a detailed explanation of the prompt construction methods based on promptly, interactive models, encompassing key elements such as object detection, multi-modal fusion, and the combination of various models. These methods provide powerful tools for harnessing the generalization capabilities of large models, thereby offering feasible solutions for achieving downstream tasks.

In the task of object detection, the role of prompt-based methods in attaining the ability to generalize is of utmost importance to achieve general artificial intelligence. These methods are considered a fundamental basis for achieving such generalization. Nevertheless, despite SAM claiming its ability to segment any object, its practical application has been questioned. In particular, concerns have been raised regarding the efficacy of SAM for applications such as medical image segmentation, camouflage object detection, mirror and transparent object detection, and other similar scenarios. As a result, recent studies have concentrated on evaluating SAM’s performance in various settings. These studies have demonstrated that point or box prompts are highly effective in various practical scenarios. SAM has achieved robust zero-shot performance in natural images, remote sensing, and medical imaging domains. However, its ability to generalize in complex application scenarios, particularly where semantic information is ambiguous or environments with low contrast, may not meet task requirements. Therefore, additional research is necessary to enhance SAM’s performance in complex environments. In the realm of crater detection, SAM is leveraged for automated image segmentation. Subsequently, the shape of each segmented mask is assessed, and additional processing steps, such as filtering and boundary extraction, are carried out as well.

In the domain of object counting , researchers adopt SAM by employing bounding boxes as prompts to generate segmentation masks. Dense image features obtained from an image encoder are multiplied and averaged with a reference object’s feature vector. Subsequently, a point grid consisting of 32 points per edge is employed as a prompt for segmentation. The resulting mask feature vector is obtained by multiplying and averaging it together with the dense features. Eventually, to determine the total count, the cosine similarity between the predicted mask and the reference example’s feature vectors is calculated. When the cosine similarity exceeds a predefined threshold, the target object is regarded as recognized. By calculating all of the target objects, the total count is finally obtained.

Remote Sensing SAM

In the domain of remote sensing image segmentation , due to the top-down perspective of remote sensing images, objects within the scene can have arbitrary orientations. Consequently, a technique was proposed to use the Rotated Bounding Box (R-Box) minimum enclosing horizontal rectangle as guidance for SAM segmentation when designing prompts. For the mask prompt, it is defined as the corresponding area enclosed by the bounding box. Previous studies have also demonstrated the suitability and effectiveness of bounding boxes in designing prompts for efficient annotation purposes .

SAM-Adapter

SAM-Adapter has been developed to infuse specialized domain knowledge into the original SAM model, thereby enhancing its capability to generalize across various downstream tasks. The integration has yielded promising outcomes. The adapter is engineered to acquire relevant knowledge and generate task-specific prompts at the preliminary stage. Notably, prompts can be outputted at each Transformer layer. By adopting this approach, prompts included in the segmented network significantly improves the performance of SAM in challenging tasks. The approach has achieved impressive results in disparate domains such as pseudocolor object detection, shadow detection, and medical image segmentation.

2 Multi-modal Fusion

In recent research, the introduction of textual information into visual models, such as images, paintings, frames, and videos, has significantly diversified and improved the fidelity of tasks. Inspired by the concept of generative models, the incorporation of textual prompts into CLIP has emerged as an effective approach.

Text2Seg introduces a vision-language model that relies on text prompts as the input. The model operates as follows: First, the text prompt serves as an input for Grounding DINO, which generates bounding boxes. These bounding boxes guide SAM in generating segmentation masks. Following this, the CLIP Surgery process then generates heatmaps using the text prompts, and the point prompts derived from these heatmaps are fed into SAM. Lastly, a similarity algorithm is applied to obtain the ultimate segmentation map.

SAMText

SAMText introduces a versatile methodology for generating segmentation masks aimed at scene text in images or video frames. The process initiates, once the input is provided, by extracting bounding box coordinates from a scene text detection model, using existing annotations. These extracted bounding box coordinates serve as prompts for SAM, which facilitates the subsequent generation of masks. If the bounding boxes exhibit orientation, SAMText computes their minimum bounding rectangles to obtain horizontal bounding boxes, which, in turn, serve as SAM’s prompts for mask generation.

Caption Anything

Caption Anything has introduced a fundamental model-enhanced image captioning framework, which facilitates multi-modal control encompassing both visual and linguistic aspects. The framework seamlessly integrates SAM and ChatGPT, merging the visual and language modalities such that users can interactively model with the framework. During usage, users initially utilize various prompts, specifically points or bounding boxes, to flexibly control the input image, thereby enabling interactive user manipulation. The framework further refines the output instructions using large language models, ensuring effective alignment with the user’s intended meaning and achieving significant consistency with the user’s intent.

SAA+

Segment Any Anomaly + (SAA+) introduces a novel technique of zero-shot anomaly segmentation that utilizes hybrid prompt regularization to enhance the adaptability of existing foundational models. The proposed regularization prompt incorporates domain-specific expertise and contextual information from the target image, thereby reinforcing more robust prompt and facilitating more accurate identification of anomalous regions. In addition, many works have similarly concluded that incorporating domain expert knowledge as prior support might offer a potential solution to segmentation problems in complex scenes .

3 Combination of Various Models

In complex scenarios, SAM’s performance often lacks robustness, necessitating a new solution that combines interactive methods with efficient tools. The combination of these functionalities presents several potentials for a wide range of fields and exhibits outstanding performance in various tasks.

Image inpainting, a pathological inverse problem that involves restoring missing or damaged parts of an image with visually plausible structures and textures, has been extensively studied in the field of computer vision. Inpaint Anything (IA) has proposed a conceptual pipeline based on the combination of various foundational models. By leveraging the strengths of these models, IA introduces three key functionalities in image inpainting: Remove Anything, Fill Anything, and Replace Anything. The pipeline follows a precise sequence, as shown in the diagram. Initially, a click prompt is employed to automatically segment designated regions, creating masks. Next, state-of-the-art inpainting models such as LaMa and Stable Diffusion (SD) are utilized to fill these masks, effectively completing the removal task. Following this step, a robust AI model like SD leverages a meticulously designed text prompt to generate the specific content required for filling or replacing the voids, facilitating the successful completion of the entire operation.

Edit Everything

Edit Everything introduces a generative system that combines SAM, CLIP, and SD to edit images guided by both image and text inputs. The original image is first segmented into multiple fragments using SAM. Then, the image-editing process is guided by text prompts such that it transforms the source image into the target image, aligning with the provided source and target prompts.

SAM-Track

SAM-Track proposes a video segmentation framework that combines Grounding-DINO, DeAOT, and SAM to enable interactive and automated object tracking and segmentation across multiple modalities. The framework incorporates interactive prompts in the form of click-prompt, box-prompt, and text-prompt in the first frame of the video to guide SAM’s segmentation process. Subsequently, text-prompts are subsequently utilized in the following frames for further result refinement. This versatile framework finds applications in a wide range of domains, including unmanned aerial vehicle technology, autonomous driving, medical imaging, augmented reality, and biological analysis.

Explain Any Concept

Explain Any Concept (EAC) proposes a novel approach to explain concepts based on three pipelines. While SAM excels in instance segmentation, its integration into Explainable AI poses the computational challenge of excessive complexity. EAC solves this by employing SAM for initial segmentation and introducing a surrogate model for efficient explanation. The first stage of the process employs SAM for instance segmentation, followed by the utilization of a surrogate model that approximates the target deep neural network. In the final stage, the trained network is applied to the results obtained in the first stage, facilitating the effective interpretation of the model’s predictions.

Future Directions and Implications

With the continuous advancement of powerful large vision models, the significance of prompts within these models has become more prominent. Designing well-crafted prompts to effectively guide downstream tasks has emerged as a burgeoning avenue for tackling this issue. However, the performance of general artificial intelligence remains constrained by its reliance on domain-specific knowledge. To overcome this limitation, future endeavors should focus on expanding the breadth of knowledge by integrating diverse and comprehensive datasets, employing interdisciplinary methodologies, and fostering fruitful collaborations among experts from various domains. These endeavors will contribute to enhancing the capabilities of AI systems and addressing the challenges associated with leveraging prompts in a more holistic manner.

Large vision models have emerged as a prominent trend in the field of AI, prompting the need to address the challenge of effectively adapting these models to downstream tasks. Several key techniques offer potential solutions in this regard.

Prompt fine-tuning, an essential tool in general artificial intelligence, serves as a crucial step in enabling models to better apply to downstream tasks . By designing suitable prompt examples and predefined inputs, models can be guided to better align with the target task, thus enhancing performance through fine-tuning for improved downstream task completion.

Reinforcement learning enables models to continuously learn and adjust their parameters by leveraging feedback signals obtained from experimentation and errors, thereby maximizing their performance . When combined with prompt fine-tuning, reinforcement learning demonstrates outstanding effectiveness in optimizing adaptive model performance.

Adapter modules enable efficient adjustments for specific tasks within large models by introducing small, functional modules . This approach selectively modifies only certain parts of the model’s structure without requiring significant changes to the overall architecture. Incorporating adapter modules in prompt engineering, not only maintains the integrity of the larger model but also introduces task-specific functional structures, enabling more targeted prompt construction.

Knowledge distillation is a technique that transfers knowledge from a large model to a smaller one. The key lies in compactly representing the knowledge from the larger model and applying it to new tasks, preserving the essential performance and generalization capabilities while facilitating natural deployment in new environments . While prompt fine-tuning is a natural choice for large models, the effectiveness of prompt engineering in smaller models remains an open question. Knowledge distillation can assist in applying prompts to small models by transferring the foundational knowledge and generalization abilities from large models, thereby achieving local deployment of smaller models.

Researchers have already devised a range of model adaptation methods, which have proven effective in various domains. These methods encompass techniques such as transfer learning, domain adaptation, and fine-tuning. While these approaches have their merits, they may not fully address the unique challenges posed by different domains. As the pursuit of effective model adaptation continues, the future holds promise for even greater achievements in applying models to specific tasks across various domains.

2 Challenges and Considerations for Visual AGI

Pioneered by advancements in the field of NLP, the development of a unified framework for large models is poised to become a critical undertaking in the future of visual models. However, in stark contrast to the progress made in NLP, the challenges faced by large models in the computer vision domain are more pronounced.

First and foremost, a key aspect of achieving AGI lies in interactive engagement with the environment and maximizing rewards . While NLP benefits from well-defined learning environments, enabling textual communication and task completion through multi-turn dialogues, the CV domain lacks a clear path and lacks interactive environments . Building a realistic environment for a CV proves exceedingly difficult due to the high costs and risks associated with human-agent interactions. On the other hand, constructing a virtual environment poses challenges when it comes to transferring trained agents to real-world scenarios .

Moreover, the image space exhibits stronger semantic sparsity, domain variations, and infinite granularity compared to the text space . Building upon the immense success of NLP, a foundation for achieving a unified paradigm in the CV domain has been established. Future research endeavors can take inspiration from the development of large NLP models, employing generative pre-training techniques combined with fine-tuning through instructions to achieve a unified approach in the CV field. Additionally, incorporating the capabilities of NLP, and applying multimodal techniques to large vision models can enable the fusion of language and images as an interactive mode for generative pre-training. This, in turn, opens up new avenues for human-machine interaction as a novel pre-training interactive mode.

3 Applications Across Multiple Domains

Visual prompts and large visual models have enabled significant progress in fields where visual understanding and analysis are critical. For example, the prompt-driven SAM has unlocked new opportunities in domains like medical imaging, agriculture, image editing, object detection, audio-visual localization, and beyond . In the medical domain, visual prompts such as segmentation masks, bounding boxes, and key points are used to help detect diseases, quantify the severity of lesions, and analyze medical scans . For the typical treatment sites in radiation oncology, Zhang et al. compared the Dice and Jaccard outcomes between clinical manual delineation and automatic segmentation using SAM with box prompt and proved SAM’s robust generalization capabilities in automatic segmentation for radiotherapy . In agriculture, visual prompts could be used to monitor crop growth, detect weeds or pests, and estimate crop yields . Yang et al. assessed the zero-shot segmentation performance of SAM on representative chicken segmentation tasks and proved that SAM-based object tracking could provide valuable data on the behavior and movement patterns of broiler birds .

By harnessing the advancements in LVM prompts, numerous fields stand to benefit from their integration. The ability to align language and visual data opens doors to improved medical diagnoses , empowering healthcare providers with valuable insights from image-based information. Additionally, the flexibility of LVM prompts enables transformative applications in the realm of natural images, facilitating creative image manipulations and empowering users with powerful editing capabilities. Moreover, the utilization of prompts in video tracking introduces new possibilities for seamless human-machine interaction, allowing for enhanced object detection and precise audio-visual localization.

As research in the field progresses, the increasing utilization of LVM prompts is anticipated to revolutionize various industries and domains. The synergy between language and visual models, facilitated by prompts, paves the way for novel solutions, improved efficiency, and enhanced user experiences across multiple domains. In summary, visual prompts provide annotated data for visual understanding across domains. They give context and guidance, enhancing the ability of machines to interpret visual data. Visual prompts have become an important tool for optimizing visual recognition systems with the increasing use of machine vision.

Conclusion

This review paper provides a comprehensive assessment of the remarkable advancements achieved in the domain of prompt engineering within the field of computer vision. It encompasses a detailed exploration of the design of visual prompts based on the ViT network architecture and the application of prompts leveraging AGI models. From a model-centric standpoint, the study delves into the positive effects of prompts on downstream tasks, highlighting their potential to inspire and enhance the emergent capabilities of large models.

Moreover, this article offers an in-depth discussion of the significance and performance of prompt engineering in various scenarios and domains, emphasizing its pivotal role in the field of computer vision. The article emphasizes the immense potential inherent in prompt engineering, holding promise for groundbreaking advancements in this discipline.

Finally, the paper concludes by providing insights into future research avenues, highlighting the remarkable prospects of utilizing prompt engineering to completely revolutionize computer vision. This technique has the unparalleled potential of improving current models and enabling the creation of novel applications, thus enlightening further research into this area. Given the growing significance of prompts in computer vision in various fields, the outcomes of this study are highly relevant and timely.

References