ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
Mu Cai, Haotian Liu, Dennis Park, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Yong Jae Lee
Introduction
Large language models (LLMs) like ChatGPT , GPT4 , and Bard have recently gained significant attention for their strong reasoning and generalization capabilities, and their ability to chat in a human-like manner. In particular, models such as GPT-4V(ision) , which incorporate visual information, have demonstrated human-level perception and reasoning capabilities . This has spurred the development of similar open-source models that aim to replicate or even surpass the proprietary models’ performance.
Despite their capabilities, current models, including seminal ones like LLaVA and MiniGPT-4 , focus predominantly on whole-image understanding; in other words, they lack the capability to process region-specific information in complex scenes. This limitation becomes particularly apparent when attempting to describe specific objects within an image using only language prompts, which can be difficult when there is ambiguity (e.g., when there are multiple people in the image, and the question relates to a specific person), as shown in Figure 1.
To address this issue, recent work explores spatial references in multimodal models. Existing efforts have primarily focused on using textual representations of coordinates , learned positional embeddings , or ROI features . However, they often lack user-friendliness, as they are limited to fixed-format visual references like bounding boxes and the spatial coordinates of a mask contour. Most of these approaches, including those by Zhang et al. and Chen et al. , only employ bounding box inputs for visual referrals. While effective in structured scenarios, this method proves less versatile in natural, user-driven interactions where the visual prompts may not conform to clean geometric shapes.
In this paper, we propose a simple yet highly effective solution to this problem: a large multimodal model that can process arbitrary visual prompts. This allows a user to intuitively mark up images and interact using natural cues such as a “red bounding box” or “pointed arrow”. Our model recognizes these visual prompts, offering a user-friendly way to integrate visual references into the language dialogue. Based on our own observation and prior work , which shows that CLIP can understand visual markers, we directly inject the visual prompts into the original image space without any additional region-specific model designs. Although our approach is deceptively simple, it yields an unexpected benefit: our model sets new state-of-the-art performances on tasks demanding precise region-specific perception and complex reasoning. It surpasses the capabilities of existing related models with specialized region encoding techniques, as evidenced by our superior performance on region reasoning tasks on Visual7W and PointQA .
To further support research in this area, we introduce ViP-Bench, a benchmark for evaluating multimodal models’ region understanding capabilities with arbitrary visual prompts. By collecting a diverse set of 303 images and questions, we provide a comprehensive assessment of visual understanding capabilities across six aspects at the region level: recognition, OCR, knowledge, math, object relationship reasoning, and language generation. We believe that ViP-Bench will provide a solid foundation for future research into multimodal models with arbitrary visual prompts.
We introduce a novel multimodal model for intuitive interaction with images using natural language and arbitrary visual prompts, enhancing user accessibility and model flexibility.
We develop a visual referal approach that overlays visual prompts directly onto images, simplifying the model’s architecture without compromising performance.
Our model, ViP-LLaVA, achieves state-of-the-art results on region understanding tasks on established benchmarks, surpassing specialized region encoding models.
We introduce ViP-Bench, a benchmark for evaluating visual prompt interpretation, setting a foundational platform for future research.
Related Work
Large language models like ChatGPT , GPT4 , and LLaMA have shown impressive reasoning and generalization capabilities. The landscape of LLMs has been markedly transformed by the recent introduction of models that integrate visual information, such as GPT-4V(ision) . Building upon open-source LLMs , a vast number of multimodal vision-language models have made significant strides, spearheaded by LLaVA and MiniGPT-4 , which combine LLaMA’s language prowess with a CLIP based image encoder. While these models excel at whole-image understanding, a key challenge has been region-specific comprehension within complex visual scenes. This has led to the exploration of spatial referrals in multimodal contexts. Existing models utilize textual coordinate representations , learned positional embeddings , or Region of Interest (ROI) features to anchor language to specific image regions. However, they often employ rigid visual referral formats that are not as intuitive for users.
Visual Prompting as a User-Friendly Solution.
Our focus is on making the interaction with multimodal models more natural and intuitive. Traditional models have employed regular shapes for visual prompting, but our research is motivated by the need for a system that can interpret a wider range of visual prompts. For example, in visual perception, interactive segmentation methods have been proposed that can take in points or scribbles . Drawing inspiration from recent findings that show GPT-4V’s ability to understand a variety of markers , we advocate for a model that can handle arbitrary visual cues, such as scribbles and arrows. In our model, ViP-LLaVA, we overlay these visual prompts directly onto the image canvas. This is accomplished by fine-tuning on a dataset specifically designed for arbitrary visual prompt instructions.
Evaluating LMM’s Region Understanding Capabilities.
Existing works evaluates the model’s region understanding capabilities on regional multichoice or captioning tasks with metrics such as accuracy, recall, and CIDer . However, these metrics fall short when it comes to evaluating visual dialogue for large multimodal models in an open-world setting. To evaluate LMM’s capability in engaging in visual conversations for image-level understanding, two families of evaluation are proposed: multiple-choice or using GPT4 as a judge for free-form answers . However, a gap still exists in the evaluation of LMM’s capabilities for comprehending arbitrary visual prompts. To address this, we introduce ViP-Bench, a comprehensive benchmark tailored to evaluate how well the LMMs can interpret various visual prompts across multiple dimensions, including recognition, OCR, knowledge, math, relationship reasoning, and language generation.
Approach
Our research hinges on the premise that a large multimodal model should not only perceive the visual content of an image but also interpret arbitrary visual markers as part of the user interaction. In this section, we describe our approach that achieves this goal, highlighting the pivotal role of CLIP in understanding visual markers and the construction of a new instruction tuning dataset tailored to train ViP-LLaVA to understand arbitrary visual prompts.
To effectively recognize the visual prompts, we balance low-level and high-level visual features in ViP-LLaVA.
To address the tendency of CLIP’s deeper features to overlook low-level details , we selectively extract features from multiple CLIP layers. Specifically, we use one early layer (6-th) to encode detailed geometric shapes and four deeper layers (15, 18, 21, 24-th) to capture broader semantic information. These multi-level features are then concatenated, normalized using LayerNorm for training stability, and finally passed through an MLP layer. This process ensures ViP-LLaVA effectively integrates diverse visual cues, a strategy validated through our ablation studies detailed in Sec. 5.4.
Our design’s simplicity of directly overlaying visual prompts offers several advantages. It reduces model complexity by bypassing additional processing modules and aligns closely with natural human interactions, as users often employ diverse and spontaneous visual markers. This flexibility allows ViP-LLaVA to interpret a wide range of user-generated visual cues, enhancing its applicability in real-world scenarios.
This training objective enables the model to generate contextually accurate responses by comprehending the visual content, language instruction, and the overlaid prompts. It fosters the model’s ability to interpret visual markers in unison with the image, thereby enhancing its proficiency in addressing complex, region-specific language inquiries. This capability is crucial for tasks requiring nuanced understanding of both the visual elements and user intentions conveyed through arbitrary visual prompts.
2 Visual Prompting Design
To train the model to recognize and interpret arbitrary visual prompts, we develop a new visual prompt instruction tuning dataset, as there are no prior datasets with arbitrary visual prompts and instruction-output text pairs that we can use.
Our dataset comprises a diverse collection of 520k image-text pairs marked with visual prompts, sourced from publicly available datasets, including (1) single region reasoning data: 80k referring comprehension and generation data from RefCOCOg , and 37k object counting data from PointQA-LookTwice , (2) two-region reasoning data: 80k triplet relationship data from Visual Genome , (3) multi-region reasoning data: 30k grounded image captioning data from Flicker 30k Entities , 213K data from Visual Commonsense Reasoning dataset , and 82k data from Visual7W . Note that all those data are collected from the training split of the aforementioned datasets.
We automatically annotate each image with various visual prompts. For the data that only comes with bounding box annotations, we sample the visual prompts from three possible categories: rectangle, ellipse, and arrow. Here we make sure that the head of the arrow lies within space, where are the width and height of the image, respectively. For ellipse, the lengths along the semi-major and semi-minor axes are inherited from the bounding box size, where we enlarge the ellipse with a ratio between . On the other hand, for regions that come with ground truth pixel-level mask annotations, we annotate each region with visual prompts sampled from the following 8 possibilities: rectangle, ellipse, point, triangle, mask, mask contour, arrow, and scribble created using Bézier curves; see Figure 3. We make sure that the head of the arrow, entire point, triangle, and scribble lies within the provided mask. These annotations simulate natural human interactions with images, where users often use spontaneous markers to highlight areas of interest.
For scribbles, we simulate human-like drawings using Bézier curves . This process begins by randomly selecting three points within the object mask, which serve as the anchors for the quadratic Bézier curve. The generated Bézier curve is then composited onto the image using the previously mentioned alpha blending technique to produce a merged image with the scribble serving as a visual prompt.
Humans naturally use various markers to highlight objects within their environment. For instance, in educational settings, teachers often use arrows or underlining to draw students’ attention to specific parts of an image or text. Similarly, in everyday communication, people might circle items in a photograph to point out something of interest or use scribbles to obscure sensitive information before sharing. Through our design, we create a visual instruction following dataset that mirrors the way humans visually interact with objects, thus fostering a more intuitive and natural interaction with the model.
3 Optional Region-level Instruction Tuning Data
Our training data comes from two sources: (i) region-level visual prompting data described in Section 3.2, and (ii) image-level data devoid of visual prompts, sourced from LLaVA v1.5 . This strategy enables ViP-LLaVA to engage in human-like conversations, primarily due to the image-level LLaVA instruction data from Liu et al. . Optionally, to further enhance ViP-LLaVA’s capability in multimodal conversations at the region-level, we design region-specific instruction data with the help of GPT-4V.
Prior approaches like Shikra attempted to generate region-level instruction data using text-only models like GPT4. However, this method is inherently limiting, particularly in object-level tasks where the model, lacking visual context, cannot accurately reference multiple objects of the same class within a single scene. To overcome this, we develop an instruction data curation method using GPT-4V. Unlike text-only models, GPT-4V can interpret visual prompts displayed in images . Our method involves feeding two images into GPT-4V: the original image and a modified version with annotated visual prompts. Alongside these images, we provide the model with the ground-truth (text) annotation in the original dataset and system messages. This process is used to curate
We introduce specific textual representations such as
Although ViP-LLaVA works well even without this enriched data for standard visual reasoning benchmarks, we find that it helps to further improve the model’s ability to have human-like conversations in open-world settings.
ViP-Bench for Evaluation
In order to rigorously evaluate the capabilities of multimodal models in interpreting and responding to visual reasoning queries, we introduce ViP-Bench, a benchmarking suite for evaluating multimodal region-understanding capabilities under various visual prompts. ViP-Bench consists of 303 unique image-question pairs, where images are collected from MM-Vet , MMBench , and Visual Genome . Each pair consists of an image coupled with a diverse visual reasoning question designed to test a model’s understanding and interpretation capabilities. We reuse the questions in MM-Vet and MMBench (but make minor adjustments so that they take into account the region-specific visual prompts), while in Visual Genome, we design the questions and answers by ourselves. We use bounding boxes and masks produced by the Segment Anything Model (SAM) to annotate the location of the objects.
Key to the design of ViP-Bench is its comprehensive coverage of six crucial aspects of visual understanding at the region level: recognition, OCR (Optical Character Recognition), knowledge, math, object relationship reasoning, and language generation. This range ensures a holistic assessment of a model’s performance in various facets of region-level visual reasoning.
ViP-Bench employs a similar grading mechanism as MM-Vet . We employ the GPT-4 text model, a state-of-the-art language model, to evaluate the responses of multimodal models. Specifically, we feed the response from the multimodal model, the human annotated answer, and several in-context scoring examples to GPT-4. The responses are scored by GPT-4 on a scale from 0 to 10, offering a quantitative measure of the multimodal model’s proficiency in understanding and interpreting visual data. This grading system provides a standardized framework for comparing the performance of different models.
ViP-Bench is meticulously annotated by humans. This process involved seven rounds of validation to ensure the accuracy and relevance of the object boxes/masks, questions, and answers. Such rigorous annotation guarantees the reliability of the benchmark as a tool for model evaluation. An illustrative example in Table 6 showcases a scenario where a leading model like GPT-4V misinterprets object localization under ViP-Bench, highlighting the challenges in current multimodal understanding. We present additional visualizations and statistics of ViP-Bench in the supp.
Through ViP-Bench, we provide a valuable tool for the research community, aiding in the development and refinement of multimodal models. By offering a comprehensive and challenging testbed, we believe ViP-Bench can set the stage for future advancements in the field of visual reasoning and multimodal interaction.
Experiments
In this section, we compare ViP-LLaVA to state-of-the-art multimodal models, including those that explicitly design region-specific modules, perform in-depth analysis to assess ViP-LLaVA’s capabilities, and perform ablation studies.
For the visual model, we choose CLIP-336px to preserve more information from the raw pixel space. We use Vicuna v1.5 as the language encoder. For the multimodal connector, a 2-layer MLP is utilized.
Training and data.
During the initial stage of training, we employ 558k BLIP captioned image-text pairs to pretrain the multimodal connector. The second stage utilizes LLaVA v1.5 instruction data alongside our region-level visual prompting dataset from Section 3.2. Both stages train the model for 1 epoch, with an overall training time of around 20/40 hours for the 7B/13B model using 8 NVIDIA A100 GPUs. Finally, we mix the 13k GPT-4V instruction data with 13k sampled data from stage 2 to get 26k stage 3 training data, and then fine-tune our stage-2 model (referred to as ViP-LLaVA-Base) for one epoch to get our model ViP-LLaVA, which requires approximately 0.5 hours for the 7B model and 1 hour for the 13B model on 8 NVIDIA A100 GPUs.
Visual prompts.
ViP-LLaVA uses 8 visual prompts: rectangles, ellipses, points, scribbles, triangles, masks, mask contours, and arrows. Their attributes, such as color, thickness, and alpha value for alpha blending (in [0.5, 1]) are randomized. The arrow’s direction and length are randomized, with the endpoint remaining within the mask. For referencing specific regions, we replace the
2 Evaluation on Region Reasoning Benchmarks
We first quantitatively evaluate ViP-LLaVA on three region reasoning benchmarks.
The Visual7W dataset tests models’ spatial perception by requiring them to match text descriptions with the correct bounding boxes from a set of choices. We differentiate between ‘generalist’ models, which are not specifically trained on the target dataset, and ‘specialist’ models, which are. For a fair comparison, we use image overlays as visual prompts for the LLaVA model and textual coordinates for Shikra’s text prompts. The results in Table 1 shows ViP-LLaVA-7B outperforming recent state-of-the-art methods, including GPT4RoI and Shikra , despite having fewer parameters, and ViP-LLaVA-13B producing even higher gains. ViP-LLaVA overlays bounding boxes directly onto the image, creating an immediate link between the image and spatial locations. This contrasts with other methods that rely on external embeddings from either textual or newly learned embedding spaces to reference specific regions, proving less effective in this context.
PointQA-LookTwice.
PointQA presents a dataset where queries are based on either a specific point or a bounding box within an image. We evaluate ViP-LLaVA under the broad-question scenario using the bounding box type, typified by the prompt How many of these are there? This requires the model to first correctly identify the object within the given region and subsequently enumerate instances of the same category across the image—essentially a test of object recognition followed by class-specific counting. In line with our methodology for Visual7W, we use the image overlaid with the bounding box for LLaVA, while for Shikra, we incorporate the bounding box coordinates into the text prompt. Table 2 shows ViP-LLaVA’s superior performance on this intricate task, surpassing other multimodal contenders. Our method of overlaying visual prompts ensures the object remains unobscured, effectively combining the original image pixels with visual cues to enhance object recognition and counting accuracy.
Visual Commonsense Reasoning.
The Visual Commonsense Reasoning (VCR) dataset is a challenging benchmark designed to evaluate a model’s capabilities in high-level cognition and commonsense reasoning in the context of visual information. The dataset presents multiple-choice questions that require an understanding of the scene depicted in an image. Each question (Q) is paired with four potential answers (A), where the model must not only select the correct answer but also provide a rationale (R) that justifies its choice, demonstrating the model’s ability to comprehend and rationalize visual elements within a given context.
We finetune ViP-LLaVA-Base-7B on VCR, similar to the protocol in GPT4RoI . As shown in Table 3, our approach exhibits state-of-the-art performance on the validation set, illustrating its proficiency in visual commonsense reasoning. This success highlights our approach’s dual strengths: adeptness in perception tasks and effectiveness in multi-region reasoning. By integrating visual prompts directly into the image, our model more effectively associates spatial locations with semantic understanding, facilitating a better interaction between spatial and semantic reasoning.
3 In-depth Analysis
ViP-LLaVA, when presented with arbitrarily drawn enclosed regions by a user, can accurately describe, e.g., a pedestrian within a small sketched area in Figure 4.
Multi-region understanding capabilities.
ViP-LLaVA demonstrates robust multi-region understanding, able to dissect complex visual scenes and infer relationships between various elements. As shown in Figure 5, ViP-LLaVA is able to infer correspondences between multiple objects in the image, and make the correct reasoning that the red and blue circles both include the train.
Arrow direction understanding.
ViP-LLaVA is able to understand arrows. Here we conduct an ablation study of the arrow direction. Given two arrows that have the same body yet different heads, as shown in Figure 6, ViP-LLaVA is able to understand the direction of the arrows, making correction descriptions about the respective regions.
Generalization to other attributes.
ViP-LLaVA also generalizes to untrained attributes, like varying visual prompt thickness or location, showcasing its adaptability beyond what was seen during training. As shown in Figure 7, ViP-LLaVA can recognize the visual prompts with different thicknesses without explicitly having been trained on them, correctly recognizing the girl in the image. Furthermore, Figure 8 shows that ViP-LLaVA is able to conduct OCR first, and then make correspondences between different regions to make a correct prediction about the content of each part.
4 Ablation Studies
To assess whether overlaying visual prompts on images obscures visual information, we conduct a comparison by inputting visual tokens from both the original and overlayed images into ViP-LLaVA-Base-7B. Using the VCR dataset, we evaluate the accuracy of the QA task with and without the additional visual tokens from the original image. Results on the VCR validation split shows an accuracy of 81.63% with the original image and overlaid image tokens, compared to 82.47% with the overlaid image tokens only. The similar accuracies suggest that the overlaid prompts do not detract from the visual information processed by our model.
Influence of CLIP multi-layer features.
We next explore the impact of using multi-layer visual features from CLIP as opposed to single-layer features, specifically focusing on the second-last layer as implemented in LLaVA . Our ablation study in Table 4 reveals a marked improvement in performance, particularly in scenarios involving multiple visual prompts, as in the Visual7W and VCR datasets. This indicates that leveraging multi-layer visual features significantly enhances the model’s ability to localize and recognize visual prompts within images.
ViP-Bench Evaluation Results
Finally, we evaluate on ViP-Bench using a set of image-level and region-level LMMs, including InstructBLIP , GPT-4V , LLaVA v1.5 , Qwen-VL , Shikra , GPT4ROI and Kosmos-2 . For open-source models, we evaluate with greedy decoding (temperature=0). As shown in Table 5, we first see that the performance of all models, including GPT-4V, is far from perfect, demonstrating the difficulty of ViP-Bench. An illustrative case in Table 6 depicts a scenario where GPT-4V and LLaVA incorrectly predict object localization. Overall, ViP-LLaVA outperforms other models, except GPT-4V, demonstrating greater adaptability to various visual perception and reasoning tasks. By training on images overlaid with visual prompts, ViP-LLaVA becomes adept at understanding arbitrary visual cues and mimicks the natural human method of referring to objects in images. This enables it not only to better identify and interpret visual prompts but also to integrate these prompts into its reasoning process, enhancing its overall comprehension and response accuracy.
In zero-shot evaluation, when visual prompts are represented as a simple list of four textual numerical values, models like Qwen-VL and LLaVA underperform compared to ViP-LLaVA. This underscores the effectiveness of visual prompts over basic textual representations.
Language tasks: A challenge for current LMMs.
The ViP-Bench results reveal that, compared to GPT-4V, open-source LMMs show a significant gap in OCR, math, and language generation tasks, while they perform decently in recognition, knowledge, and object relationship reasoning. This suggests that future VLM developments should prioritize enhancing language reasoning capabilities. For OCR, the results indicate a need for higher resolution inputs or a more robust backbone model, moving beyond the existing capabilities of models like CLIP.
Overfitting Concerns in Region-Level LMMs.
Current region-level LMMs, including Shikra , GPT4ROI and Kosmos-2 , tend to struggle with tasks involving mathematics, relationship reasoning, and language generation. This trend suggests a potential overfitting issue with these models to existing public region-level datasets, which predominantly feature brief descriptions.
Conclusion
In summary, ViP-LLaVA shows that visual prompts are promising for region-specific image understanding. By integrating arbitrary visual prompts, we bridge the gap between user-friendly interfaces and the precision required for region comprehension. ViP-LLaVA’s intuitive design leverages natural linguistic interactions coupled with visual markers, simplifying the process of image annotation while enhancing the clarity of visual references. Our state-of-the-art performance on established benchmarks including Visual7W, PointQA, and VCR, underlines the efficacy of ViP-LLaVA. Notably, the introduction of ViP-Bench as a comprehensive evaluative platform sets a new standard for assessing multimodal models’ region reasoning abilities. ViP-LLaVA establishes a foundation for further exploration in the field of intelligent visual systems. We believe that ViP-LLaVA can motivate how visual and linguistic modalities are integrated, enabling more sophisticated and nuanced human-machine interactions.
This work was supported in part by NSF CAREER IIS2150012, and Institute of Information & communications Technology Planning & Evaluation(IITP) grants funded by the Korea government(MSIT) (No. 2022-0-00871, Development of AI Autonomy and Knowledge Enhancement for AI Agent Collaboration) and (No. RS2022-00187238, Development of Large Korean Language Model Technology for Efficient Pre-training).
References
In-Depth Analysis
ViP-LLaVA, having been trained on eight types of visual prompts—namely mask contour, ellipse, bounding box, triangle, scribble, point, arrow, and mask—exhibits notable generalization capabilities. As demonstrated in Figures 7 and 8 of the main paper, ViP-LLaVA adeptly handles visual prompts with varying thicknesses and diverse markers, even though it was not explicitly trained on such variations. Furthermore, it effectively interprets text markers as visual prompts, a feature inspired by the Set-of-Mark .
Figures 9, 10, and 11 present qualitative examples. In Figure 9, ViP-LLaVA accurately localizes objects tagged with the digits “1”, “2”, and “3”, and generates precise descriptions for each. Figure 10 showcases the model’s ability to recognize digit markers and describe the color of vehicles accurately, despite the markers displaying counterfactual colors relative to the actual vehicle colors. Figure 11 illustrates the model’s competency in localizing a lemon within a scene densely populated with markers.
2 Effect of Optional GPT-4V Region-Level Instruction Data
As mentioned in Section 3.3 of the main paper, incorporating GPT-4V as an additional source of instruction data can enhance ViP-LLaVA’s performance. An example of the curation process is shown in Figure 12. For this purpose, we combine 13K data entries from the original stage 2 instruction dataset with an equal number of GPT-4V region-level instruction data entries, forming a comprehensive 26K-entry stage 3 fine-tuning dataset. We fine-tune our stage-2 model for one epoch, which requires approximately 0.5 hours for the 7B model and 1 hour for the 13B model on 8 NVIDIA A100 GPUs. As shown in Table 8, the fine-tuned model, designated as ViP-LLaVA, demonstrates improvements across nearly all datasets for both the 7B and 13B models, underscoring the efficacy of the GPT-4V instruction data curation process. Notably, even without the GPT-4V instruction data, ViP-LLaVA outperforms contemporary methods on benchmarks such as Visual7W, PointQA-LookTwice, and ViP-Bench. The inclusion of GPT-4V instruction data further amplifies this performance advantage.
3 Understanding Arrow Direction
To rigorously evaluate ViP-LLaVA’s capacity for interpreting arrow directions, we next construct a challenging dataset of examples derived from the COCO validation set . Specifically, we generate multiple scenarios with arrows: each arrow originates from the center of one object’s bounding box and points towards the center of another, and vice versa. These visualizations are depicted in Figure 13. The typical prompt used is as follows: Determine whether object A (category1) or object B (category2) is at the head of the arrow, with the other object representing the tail. It is important to note that we ensure each pair of objects belong to distinct categories. A total of 3520 such paired examples are collected and analyzed. Impressively, ViP-LLaVA-13B achieves an accuracy of 90.28%, demonstrating a robust understanding of arrow directionality and ruling out the possibility of random guessing.
Additional Ablation Studies
To ensure a fair comparison, we conduct ablation studies using the same image encoder (CLIP ViT-L from Radford et al. ), input resolution (224 pixels), and language model (Vicuna v1.1 ) as employed by GPT4ROI . Table 9 presents the results of this analysis. Despite utilizing the same underlying technologies, ViP-LLaVA consistently outperforms on the ViP-Bench evaluations and achieves comparable results on the Visual7W dataset, notwithstanding the fact that GPT4ROI was specifically fine-tuned for Visual7W. These results further reinforce the potential of visual prompting as a more effective approach for region-specific referencing compared to embedding coordinates directly into the language model.
2 Comparing Visual Prompts with Coordinates
To rigorously evaluate the effectiveness of visual prompts versus coordinate-based region referring formats, we next replace visual prompts with textual coordinates embedded in language descriptions. We train a 7B model using identical data and training schedules. The results, as shown in Table 10, indicate that visual prompts significantly outperform coordinate formats on the PointQA-LookTwice and ViP-Bench@Box datasets. Performance on the Visual7W dataset remains comparable between the two formats. These comparisons highlight the superiority of visual prompts as a more effective format for region-specific referencing in complex visual tasks.
Additional Experimental Results
Expanding upon the region perception and reasoning tasks discussed in the main paper, we further evaluate ViP-LLaVA’s region captioning capabilities on the RefCOCOg dataset . This involves fine-tuning the ViP-LLaVA-Base-7B for one epoch subsequent to stage 2 training. As Table 11 illustrates, ViP-LLaVA-Base-7B demonstrates strong performance in region captioning, as evidenced by its scores in both CIDEr and METEOR metrics. These results indicate that visual prompting is not only effective for region-specific referencing and reasoning tasks but also shows promising potential in generating precise and contextually relevant captions for specific image regions.
To evaluate the consistency of ViP-LLaVA-Base-7B, we employ the GPT-4 text model as a judge, conducting five separate assessments. The observed variance in the overall score is a minimal 0.1, indicating stable performance by the GPT-4 judge across multiple evaluations.
Potential of Visual Prompt Augmentation
A key advantage of ViP-LLaVA approach is the ability to very easily employ prompt augmentation during testing. This entails using various sets of visual prompts and aggregating the predictions for a more accurate final answer. For instance, we can modify the prompt from “the woman within a red rectangle” to “the woman marked with a red scribble”, along with corresponding changes in the overlaid image. As shown in Table 12, ViP-LLaVA-Base-7B achieves further improvements through visual prompt augmentation. This process is lossless, unlike textual coordinate representation, where e.g., perturbing coordinates can reduce localization accuracy.
Further Insights into ViP-Bench
Table 13 presents the statistical breakdown of ViP-Bench. The majority of examples focus on recognition capabilities, with a notable proportion (89 examples) requiring Optical Character Recognition (OCR). The proportion of each capability and the combined capabilities are shown in Figure 14 and Figure 15 respectively.
2 Visualizations of ViP-Bench
Figure 16 showcases examples from ViP-Bench, comparing synthesized and human-annotated visual prompts. Panel (a) illustrates tight bounding boxes as synthesized prompts, while panel (b) features human-annotated bounding boxes, highlighting the diversity in human-driven region referring methods. The text prompt that we use to evaluate ViP-Bench performance using GPT4 text model is similar to that used in MM-Vet, which is shown in Table 14. Some examples are shown in Table 7.
3 Examples of capability requirements.
Table 7 presents a selection of examples from our benchmark, demonstrating the diverse capabilities required to complete various tasks, whether they involve single-region or multi-region analysis.
4 Failure cases of GPT-4V
Tables 12.4 to 12.4 display various instances where GPT-4V encountered challenges on ViP-Bench. For instance, Table 12.4 illustrates a case where both GPT-4V and LLaVA-1.5 incorrectly interpret a yellow scribble, with GPT-4V mistaking a yellow circle for the scribble, leading to erroneous responses. In contrast, ViP-LLaVA accurately answers the questions. Another example in Table 12.4 (a) shows GPT-4V incorrectly identifying a person marked by a pink point as holding ski poles and LLaVA-1.5 as holding a green flag, while ViP-LLaVA successfully makes the correct prediction.