LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, Jianfeng Gao, Chunyuan Li
Introduction
A long-standing aspiration in artificial intelligence is to develop general-purpose assistants that can effectively follow users’ (multimodal) instructions to complete a wide range of real-world tasks (Askell et al., 2021; Li et al., 2023c). Recently, the community has witnessed a growing interest in developing foundation models with emergent abilities of multimodal understanding and generation in open-world tasks (Gan et al., 2022; Li et al., 2022). While the recipes of using Large Language Models (LLMs) such as ChatGPT (OpenAI, 2023a) to develop general-purpose assistants for natural language tasks have been proved effective, the recipes of building general-purpose, multimodal assistants for computer vision and vision-language tasks remain to be explored.
Ongoing efforts of developing multimodal agents can be broadly categorized into two classes (Li et al., 2023c): End-to-end training with LLMs, where image-text data and multimodal instruction-following data are collected to continually train LLMs to acquire the ability of processing visual information, resulting in a series of Large Multimodal Models (LMMs). Impressive visual understanding and reasoning performances have been demonstrated by both proprietary models such as Flamingo (Alayrac et al., 2022) and multimodal GPT-4 (OpenAI, 2023c), and open-sourced models such as LLaVA (Liu et al., 2023a) and MiniGPT-4 (Zhu et al., 2023). Although these end-to-end training methods are effective in helping LMMs to gain emergent abilities (such as in-context learning), it remains challenging to develop a unified architecture that can seamlessly incorporate a wide range of skills, such as image segmentation and generation, which are crucial for real-world multimodal applications. Tool chaining with LLMs, where the prompts are meticulously crafted to enable LLMs (e.g., through LangChain lan (2022)) to invoke different tools (e.g., pre-trained vision models) to perform desired (sub-)tasks, without the need of additional model training. Some prominent works include VisProg (Gupta & Kembhavi, 2022), ViperGPT (Surís et al., 2023), Visual ChatGPT (Wu et al., 2023), X-GPT (Zou et al., 2023a), and MM-REACT (Yang et al., 2023b). The strength of these methods is the ability to perform a broad spectrum of visual tasks through the use of (new) tools, which can be incorporated into an AI agent with very low development cost. However, prompting is neither adaptable nor robust enough to allow multimodal agents to always accurately select and activate appropriate tools (from a large and diverse toolset) and compose their results to generate final answers on the fly for real-world multimodal tasks.
In this paper, we present LLaVA-Plus (Large Language and Vision Assistants that Plug and Learn to Use Skills), a general-purpose multimodal assistant that learns to use tools using an end-to-end training approach that systematically expands the capabilities of LMMs via visual instruction tuning. To the best of our knowledge, this is the first attempt reported to combine the strengths of the end-to-end training and tool chaining methods mentioned above. LLaVA-Plus is equipped with a skill repository that contains a wide range of vision and vision-language tools. The design is an embodiment of the “Society of Mind” scheme (Minsky, 1988), where each tool is originally designed for a specific skill and by itself is only useful for specific scenarios, but the combinations of these tools lead to emergent abilities that show signs of higher intelligence. For example, LLaVA-Plus is able to construct a new workflow on the fly, given users’ multimodal inputs, select and activate relevant tools from the skill repository, and compose their execution results to fulfill many real-world tasks that are unseen during model training.
LLaVA-Plus can be continually improved by incorporating new skills or tools via instruction tuning. Consider a new multimodal tool that has been developed for a specific scenario or skill. We collect pertinent user instructions that request this tool and their execution results (or following) to form instruction-following data for tuning. After instruction tuning, LLaVA-Plus expands its abilities as it learns to use this new tool to deal with the tasks that it cannot handle before. LLaVA-Plus also differs from those existing works on teaching LLMs to use tools (e.g., Yang et al., 2023a; Patil et al., 2023), where visual signals are only used when the multimodal tools are activated. In contrast, LLaVA-Plus uses the raw visual signals through the entire human-AI interaction sessions to improve LMM’s ability of planning and reasoning.
In summary, our paper makes the following contributions:
New multimodal instruction-following tool use data. We present a new pipeline for curating vision-language instruction-following data, dedicated for tool use in human-AI interaction sessions, leveraging ChatGPT and GPT-4 as labeling tools.
New large multimodal assistant. We have developed LLaVA-Plus, a general-purpose multimodal assistant that extends LLaVA (Liu et al., 2023a) by incorporating a large and diverse set of external tools that can be selected, composed, and activated on the fly for performing tasks. As shown in Figure 1, LLaVA-Plus significantly extends LMM’s capabilities. Our empirical study validates the effectiveness of LLaVA-Plus with consistently improved results on multiple benchmarks, and in particular, new SoTA on VisiT-Bench with a diverse set of real-life tasks.
Open-source. We will release the following assets to the public: the generated multimodal instruction data, the codebase, the LLaVA-Plus checkpoints, and a visual chat demo.
Learning to Use Tools with Visual Instruction Tuning
Inspired by the impressive performance of multimodal GPT-4 and the open-source LMMs such as LLaVA/MiniGPT-4, the community has witnessed a surge in developing LMMs and the multimodal instruction-following data, following the instruction tuning paradigm (e.g., Liu et al., 2023a; Peng et al., 2023a). In this paper, we use LLaVA as a running example. But note that the proposed recipe can be easily applied to other LMMs. Starting with a user input image query , existing LMMs such as LLaVA typically accept a natural language instruction input from the user, and output a natural language response . Therefore, we can use a unified scheme to represent multimodal instruction-following data as:
where Human and Assistant are special role tokens, and
We propose a modularized system architecture that allows an LMM, working as a planner, to learn to use a wide range of skills at scale, and thus facilitating easy expansion of its capabilities and interface. Specifically, we build a skill repository, where the LMM can leverage a broad range of existing vision and vision-language specialist models as tools for their respective skills when needed, to complete various tasks in the wild. The LMMs in most existing multimodal agents typically perform user-oriented dialogues, where the LMMs are required to immediately respond to user instructions based solely on the knowledge encoded in model weights, as shown in equation 1 and the left part of Figure 2. In addition to this, the LMM in LLaVA-Plus also performs skill-oriented dialogues, where the LMM initiates requests to call appropriate tools from the skill repository, and subsequently aggregate the tool execution results after applying proper skills, as shown in the right part of Figure 2.
We illustrate how LLaVA-Plus works with a full dialogue session in Figure 2. It proceeds in four steps: \raisebox{-0.9pt}{1}⃝ Humans provide a task instruction related to an image . \raisebox{-0.9pt}{2}⃝The LMM-powered assistant analyzes both and , and outputs that chooses the tool from skill repository and writes the appropriate prompt as the tool argument. \raisebox{-0.9pt}{3}⃝ By executing the tool, the result is returned to the assistant. \raisebox{-0.9pt}{4}⃝ The assistant aggregates with and , and outputs to humans. The interaction can be represented as:
Compared with equation 1 which is used to train LLaVA, the only newly introduced component for LLaVA-Plus training is the skill-oriented dialogue. Table 1 illustrates one sequence example of calling detection and segmentation skills in human-AI interactions. LLaVA-Plus is trained with an auto-regressive objective on the sequence of equation 2, where only the green sub-sequences (or tokens) are used to compute the loss, and thus the model learns to predict skill use, answers, and when to stop.
Figure 2 shows that the LMM of LLaVA-Plus needs to perform both user-oriented and skill-oriented dialogues. To this end, we use a unified model prediction format to represent dialogues with and without the need of calling the skill repository. Inspired by Yao et al. (2022), the format consists of three fields, as illustrated in Table 1: Thought is a text sequence representing a reasoning process, which determines whether the skill repository is needed to follow the user instruction, and if so, which tools to use. Action is a list of function calls for the tools to execute the thought. The list is in the JSON format, with each item consisting of two sub-fields: API_name to call the tool and API_params for the corresponding function arguments if applicable. When action is an empty list, no skill is invoked. Value is a natural language response that LLaVA-Plus generates by aggregating tool execution results and the human-AI session history. When presented in of user-oriented dialogues, it is the final response returned to human users. When presented in of skill-oriented dialogues, it is a natural language description about tool execution. In the serving stage, we find it important to ensure a good user experience that we only return the content in the value field of to human users, but hide the entire skill-oriented dialogues unless we need to debug the system.
2 Skill Repository: Multimodal Tool Use Instruct Data Generation
The skill repository of LLaVA-Plus consists of multimodal tools of different skills. To allow the LMM to always activate the most appropriate tools to complete a task, the corresponding tool-use multimodal instruction-following data is needed for LMM tuning. We follow the self-instruct method to curate the data by using GPT-4 as the labeler. Without loss of generality, in this study we want LLaVA-Plus to deal with the scenarios that requires novel skills that LLaVA does not have, e.g., the individual skills for visual understanding, generation, and external knowledge retrieval and the compositions of these individual skills, as summarized in Table 2. In what follows, we treat visual understanding skills as core skills and the others as extended skills, and describe the way instruction data is curated.
Visual understanding skills enable machines to interpret and comprehend visual signals. Existing LMMs have only a limited subset of visual understanding skills, constrained by language inputs and outputs. We expand them to a broader skill set with visual input prompts and visual outputs, including open-set detection and grounding, semantic/instance/interactive segmentation, tagging, captioning, OCR and their compositions, and so on. These understanding skills can be grouped into two categories, depending on whether additional function arguments are required.
Skills with Image-only. The skills without additional function arguments include captioning, tagging, semantic segmentation, caption+grounding, tagging+grounding, and OCR. We have curated training samples for each tool individually. To collect the training samples for a given skill, we fill in the four data variables in equation 2 using different strategies. For , we use GPT-4 to generate a set of instructions that require the use of tools for proper answers. For each sample, we randomly select a question and rewrite it to enhance data diversity. An rewriting example is shown in Table 9 in Appendix. For , its thoughts and value are generated by randomly selecting from some preset responses with rewriting. The actions is known, so it can be directly assigned. is generated with a fixed rule: first presenting the tool outputs and then repeating the initial question. For , its thoughts is created in a similar way to thoughts in , and action is set empty. The value of is the most important field, as it is the visible response to humans in chat. We feed all previous information, including previous questions, the previous tool outputs, and context of the image to language-only GPT-4, which then generates responses to form instruction-following data. Inspired by LLaVA, we consider the ground-truth captions, object coordinates, and object categories as image contexts.
Skills with Additional Function Arguments. Visual skills such as object detection and instance segmentation often require humans to provide very specific instructions regarding the concepts of interests. Their instruction-following data is more challenging to create. We use two methods in this study. The first method is similar to that in the image-only skill setting, where the initial contains a placeholder concept, one or more categories presented in the image are randomly chosen to replace this placeholder, and the final is obtained via rewriting, as shown in Table 9. To allow the LMM to learn more diverse prompts beyond category information, we use GPT-4 to generate questions. Specifically, we manually create two seed samples following the full dialogue in equation 2, send them, together with image contexts, to GPT-4, and ask GPT-4 to generate a full dialogue based on a new image context. An example is shown in Table 10 in Appendix.
2.2 Extended Skills
The LLaVA-Plus recipe can be applied to any tools to improve the system capabilities. We demonstrate its versatility by onboarding multimodal tools of different categorizes. Due to the limited space, we describe the instruction-following data creation process in Section B in Appendix, and summarize the extended skills we have enabled.
External Knowledge. To enable LMMs to use knowledge beyond that encoded in pre-trained model weights, we use the CLIP search API to retrieve external knowledge from LIAON.
Generation. To allow LLaVA-Plus to output images, we use Stable Diffusion (SD) and Instruct-Pix2Pix for image generation and editing, respectively.
Visual Prompts. To better follow human intents, we support various visual prompts for human-AI interaction, such as user-drawn points, sketches and boxes. SAM, Semantic-SAM and SEEM are used for different interactive segmentation tasks.
Skill Composition. To allow LLaVA-Plus to deal with real-world compositional tasks. We curate data for the following scenarios: The scenarios where various visual understanding results of the same image in a multi-turn human-AI interaction session are required. We generate instruction data by applying different tools (including detection, segmentation, tagging, and captioning). Interactive Segmentation + Inpainting. By combining the SAM segmentation results from the user pointing and SD, we enable inpainting with visual interaction. Semantic Segmentation + Generation. By combining the spatial layout from OpenSeed semantic segmentation and ControlNet, we enable instructional visual-conditioned generation. Image Generation/Editing + Social Media Post. It is time-consuming for human users to generate posts that contains both images and text. Thus, we use SD to generate an image, or Instruct Pix2Pix to edit an image, then combine the image with its description generated by a pre-trained LMM to create a multimodal post.
3 Model Training and Serving
To train LLaVA-Plus, we combine the curated tool use instruction data, as shownin Table 2, with the LLaVA-158K dataset. To convert LLaVA-158K into the unified prediction format as described in Section 2.1, we treat the responses in LLaVA-158K as value, and add the fields of thoughts and actions with templates, as illustrated in the example in Table 8 in Appendix. LLaVA-Plus are built in two settings. () LLaVA-Plus (All Tools), where tool use is cast as external knowledge. All visual understanding tools except segmentation in Table 2 are utilized to process the input image, and the extracted recognition results are organized as symbolic sequence representations to enrich the image features in both the training and evaluation stages. () LLaVA-Plus (Fly), where tools are used on the fly. To reduce the cost of calling all tools, we only provide the execution results of related tools for a given instruction. When reporting quantitative numbers, we train models on the 81K understanding instruction data, because existing benchmarks focus mainly on understanding capabilities. When building demo systems, we train our models on the full dataset.
LLaVA-Plus is served using the FastChat (Vicuna, 2023) system, which is composed of web servers that interface with humans, model workers that host the LMM and multiple tools, and a controller to coordinate the web-server and model workers. The 7B LLaVA-Plus and all the tools can be loaded and served in a 80G GPU.
Related Works
We summarize the connections and differences between LLaVA-Plus and existing general-purpose multimodal systems in Table 3, where only representative methods are shown due to space constraint. They can be broadly categorized into two classes as discussed below.
There is a growing interest in exploring a paradigm of building general-purpose AI agents that synergistically leverage multiple tools with LLMs to solve sophisticated, open-world problems. The idea is originated in NLP to invoke general tools whose skills are lacked from LLM (e.g., ToolFormer (Schick et al., 2023), ChatGPT-Plugin (OpenAI, 2023b)), and is recently extended to the multimodal space. There are two ways to leverage multimodal tools with the LLM as a planner to determine which tools to invoke: tool chaining by prompt engineering and in-context-learning, such as Visual ChatGPT (Wu et al., 2023), MM-ReAct (Yang et al., 2023b), and instruction tuning of LLM with a focus on multimodal tool use, such as GPT4Tools (Yang et al., 2023a) and Gorilla (Patil et al., 2023). LLaVA-Plus represents the first work of utilizing the LMM as the planner for tool use, where image inputs are considered throughout the entire interaction sessions for improved user experience.
Inspired by the success of a unified architecture of LLMs to complete many language tasks, the AI community has witnessed an increasing interest in building unified models with versatile multimodal capabilities. Proprietary models such as Flamingo (Alayrac et al., 2022) and multimodal GPT-4 (OpenAI, 2023c) (or GPT-4V (OpenAI, 2023d)) have demonstrated strong multimodal performance on zero-shot task transfer, which quickly inspired their open-source counterparts: LLaVA, MiniGPT-4, Open-Flamingo (Awadalla et al., 2023), Otter (Li et al., 2023a), to name a few. These LMMs can deal with the tasks with image-text input and text output. The capabilities have been extended to support the tasks with image-text output, such as image editing and segmentation, as demonstrated in CM3Leon (Yu & et al, 2023), Emu (Sun et al., 2023), and GILL (Koh et al., 2023). Bounding box outputs for grounding are recently supported, as shown in Kosmos-2 (Peng et al., 2023b), Shikra (Chen et al., 2023a) and DetGPT (Pi et al., 2023). GPT4ROI (Zhang et al., 2023c) allows users to select regions of interest with bounding boxes for human-AI visual chat. BubaGPT (Zhao et al., 2023) and LISA (Lai et al., 2023) use an extra referring segmentation model to enable the mask prediction capability. Compared with them, LLaVA-Plus enables a much wider range of multimodal skills and their compositions, as illustrated in Table 3.
Experiments
We consider two benchmarks. LLaVA-Bench (Liu et al., 2023a) evaluates the visual chat of LMMs, with three types of questions: conversation, detailed description and visual reasoning. It consists of two datasets: the COCO set containing 30 COCO images and 90 chat questions, and the In-the-Wild set containing 24 web images with 60 questions. Language GPT-4 (gpt4-0314) is used to score the generated answers. The relative scores between the model output and gold response are reported. SEED-Bench (Li et al., 2023b) evaluates the image-level and instance-level perception and reasoning of LMMs, with 19K multi-choice questions. The results are shown in Table 4. Both LLaVA-Plus variants outperform LLaVA on these two benchmarks, demonstrating the effectiveness of adding visual recognition results of applying new skills in the LMM pipeline. LLaVA-Plus (All Tools) shows superior performance to LLaVA-Plus (Fly) because the former leverages more tools as additional contexts. We further conducted several ablations: We tried to directly add the skill execution results in the testing stage of LLaVA, shown as the row of LLaVA (Tools in Test). The degraded performance compared with LLaVA demonstrates the necessity of learning to use skills in training. We removed thoughts in the unified data format and observed a performance drop, indicating chain-of-thoughts style data format is beneficial. GPT4Tools trains an LLM for multimodal tool use. Its lower performance indicates that visual instruction tuning of tool use in LLaVA-Plus is important.
To study the novel capabilities enabled by learning to use skills, we create an evaluation set LLavA-Bench (Tools), which measures four capabilities (grounding, tagging, caption, and OCR) with 10, 12, 12, and 10 samples in each. In Table 5, we also compare against the commercial visual chat systems such as Microsoft BingChat and Google Bard. LLaVA-Plus significantly outperforms the others on this benchmark, mainly because the other systems are not equipped with some of these capabilities. By comparing with chaining tools with GPT-4 (row of “All tools + GPT4”) and MM-REACT, we demonstrate the advantage of training an open-source LMM as a planner for tool use.
2 Comparisons with SoTA LMM systems
MMVet (Yu et al., 2023) contains 200 images and 218 questions, aiming to evaluate six core vision-language (VL) capabilities and their combinations. For evaluation, an LLM-based evaluator (gpt4-0613) is used to score open-ended outputs of different forms. The results are reported in Table 6. LLaVA-Plus consistently outperforms LLaVA on both 7B and 13B model sizes. The categories with most significant improvements are OCR and spatial, indicating the positive impact of the corresponding visual skills on LMM outputs.
VisIT-Bench (Bitton et al., 2023) is a real-world use oriented LMM benchmark, comprising 592 questions and 1,159 public images categorized into 70 instruction families. The results are shown in Table 7, which summarizes the battles between LMMs with GPT-analog human judgment. Elo ratings are computed by treating each pairwise human judgment as a “match”. The difference between the Elo ratings of two models provides an estimate for the win probability when pitting model A vs. model B. The “#matches” column indicates the number of total matches in which a particular model participates. Win-rate indicates the win rate of a model against the human-verified reference outputs. LLaVA-Plus significantly outperforms the leading method LLaVA by 100+ ELO score, achieving a new SoTA on the leaderboard.
3 Visual Examples of New Capabilities
In Table 3, we illustrate new capabilities of LLaVA-Plus with visual examples. Please see Section D in Appendix for many other interesting scenarios that demonstrate the versatile capabilities of LLaVA-Plus by learning to use skills and their compositions.
In the left example, the questions require identifying the precise object locations. LLaVA-Plus can successfully detect the frisbee’s coordinates, which help determine its status of flying in the air and thus describe the outdoor scene/activity. The same example is shown to Bard, Bing Chat, MM-REACT and LLaVA in Figure 6 in Appendix. They all fail, revealing the lack of grounding ability.
In the right example, we illustrate an interactive image editing scenario, where users aim to see the spatial layout of the scene first and then generate an image of a similar layout, but with a new “under water” scene. The LMM not only applies the correct skills, but also generates a function argument “A bicycle parked next to a bench under the sea” for conditional image generation. This reveals the appealing property of LMM as a planner, as it can see the raw image, and provide necessary image analysis results throughout the human-AI interaction process. More such examples are in Appendix Figure 11.
In the bottom example, we show that LLaVA-Plus can be used to help create multimodal social media posts. For example, when capturing an image, the user wants to post the same image in an autumn scene and associate the image with some attractive text to post Instagram. LLaVA-Plus can use the editing skills to revise the image, and combine the context of visual images and their related language topics to suggest several caption options. In Appendix Figure 12, we create all four seasons for the same scenarios, and observe that LLaVA-Plus can follow the instruction to easily switch among them while consistently maintaining the original image cue.
Conclusion
We have presented LLaVA-Plus, a general-purpose, multimodal assistant which is based on an LMM that plugs and learns to use skills to complete a wide range of vision-language tasks in the wild. The first visual instruction dataset specifically designed for multimodal tool use has been collected for model training. By incorporating the execution results of new skills, LLaVA-Plus consistently outperforms LLaVA across many benchmarks, creates a new SoTA and shows emergent multimodal interaction capabilities. However, LLaVA-Plus is limited due to hallucinations and tool use conflicts in practice. There are interesting problems yet to be addressed in future research on building reliable general-purpose multimodal AI agents.
To ensure the reproducibility of our research, we will publicly release a comprehensive set of assets including the generated multimodal instruction data, our codebase, the LLaVA-Plus checkpoints, and a visual chat demo. Additionally, we have ensured complete transparency by elaborating on every facet of our training data collection and model training within this paper, as shown in Sec. 2.
References
Appendix A Data
Augmenting LLaVA data. The original LLaVA data only consists of questions and answers. We need to augment this data to make it match with our regular data format. We transformed the original answers in LLaVA into a part of the values field, then added an empty list for actions, and generated a thoughts using ChatGPT. The thoughts should indicate that the model can answer the question without invoking any tools. An example is shown in Table 8 in Appendix. We found the model cannot invoke tools if we did not unify the two data formats.
Details on data generation. The pipeline to generate questions for visual prompts is shown in Table 4. The pipeline to generate questions with image-related parameters is shown in Table 5. An example of rewriting questions using GPT4 is shown in Table 9. The self-instruct example to generate multi-turn conversation for detection is shown in Table 10.
Appendix B Extended Skills
To enable LMMs to gain knowledge beyond that encoded in pre-trained model weights, we use the CLIP search API to retrieve external knowledge from LIAON. We utilize the images and questions from the InfoSeek dataset, and generate the other fields of the training sequence by following the image-only skill data creation pipeline. Input images are considered as queries, and image-to-text retrieval is performed to get top-K items for each query. To encourage the LMM to leverage external knowledge, we only consider the subset of questions whose ground truth answers can be extracted or derived from the retrieved knowledge. This subset can be selected using ChatGPT that compares the answers and retrieved knowledge.
For image generation, we employ Stable Diffusion (SD) as the tool, and generate instruction data based on the JourneyDB dataset due to its high quality in language prompt and images. We ask ChatGPT to generate human-like instructions based on the original, detailed prompt for image generation, focusing on the scenarios where human-specified instructions are ambiguous and short, and thus cannot easily align with the prompt distribution of SD. Similarly, we use Instruct-Pix2Pix for image editing. The Instruct Pix2Pix dataset contains both instructions and prompts of source and target images. We directly use their editing instructions and follow the image-only skill data creation pipeline to fill the other fields.
The visual prompt data is constructed similarly to that for visual understanding skills, except that additional visual inputs, such as user-drawn points, sketches and boxes, are required. Take SAM as an example. A point is required as input for interactive segmentation. We simply generate a random point and then convert it into a text sequence, and append it to a user question to form a concatenated text sequence , which is a standard format that LMMs such as LLaVA can deal with. Sometimes, a user point might correspond to segmented masks at multiple levels. To support this skill, we use Semantic-SAM (Li et al., 2023d) to create training data where the multi-granularity segmentation functionality is explicitly specified by instructions.
The scenarios described so far are designed to create training samples for single-skill tasks. However, many real-world scenarios often require some compositions of several skills. To allow LLaVA-Plus to deal with such compositional tasks, we have curated instruction-following data for compositional skills as follows. Various visual understanding results of the same image can be requested. To teach an LMM to learn to use multiple skills in a multi-turn human-AI interaction session, we generate instruction data by applying different tools (including detection, segmentation, tagging, and captioning) to the same image from COCO, combining the results with LLaVA instruction data, and then randomly mixing these datasets. This produces instruction data that simulates users’ behavior of using multiple tools to deal with real-world tasks. Interactive Segmentation + Inpainting. In one editing scenario, we ask a user to specify an area of an image with visual pointing along with language instruction. We then combine the SAM segmentation results and the SD inpainting results to create an instruction-following sample. Semantic Segmentation + Generation. In another image editing scenario, we ask a user to specify the spatial layout of an image, using an user-provided image and a language instruction. We then combine the OpenSeed semantic segmentation results and ControlNet conditional generation results to create an instruction-following sample. Image Generation/Editing + Social Media Post. It is time-consuming for human users to generate posts that contains both images and text. Thus, we use existing tools to create large amounts of multimodal posts for model tuning as follows. We use SD to generate an image, or Instruct Pix2Pix to edit an image. We then combine the image with its description generated by a pre-trained LMM to create a multimodal post.
Appendix C Results
We aim to investigate the potential of the LMM in enhancing existing tools. A comparison of three distinct models on the COCO caption benchmark is presented in Table 11. We employed BLIP2 as our primary captioning tool and hence, use it as the benchmark model. Additionally, the original LLaVA is also included for reference. The enhanced LLaVA-Plus model refines BLIP2’s outputs, leading to richer details.
The table reveals that LLaVA-Plus outperforms the others in terms of the CLIP score. Intriguingly, both language models exhibit subpar performance on language-language metrics. A striking observation is the significantly lower CIDEr scores for these models when juxtaposed with BLIP2.
Grounding DINO, despite its commendable object detection prowess, occasionally exhibits hallucinations, leading it to generate false positive instances. Our LLaVA-Plus model, capable of simultaneously analyzing model outputs and image content, holds the potential to reduce such false positives.
To harness this potential, we crafted examples using negative prompts from COCO and directed the model to eliminate false positive outputs. We subsequently evaluated the model on the first 100 images from the COCO validation set. By using all negative categories of an image as prompts, we gauged the presence of false positive objects. The results are tabulated in Table 12.
The results show that Grounding DINO has a high possibility of resulting in false positive examples. With the LLaVA-Plus model, it can help to reduce the false positive rate significantly.
Appendix D Example Scenarios
We show more scenarios of LLaVA-Plus in leveraging new skills to improve visual chat experience.
Figure 6 compares object localization capability of LLaVA-Plus with Bard, Bing Chat, MM-REACT and LLaVA. It turns out the commercial visual chat do not have the ability to tell the object spatial location, while LLaVA-Plus can successfully identify the object location and thus describe the outdoor scene and activity correctly.
Figure 7 (a) shows an example to detect and count the number of objects. Figure 7 (b) shows a real-life scenarios to pick up the appropriate tools and teach the users how to use them. Compared langauge-output-only LMM such as LLaVA/GPT-V, identify and visualization the location of object is an more intuitive approach for users to comprehend. Figure 8 provides object segmentation results, but enriched with language descriotion at the instance level. It is the synergy of LMM and segmentation that improve the enhanced fine-grained understanding.
In Figure 9, we compare LLaVA-Plus and LLaVA in terms of generating response with detailed facts and entities. The retrieval external knowledge of LLaVA-Plus introduces more relevant information that allows LLaVA-Plus to ground in generation.
In Table 10, we show that LLaVA-Plus can produce detailed SD-favored language prompts for image generation, based on the high-level and brief requests. This can help improve image generation quality.
Figure 11 demonstrate the multi-turn interactive image segmentation and editing capabilities. By leveraging OpenSEED, LLaVA-Plus can apply the skill of full-image semantic segmentation to group pixels of the same object together, providing the spatial layout of the scene. With further requests to produce new images that follow the same layout but change other aspects, the corresponding editing skills can be executed, through InstructPix2Pix and ControlNet.
In Figure 12, the four seasons of the same scene are used as instructions to ask LLaVA-Plus to provide the edited images and attractive texts. Another example on fireworks is shown in Figure 13
Figure 14 demonstrates the use of semantic SAM to support visual pointing on the image from humans, after which multiple segmentation masks at different levels are shown. Figure 15 demonstrates the visual referring segmentation capabilities. LLaVA-Plus allows humans to specify the segmentation intents on the object of interest with the selected regions from another image. This is useful because some concepts can be hard described in language, but easier to express with reference visual regions.