GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, Ying Shan
Introduction
Recent advances in large language models (LLMs), such as GPT-3 , InstructGPT , and ChatGPT , have demonstrated substantial potential in the area of zero-shot learning and logical reasoning. These models are typically trained on a large volume of text-only data, primarily sourced from the internet. However, as promising as they may seem, these advanced proprietary LLM have significant limitations. One of the major hindrances is the high computational cost associated with these models, which may not be affordable or accessible to many scenarios. Additionally, these models typically depend on specialized data, such as source code and conversation history, which are not easily available to the public.
Instead of solely focusing on language processing, many recent researches attempt to bridge the gap between language models and multi-modal tools. The intelligent agents like Visual ChatGPT and MMREACT have made efforts to meet this goal by sophisticated prompt engineering. These agents utilize a pre-defined template to create instructions that can be executed by vision-language foundation models. Although these approaches have led to impressive results, the primary process of instruction decomposition is heavily based on GPT-3.5 , which is not publicly available, thus limiting further advancements. In addition, equipping these agents with the capability to use tools requires a large amount of data . This brings up an open question: how to efficiently enable a primitive language models to use multi-modal tools?
To achieve it, different from previous studies , we explore a new perceptive as illustrated in Table 1. We propose a simple yet effective method, called GPT4Tools, designed to empower open-source LLMs with the ability to use tools via self-instruct from advanced LLMs. To achieve this, we construct an instruction dataset by prompting advanced teachers (for example, ChatGPT ) conditional on visual content and tool descriptions, which leads to the generation of tool-related instructions. Unlike Toolformer , our method can utilize visual content description to significantly improve data diversity. Furthermore, with the generated instruction-following dataset, we incorporate Low-Rank Adaptation (LoRA) to fine-tune the primitive language models, such as Vicuna and OPT . Besides intrinsic language abilities, by using GPT4Tools, language models can also have the capability to use tools to solve a variety of visual problems. The tasks include visual comprehension and image generation, such as object grounding and segmentation, generating and instructing images, and visual question answering (VQA). With the proposed GPT4Tools, the LLMs not only significantly improves the accuracy of invoking seen tools, but also enables the zero-shot capacity for unseen tools in a zero-shot manner.
We propose an evaluation metric to assess the effectiveness of LLMs in utilizing tools across diverse tasks. With this metric, two human-curated validation sets are constructed to evaluate the LLMs in zero-shot and fine-tuning ways, providing a comprehensive measure of the ability to use tools. To demonstrate the effectiveness of GPT4Tools, we conduct extensive experiments on various language models. The results show the efficacy in teaching LLMs when and how to use tools. Specifically, with the GPT4Tools, the fine-tuned Vicuna-13B achieves 9.3% absolute gains in successful rate over GPT-3.5 , which acquires tool priors in context. Moreover, the fine-tuned Vicuna-13B shows the strong capacity to invoke unseen tools, which is comparable to GPT-3.5 in successful rate.
Our GPT4Tools stands distinct from previous and concurrent studies in three ways. First, our method enables primitive open-source language models to use tools, eliminating the dependence on advanced proprietary LLMs like ChatGPT. Second, we design a new approach based on multi-modal contexts for self-instruction and augmentation, which significantly promote the multi-modal tool usage and can be deployed in different approaches. Third, we propose a new benchmark to assess the effectiveness of using tools, and our method shows remarkable improvements.
Related Work
Vision and Language Model. In the quest to achieve multimodal models capable of addressing both language and vision tasks, several studies have explored methods to enable language models to comprehend visual input. These include techniques such as transforming images into discrete textual representations or projecting continuous image features into the textual feature space . Concurrently, other research has been dedicated to the development of generalist models , which permit a model to simultaneously input images and text, eliminating the necessity for a projection process. For instance, OFA devised a unified sequence-to-sequence decoding architecture applicable to language and object detection tasks. Similarly, Pixel2Pixel converted the outcome of visual comprehension tasks into a series of discrete tokens akin to language tasks. Gato brought together a range of vision and control tasks into a sequential prediction issue, while UViM and Unified-IO advocated for the learned discrete codes as a means to unify an array of vision tasks. By contrast, we in this paper equip the language model with a diverse array of specialized multi-modal tools to process distinct vision tasks. This approach not only promotes the scalability of the model for various tasks, but also avoids the issue of forgetfulness stemming from repeated fine-tuning.
Instruction Tuning. Recent studies have turned out that pre-trained language models could follow natural language instructions and complete various real-world tasks if they are tuned on specific instruction-following data. Notably, InstructGPT , FLAN-T5 , OPT-IML demonstrated remarkable performance on specific tasks after being fine-tuned with instruction data. In order to release the cost of human-written instructions, Self-Instruction found that the instruction-following capabilities of language models can be enhanced by turning on their own generated instruction data. More importantly, this approach inspired a feasible means to improve the zero- and few-shot abilities of language models, i.e., distilling off-the-shelf language models using instructional data from strong ChatGPT or GPT-4 . As a result, many recent works tried to construct excellent language models for various applications based on the LLaMA . For instance, Stanford-Alpaca has employed 52K instructions generated by GPT-3.5 to construct an exceptional dialogue model. LLaVa has adopted GPT-3.5 and GPT-4 to incorporate instruction-following data related to visual content. In this paper, we use GPT-3.5 to construct tool-related instruction datasets, thereby allowing other language models to acquire tool usage capabilities.
Tool Usage. In the Natural Language Processing (NLP) community, several arts sought to endow language models with the ability to use tools. For instance, Komeili et al. proposed to generate conversation responses conditioned on the results of the search engine. LaMDA created a set of tools (comprising an information retrieval system, a calculator, and a translator) to avoid plausible outputs. Lazaridou et al. utilized few-shot prompting on Gopher-280B to enable the search engine to ground its output in factual and current information. Similarly, Visual ChatGPT and MMREACT prompted ChatGPT to invoke visual foundation models. In addition, ToolFormer used self-instruction and bootstrapping to teach GPT-J (6B) using five tools, which include a question and answer system, a calculator, a search engine, a machine translation system, and a calendar. On the contrary, we focus on using the GPT-3.5 model as a powerful teacher to distill off-the-shelf language models and enable them to access many visual models.
Method
Large language models (LLMs) have shown remarkable in-context learning abilities. Among them, ChatGPT and GPT-4 are proven to effectively perform text-annotation tasks or instruct other models to follow instructions of specific domains . Inspired by these findings, we propose leveraging ChatGPT as a powerful teacher to enable off-the-shelf language models to acquire tool usage capabilities. Specifically, we utilize ChatGPT to generate tools-related instruction-following data, which is then used to tune the language model. This process enables the language model to access multimodal information by invoking visual models. Furthermore, we propose an evaluation metric to assess the tool-use ability of the given language model. In the following, we present the data generation, instruction tuning, and evaluation metric in turn.
Data Formation. Upon the collected raw dataset ( items), we apply a filtering process to remove similar instructions, resulting in retained items. Subsequently, we transform the retained data into an instruction-response format utilizing a standardized template as shown in the bottom-left corner of Figure 1. This procedure produces a new dataset, denoted as . The instruction component of incorporates a prefix prompt that encompasses system messages and tool definitions,
Data Augmentation. Although we have successfully acquired instruction-following data related to the tool usage, this simplistic format lacks complexity and depth in both instructions and responses. To mitigate this issue, we augment the generated data from two perspectives:
Negative samples. The generated instructions primarily focus on tool usage, i.e., the decision after the Thought is always "Yes". Consequently, there is a potential risk that the fine-tuned model overfits such a decision. When the user instruction does not connect with the tool usage, the fine-tuned model may erroneously execute irrelevant actions by invoking unnecessary tools. To mitigate this issue, we synthesize negative samples by selecting conversation data from the existing dataset and converting them into the required template, as illustrated in Figure 3 (b). By tuning the model with , it can decide when to use tools.
Context samples. The generated instructions adopt a standard and fixed single-tune format, which lacks a contextual structure. Thus, we augment the dataset by cutting off the chain of action, as shown in Figure 3 (c). Furthermore, we randomly select multiple instructions from and reformat them into multi-turn conversation data. In this way, we synthesize the contextual instruction-following data , which enables the tuned model to call tools within the given context.
So far, we have constructed the tool-related instructional dataset, including positive samples, negative samples, and context samples: .
2 Instruction Tuning
Based on the , we tune the off-the-self language model using its original auto-regressive training objective. To make the tuning feasible, we leverage LoRA optimization,, which freezes the language model and optimizes rank decomposition components of the Transformer layers. For a sequence with tokens, we compute the probability of the target response by:
where denotes the instruction tokens; and is the trainable parameters. In practice, prefix prompt and suffix prompt are also involved, but we here skip them for better readability.
3 Evaluation Approach
Numerous benchmarks typically utilize human-annotated datasets to evaluate the performance of a model. For the purpose of measuring the tool-usage capacity of the language model, we construct a evaluation dataset following the same procedures detailed in § 3.1, and and manually verify the accuracy of each constituent item. This evaluation dataset is partitioned into two components: the first part (validation set) has the same ingredients as the training set, encompassing 23 tools; the second part (test set) comprises 8 novel tools that are absent from the training set. We will use the validation set to validate whether the model can adhere to user commands correctly after tuning with the training set. The test set will be employed to verify whether the model can generalize to new tools after tuning. Based on the human-annotated evaluation dataset with instructions, we design a successful rate to measure the model’s performance from three aspects:
Here, denotes a sequence of arguments, encompassing both the image path and the input text. For instance, ControlNet needs the image path saved conditions (e.g. pose map, depth map, or segment map) and the input text described the user command. represents the quantity of arguments in . When the argument belongs to the image path, equals if the predicted and ground-truth image paths share the same suffix, and otherwise. When the argument is the input text, is equal to the BLEU score between the predicted and the ground truth text.
Experiments
We employ the ChatGPT (gpt-3.5-turbo) as the teacher model to generate the raw instruction-following data. Since this study focused on teaching off-the-self language models to use tools instead of prompt engineering, we adopted the methodology outlined in the Visual ChatGPT to construct tool-related prompts. Our tool pocket consists of 31 tools, including the 23 tools defined in Visual ChatGPT and 8 additional tools (refer to Table 6 in Appendix B for detailed tool names). The training set comprises 71K instruction-response pairs, wherein all instructional data is related to the 23 tools from Visual ChatGPT. We divided the human-annotated evaluation dataset into two parts: validation set and test set. The validation set contains the same tools as the training set, with approximately 50 items associated with each tool. The test set includes tools that are not present in the training set. (further details provided in Appendix A)
Based on the collected data, we tuned language models (LLaMA , Vicuna , and OPT ) with LoRA technology. Specifically, we equipped the projection layers of query, key, value, and output with LoRA layers. The LoRA attention dimension and scaling alpha were set to . While the language model was kept frozen, the LoRA layers were optimized using the AdamW . All models were fine-tuned over 3 epochs, with a batch size of 512. The learning rate was set to , and the maximum length of new tokens was restricted to 2048. Unless otherwise specified, we used Vicuna-13B for the ablation experiments.
2 Main Result
3 Ablation Study
4 Case Study
Figure 5 presents a comparative analysis of our model with Visual ChatGPT and LLaVa . When an image is submitted by the user alongside the instruction "Generate a picture of real people based on the edge", Visual ChatGPT delivers an image that exhibits a weak correlation with the given instruction. Owing to its inability to generate images, LLaVa only returns a caption. In contrast, our model produces an accurate result, thereby evidencing that the tool-related instruction tuning method proposed in this paper can effectively instruct language models in the correct usage of tools. In the Figure 6, we further demonstrate that the Vicuna-13B fine-tuned on GPT4Tools is capable of finishing some visual commands by invoking visual tools. This finding indicates that imparting knowledge to language models regarding the tool invocation could potentially be a way toward the development of a generalist model. More case studies are present in Appendix C.
Limitation
Although the proposed GPT4Tools can teach plug-and-play language models to use tools effectively, it still has some limitations. For instance, the success rate of all models is not , thus further improvements are still necessary for practical applications. Additionally, GPT4Tools teaches the model to explicitly invoke tools using a verbose and fixed prompt (Table 9). This approach reduces the computational efficiency of the model as attention-based architectures compute the relationships between all tokens. Therefore, in the future, it should be explored how to enable the model to implicitly invoke various tools, rather than using the complex prompt. Nevertheless, our GPT4Tools provides a viable approach for equipping language models with multimodal tools.
Conclusion
This paper introduces GPT4Tools, a novel method that enables open-source LLMs to utilize multimodal tools efficiently. We construct a tool-related instructional dataset from advanced ChatGPT and augment them by introducing negative and context samples. Based on the built dataset, we employ LoRA optimization to enhance LLMs’ tool-usage capabilities, thus allowing LLMs to handle various visual tasks, e.g., visual comprehension and image generation. Moreover, we propose a benchmark to assess tool usage accuracy from the decision when to use tools, which tools to use, and arguments of invoked tools. In this benchmark, the LLMs tuned with our GPT4Tools perform comparably to GPT-3.5 on unseen tools. We desire the GPT4Tools to pave the way for one thing, i.e., to equip LLMs with multimodal tools.
Appendix A GPT4Tools Dataset
The training set of GPT4Tools has 71.4K instruction-following data, which includes 35.7K items using tools. Note that these instruction-response pairs are generated from 41K items in since some actions require two tools. The instructional data in the training set involves 23 tools whose names are shown in Table 5 (marked in gray). The distribution of these 23 tools is illustrated on the left of Figure 7. We employ this training set to instruct the language model to invoke tools.
A.2 Evaluation Set.
The evaluation set consists of two parts: validation set and test set.
Validation. The validation set has 1170 samples in total, which includes the same tools as the training set. The number of each tool is almost 50. This set contains some augmented samples as the training set. Thus, it is utilized to verify the effectiveness of the language model in understanding tools after fine-tuning with the training set.
Test. The test set includes 8 tools unseen by the training set. All unseen tool names are marked in black and shown in Table 5, and their detailed definitions are shown in Table 6. The total number of samples is 652, whose distribution is shown on the right of Figure 7. As this set only involves single-turn samples, it is used to evaluate the zero-shot capability of invoking tools by the language model.
Appendix B Prompt
Tool Prompt. The proposed GPT4Tools supports 31 tools, including 23 tools defined in Visual ChatGPT and 8 new tools. They are dependent on image generation models (e.g. ControlNet , Stable Diffusion , InstructPix2Pix , and Shape-E ), and image understanding models (e.g. SAM , BLIP , MMDetection , MMOCR , MMagic , Face Recognition https://github.com/ageitgey/face_recognition, GroundingDINO , and others .). All tool names are summarized in Table 5, where black texts are the newly defined tools. Detailed descriptions of the new tools are illustrated in Table 6, in which the prompt defines the usage scenario of the tool and its arguments.
Generation Prompt. We encouraged the GPT-3.5 (gpt-3.5-turbo) to generate instruction-following data by utilizing the prompt outlined in Table 7. Subsequently, we filtered out noisy instructions, as exemplified in Table 8. Based on the retained data, we performed augmentation following the steps described in § 3.1, resulting in the tool-related dataset.
Tool-Usage Prompt. During replying to the user command, we encouraged the fine-tuned language model to invoke tools by prompt shown in Table 9. In this prompt, the
Appendix C Case Study
Noise During the Generation of Instructions. While ChatGPT or GPT-4 have demonstrated the ability to generate high-quality data , there still are some noises in the generated data. For instance, Table 8 shows three kinds of cases with noise, including the sample with error format, the sample with error arguments, and the sample assigned error tools. Therefore, a practical and effective filtering step is necessary when using data generated by large language models.
Bad Cases of GPT-3.5. As shown in Table 10 and 11, the GPT-3.5 invokes the wrong tools to response the user command. Therefore, when using a language model as a controller to build a generalist model, it is advisable to employ our GPT4Tools to enhance the accuracy of language model actions further.
Appendix D Experiment Settings
In § 4, we benchmark tool-usage ability of the language model using a self-built dataset. The fine-tuning configuration is recorded in Table 12.