GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Yu Liu, Kai Chen, Ping Luo

Introduction

Recent advancements of large language models (LLM) have shown incredible performance in solving natural language processing tasks in a human-like conversational manner, for example, commercial products OpenAI (2022); Anthropic (2023); Google (2023); OpenAI (2023) and community open-source projects Touvron et al. (2023a; b); Taori et al. (2023); Chiang et al. (2023); Du et al. (2022); Sun & Xipeng (2022). Their unprecedented capabilities present a promising path toward general-purpose artificial intelligence models. Witnessing the power of LLM, the field of multimodal models Yang et al. (2023b); Huang et al. (2023); Girdhar et al. (2023); Driess et al. (2023) is developing a new technology direction to leverage LLM as the universal interface to build general-purpose models, where the feature space of a specific task is tuned to be aligned with the feature space of pre-trained language models.

As one of the representative tasks, vision-and-language models align the vision encoder feature to LLM by instruction tuning on image-text pairs, such as MiniGPT-4 Zhu et al. (2023), LLaVA Liu et al. (2023a), InstructBLIP Dai et al. (2023), etc. Although these works achieve amazing multimodal abilities, their alignments are only on image-text pairs Chen et al. (2015); Sharma et al. (2018); Changpinyo et al. (2021); Ordonez et al. (2011); Schuhmann et al. (2021), the lack of region-level alignment limits their advancements to more fine-grained understanding tasks such as region caption Krishna et al. (2017) and reasoning Zellers et al. (2019a). To enable region-level understanding in vision-language models, some works attempt to leverage external vision models, for example, MM-REACT Yang et al. (2023b), InternGPT Liu et al. (2023d) and DetGPT Pi et al. (2023), as shown in Table 1. However, their non-end-to-end architecture is a sub-optimal choice for general-purpose multi-modal models.

Considering the limitations of previous works, our objective is to construct an end-to-end vision-language model that supports fine-grained understanding on region-of-interest. Since there is no operation that can refer to specific regions in current image-level vision-language models Zhu et al. (2023); Liu et al. (2023a); Zhang et al. (2023b); Dai et al. (2023), our key design is to incorporate references to bounding boxes into language instructions, thereby upgrading them to the format of spatial instructions. For example, as shown in Figure 1, when the question is “what is doing?”, where the refers to a specific region-of-interest, the model will substitute the embedding of with the region feature extracted by the corresponding bounding box. The region feature extractor can be flexibly implemented by RoIAlign He et al. (2017) or Deformable attention Zhu et al. (2020).

To establish fine-grained alignment between vision and language, we involve region-text datasets in our training, where the bounding box and the text description of each region are provided. The datasets are consolidated from publicly available ones including COCO object detection Lin et al. (2014), RefCOCO Yu et al. (2016), RefCOCO+ Yu et al. (2016), RefCOCOg Mao et al. (2016), Flickr30K entities Plummer et al. (2015), Visual Genome(VG) Krishna et al. (2017) and Visual Commonsense Reasoning(VCR) Zellers et al. (2019a). These datasets are transformed into spatial instruction tuning format. Moreover, we incorporate the LLaVA150K dataset Liu et al. (2023a) into our training process by utilizing an off-the-shelf detector to generate bounding boxes. This enhances our model’s ability to engage in multi-round conversations and generate more human-like responses.

The collected datasets are categorized into two types based on the complexity of the text. First, the plain-text data contains object category and simple attribute information. It is used for pre-training the region feature extractor without impacting the LLM. Second, the complex-text data often contains complex concepts or requires common sense reasoning. We conduct end-to-end fine-tuning of the region feature extractor and LLM for these data.

Benefiting from spatial instruction tuning, our model brings a new interactive experience, where the user can express the question to the model with language and the reference to the region-of-interest. This leads to new capacities beyond image-level understanding, such as region caption and complex region reasoning. As a generalist, our model GPT4RoI also shows its strong region understanding ability on three popular benchmarks, including the region caption task on Visual Genome Krishna et al. (2017), the region reasoning task on Visual-7W Zhu et al. (2016) and Visual Commonsense Reasoning Zellers et al. (2019a) (VCR). Especially noteworthy is the performance on the most challenging VCR dataset, where GPT4RoI achieves an impressive accuracy of 81.6%, 6 points ahead of the second-place and nearing the human-level performance benchmarked at 85.0%.

In summary, our work makes the following contributions:

We introduce spatial instruction, combining language and the reference to region-of-interest into an interleave sequence, enabling accurate region referring and enhancing user interaction.

By spatial instruction tuning LLM with massive region-text datasets, our model can follow user instructions to solve diverse region understanding tasks, such as region caption and reasoning.

Our method, as a generalist, outperforms the previous state-of-the-art approach on a wide range of region understanding benchmarks.

Related Work

The field of natural language processing (NLP) has achieved significant development by the high-capability large language model (LLM). The potential of LLM is first demonstrated by pioneering works such as BERT Devlin et al. (2018) and GPT Radford et al. (2018). Then scaling up progress is started and leads to a series of excellent works, for example, T5 Raffel et al. (2020), GPT-3 Brown et al. (2020), Flan-T5 Chung et al. (2022), PaLM Chowdhery et al. (2022), etc. With the growth of training data and model parameters, this scaling up progress brings to a phenomenal product, ChatGPT OpenAI (2022). By generative pre-trained LLM and instruction tuning Ouyang et al. (2022) on human feedback, ChatGPT shows unprecedented performance on conversations with humans, reasoning and planning tasks Mu et al. (2023); Yang et al. (2023a); Bubeck et al. (2023), etc.

2 Large Vision-Language Model

To utilize high-performance LLM to build up vision-language models, LLM as task coordinator is proposed. Given the user instruction, LLM parses the instruction and calls various external vision models. Some representative works are Visual ChatGPT Wu et al. (2023), ViperGPT Surís et al. (2023), MM-REACT Yang et al. (2023b), InternGPT Liu et al. (2023d), VideoChat Li et al. (2023), etc. Although these models largely expand the scope of multimodal models, they depend on external vision models and these non-end-to-end architectures are not the optimal choice for multi-modal models. To obtain end-to-end vision-language models, instruction tuning LLM on image-text pairs is proposed to align visual features with LLM and accomplish multimodal tasks in a unified way, for example, Flamingo Alayrac et al. (2022), MiniGPT-4 Zhu et al. (2023), LLaVA Liu et al. (2023a), LLaMa-Adapter Zhang et al. (2023b), InstructBLIP Dai et al. (2023), MM-GPT Gong et al. (2023), VPGTrans Zhang et al. (2023a), etc. These models achieve amazing image-level multimodal abilities, while several benchmarks such as LVLM-eHub Xu et al. (2023) and MMBench Liu et al. (2023c) find that these models still have performance bottlenecks when need to be under specific region reference. Our GPT4RoI follows the research line of visual instruction tuning and moves forward region-level multimodal understanding tasks such as region caption Krishna et al. (2017) and reasoning Zellers et al. (2019a).

3 Region-Level Image Understanding

For region-level understanding, it is a common practice in computer vision to identify potential regions of interest first and then do the understanding. Object detection Ren et al. (2015); Carion et al. (2020); Zhu et al. (2020); Zang et al. (2023) tackles the search for potential regions, which are generally accompanied by a simple classification task to understand the region’s content. To expand the object categories, Kamath et al. (2021); Liu et al. (2023b); Zhou et al. (2022); Li* et al. (2022) learn from natural language and achieve amazing open-vocabulary object recognition performance. Region captioning Johnson et al. (2015); Yang et al. (2017); Wu et al. (2022) provides more descriptive language descriptions in a generative way. Scene graph generation Li et al. (2017); Tang et al. (2018); Yang et al. (2022) analyzes the relationships between regions by the graph. The VCR Zellers et al. (2019b) dataset presents many region-level reasoning cases and Yu et al. (2021); Su et al. (2019); Li et al. (2019b); Yao et al. (2022) exhibit decent performance by correctly selecting the answers in the multiple-choice format. However, a general-purpose region understanding model has yet to emerge. In this paper, by harnessing the powerful large language model Touvron et al. (2023a); Chiang et al. (2023), GPT4RoI uses a generative approach to handle all these tasks. Users can complete various region-level understanding tasks by freely asking questions.

Method: GPT4RoI

The overall framework of GPT4RoI consists of a vision encoder, a projector for image-level features, a region feature extractor, and a large language model (LLM). Compared to previous works Zhu et al. (2023); Liu et al. (2023a), GPT4RoI stands out for its ability to convert instructions that include spatial positions into an interleaved sequence of region features and text embeddings, as shown in Figure 2.

We adopt the ViT-H/14 architecture from CLIP Radford et al. (2021) as the vision encoder. Following Liu et al. (2023a), we use the feature map of the penultimate transformer layer as the representation of the entire image, and then map the image feature embedding to the language space using a single linear layer as projector. Finally, we employ the Vicuna Zheng et al. (2023), an instruction-tuned LLaMA Touvron et al. (2023a), to perform further processing.

To extract region-level features with spatial positions, a multi-level image feature pyramid is constructed by selecting four layers from the CLIP vision encoder. These layers are located at the second-to-last, fifth-to-last, eighth-to-last, and eleventh-to-last positions, respectively. We then add feature coordinates Liu et al. (2018) for each level to incorporate absolute position information and we adopt five lightweight scale shuffle modules Zhang et al. (2023c) to improve multi-level feature representation. Finally, we use RoIAlign He et al. (2017) to extract region-level features with the output size of 14×\times14, which maintains sufficient detailed information for caption and reasoning. Moreover, all four level features are involved in the RoIAlign operation and fused into a single embedding as the representation of the region-of-interest (RoI).

2 Tokenization and Embedding

To enable users to refer to regions of interest in text inputs, we define a special token , which acts as the placeholder that will be replaced by the corresponding region feature after tokenization and embedding. One example is depicted in Figure 2. When a user presents a spatial instruction, “What was doing before touched him?”, the embedding of and are replaced by their corresponding region features. However, this replacement discards the references to different regions. To allows LLM to maintain the original references (region1, region3) in the response sequence, the instruction is modified to “What was region1 doing before region3 touched him?”. Then, LLM can generate a reply like “The person in region1 was eating breakfast before the person in region3 touched them.”

Regardless of the user instruction, we incorporate a prefix prompt, “The provides an overview of the picture.” The is a special token that acts as a placeholder, the embedding of which would be replaced by image features of the vision encoder. These features enable LLM to receive comprehensive image information and obtain a holistic understanding of the visual context.

3 Spatial Instruction Tuning

Our model is trained using a next-token prediction loss Liu et al. (2023a); Zhu et al. (2023), where the model predicts the next token in a given input text sequence.

The training details are in Section A.2 in the Appendix.

We transform annotations into instruction tuning format by creating a question that refers to the mentioned region for each region-text annotation. We partition the available region-text data into two groups, employing each in two distinct training stages. In the first stage, we attempt to align region features with word embeddings in language models using simple region-text pairs that contain color, position, or category information. The second stage is designed to handle more complex concepts, such as actions, relationships, and common sense reasoning. Furthermore, we provide diverse instructions for these datasets to simulate chat-like input in this stage.

Stage 1: Pre-training In this stage, we first load the weights of LLaVA Liu et al. (2023a) after its initial stage of training, which includes a pre-trained vision encoder, a projector for image-level features, and an LLM. We only keep the region feature extractor trainable and aim to align region features with language embedding by collecting short text and bounding box pairs. These pairs are from both normal detection datasets and referring expression detection datasets, which have short expressions. The objective is to enable the model to recognize categories and simple attributes of the region in an image, which are typically represented by a short text annotation (usually within 5 words). Specifically, we utilize COCO Lin et al. (2014), RefCOCO Yu et al. (2016), and RefCOCO+ Yu et al. (2016) datasets in this stage.

As shown in Table 2, for COCO detection data, we first explain the task in the prompt and then convert the annotations to a single-word region caption task. For RefCOCO and RefCOCO+, we also give task definitions first and train the model to generate descriptions containing basic attributes of the region. Only the description of the region (in red color) will be used to calculate the loss.

After this training stage, GPT4RoI can recognize categories, simple attributes, and positions of regions in images, as shown in Figure 3.

Stage 2: End-to-end fine-tuning In this stage, we only keep the vision encoder weights fixed and train the region feature extractor, image feature projector, and LLM weights. Our main focus is to enhance GPT4RoI’s ability to accurately follow user instructions and tackle complex single/multiple region understanding tasks. We tailor specific instructions for different tasks. For single region caption, we construct from Visual Genome (VG) region caption part Krishna et al. (2017) and RefCOCOg Mao et al. (2016). For multiple region caption, Flicker30k Plummer et al. (2015) is converted to a multiple region caption task where the caption should include all visual elements emphasized by bounding boxes. To simulate user instruction, we create 20 questions for each caption task as shown in Table 8 and Table 9. For the region reasoning task, we modify Visual Commonsense Reasoning (VCR) Zellers et al. (2019a) to meet the input format requirements and make it more similar to human input. The details of this process can be found in Section A.3.

To improve the capability of GPT4RoI for multi-round conversation and generate more human-like responses, we also involve the LLaVA150k Liu et al. (2023a) visual instruction dataset in this stage. We employ an off-the-shelf LVIS detector Fang et al. (2023) to extract up to 100 detection boxes per image. These boxes are then concatenated with the user instructions in the format “ may feature a class_name”. LLaVA150k significantly improves the capability of GPT4RoI for multi-round conversation.

After completing this training stage, GPT4RoI is capable of performing complex region understanding tasks based on user instructions, including region caption and reasoning, as demonstrated in Section 4.

Demostrations

In this section, we compare the differences between the visual instruction tuning model LLaVA Liu et al. (2023a) and our spatial instruction tuning model GPT4RoI. We demonstrate our new interactive approach and highlight its advanced capabilities in understanding multimodality.

As shown in Figure 4.A, when we try to make LLaVA focus on the center region of the image, it only sees the boy holding an umbrella and a bag, but it misses the book. As a result, LLaVA gives a wrong answer to the question “What is the boy doing” (Figure 4.A.①), and this leads to an incorrect conclusion that “the boy’s behavior is not dangerous” (Figure 4.A.②).

In comparison, as shown in Figure 4.B, our approach GPT4RoI efficiently recognizes visual details using the given bounding box. This allows it to accurately identify the action of “reading a magazine.” Furthermore, GPT4RoI demonstrates its reasoning abilities by correctly inferring that the “boy’s behavior is dangerous”, and giving a reasonable reason that “the boy is reading a book while crossing the street”.

When there are multiple instances in the image (as depicted in Figure 4.C), we attempt to refer to the corresponding instances as “the right” and “the middle”. However, LLaVA provides incorrect information by stating that the right man is “looking at the women” (as shown in Figure 4.C.③). Even more concerning, LLaVA overlooks the actual women in the middle and mistakenly associates the women on the left as the reference, resulting in completely inaccurate information (as shown in Figure 4.C.④ & ⑤).

In comparison, as shown in Figure 4.D, GPT4RoI is able to understand the user’s requirements, such as identifying the person to call when ordering food, and accurately recognize that the person in region1 fulfills this criterion. Additionally, it correctly recognizes that the person in region3 is “looking at the menu”. Importantly, GPT4RoI can also infer relationships between the provided regions based on visual observations. For example, it deduces that the likely relationship between region2 and region3 is that of a “couple”, providing a reasonable explanation that they “are smiling and enjoying each other’s company”.

Quantitative Results

To quantitatively evaluate GPT4RoI, we have chosen three representative benchmarks to assess the region understanding capabilities. These benchmarks include the region caption task on Visual Genome Krishna et al. (2017), the region reasoning task on Visual-7W Zhu et al. (2016), and Visual Commonsense Reasoning Zellers et al. (2019a) (VCR). In order to minimize the impact of specific dataset label styles and make evaluation metrics easier to calculate, we fine-tuned GPT4RoI on each benchmark using different task prompts. More details can be found in Section A.2 in the Appendix.

We report the scores of BLEU, METEOR, ROUGE, and CIDEr for both GPT4RoI-7B and GPT4RoI-13B on the validation set of Visual Genome Krishna et al. (2017). The grounding box in the annotation is combined with the task prompt in Appendix Table 7 to get the response.

The generalist approach GPT4RoI outperforms the previous state-of-the-art specialist model GRiT Wu et al. (2022) by a significant margin, without any additional techniques or tricks. Additionally, we observe that the performance of GPT4RoI-7B and GPT4RoI-13B is comparable, suggesting that the bottleneck in performance lies in the design of the visual module and the availability of region-text pair data. These areas can be explored further in future work.

2 Visual-7W

Visual-7W Zhu et al. (2016) is a PointQA dataset that contains a which box setting. Here, the model is required to choose the appropriate box among four options, based on a given description. For example, a question might ask, “Which is the black machine under the sign?”. This type of question not only tests the model’s object recognition but also its ability to determine the relationship between objects.

To prevent information leakage, we remove overlapping images with the test set from Visual Genome Krishna et al. (2017). The results clearly demonstrate that the 13B model outperforms the 7B model by a significant margin. This finding suggests that the reasoning ability heavily relies on the Large Language Model (LLM).

3 Visual Commonsense Reasoning

Visual Commonsense Reasoning (VCR) offers a highly demanding scenario that necessitates advanced reasoning abilities, heavily relying on common sense. Given the question(Q), the model’s task is not only to select the correct answer(A) but also to select a rationale(R) that explains why the chosen answer is true.

GPT4RoI shows significant improvements over the previous methods across all Q→AQ\rightarrow A, QA→RQA\rightarrow R, and Q→ARQ\rightarrow AR tasks. Notably, in the crucial Q→ARQ\rightarrow AR task, GPT4RoI-13B achieves a performance of 81.6 accuracy, surpassing preceding methods by over 6 points, even outperforming non-open source company-level results, which may take advantage of private data. More importantly, this performance is almost reaching human-level performance of 85.0 accuracy, which shows that the multimodal ability of GPT4RoI is promising to be further developed to human intelligence.

Conclusions

In this paper, we present GPT4RoI, an end-to-end vision-language model that can execute user instructions to achieve region-level image understanding. Our approach employs spatial instruction tuning for the large language model (LLM), where we convert the reference to bounding boxes from user instructions into region features. These region features, along with language embeddings, are combined to create an input sequence for the large language model. By utilizing existing open-source region-text pair datasets, we show that GPT4RoI enhances user interaction by accurately referring to regions and achieves impressive performance in region-level image understanding tasks.

References

Appendix A Appendix

In this appendix, we provide a detailed method architecture figure. We then discuss training-related details, including hyperparameters and instruction templates used in each stage and task. Specifically, we describe how we utilize the VCR dataset. Finally, we analyze some error cases and propose potential improvements for future exploration.

Here is a more detailed framework of our approach, GPT4RoI.

1. We preprocess the input text by adding prefixes to retain both image information and pure text references for each region. 2. Next, we tokenize and embed the text. The image feature and region features will replace the placeholders and respectively. 3. The resulting interleaved sequence of region & image features and language embeddings is then fed into a large language model (LLM) for further processing.

A.2 Training Details

Dialogue model The dialogue model in the demo is trained on 8 GPUs, each with 80G of memory. During the first training stage, a learning rate of 2e-5 is used with a cosine learning schedule. The batch size is 16 for 2 epochs, with a warm-up iteration set to 3000 and a warm-up ratio of 0.003. The weight decay for all modules was set to 0. During the second training stage, the learning rate is reduced to 2e-5 and the model is trained for 1 epoch. To enable end-to-end fine-tuning of the model, which includes a 7B Vicuna, Fully Sharded Data Parallel (FSDP) is enabled in PyTorch to save memory.

Downstream tasks We finetune on three datasets with different learning schedules and task prompts (as shown in Table 7). For the region caption task on Visual Genome Krishna et al. (2017), we perform fine-tuning for 4 epochs with a learning rate of 2e-5. As for Visual-7W Zhu et al. (2016), we observe that it requires a smaller learning rate of 1e-6 to stabilize the training, which is also trained in 2 epochs. On the Visual Commonsense Reasoning Zellers et al. (2019a), we fine-tune the model for 1 epoch using a learning rate of 2e-5.

Instruction of three downstream tasks. The instructions for three downstream tasks are provided in Table 7.

Instruction of Single-Region Caption The instructions for single-region caption are provided in Table 8. We randomly select one as the question in training.

Instruction of Multi-Region Caption The instructions for multi-region caption are provided in Table 9. We randomly select one as the question in training.

A.3 Preprocess of VCR

The Visual Commonsense Reasoning(VCR) dataset Zellers et al. (2019b), comprises 290,000 multiple-choice questions obtained from 110,000 movie scenes. Each image in the dataset is annotated with a question that requires common-sense reasoning, along with its corresponding answer and the explanation for the answer. To construct a sequence of questions, we convert the explanation to a follow-up question and format them into a two-round conversation. Table 10 shows an example of the follow-up question that asks for the reasoning behind the answer.

The VCR dataset is valued for its diverse question-answer pairs that require referencing from prior question-answers to perform reasoning. Therefore, it’s crucial to assign a reference to each region in the dataset. We accomplish this by starting each conversation with a reference to all regions, e.g., There are , … in the image. This approach explicitly references every region, avoiding confusion in future analyses. Additionally, we substitute the corresponding in the answer with category_name at region{i} to ensure a plain text output sequence.

A.4 Failure Case Analysis

Due to limited data and instructions, GPT4RoI may fail in several landmark scenarios. We have conducted a thorough analysis and look forward to improving these limitations in future versions.

Instruction obfuscation As shown in Figure 6.(a), our multiple-region reasoning capability mainly relies on VCR, where we often use sentences that declare , , etc. at the beginning of the question. However, when users adopt the less common sentence structure to refer to regions, it can often be confused with region captions that have the highest proportion in the dataset. As shown in Figure 6.(b), because our data and instructions are mainly generated by rules, our training data does not include content with the "respectively" instruction in multi-region scenarios. This can be resolved by adding specific instructions. In future versions, we aim to develop more diverse instructions, while ensuring data balance.

Misidentification of fine-grained information within in region Although GPT4RoI has improved the fine-grained perception ability of images compared to image-level vision language models, the limited amount of region-level data results in insufficient fine-grained alignment within regions. For example, in Figure 7.(a), the model incorrectly identifies the color of the helmet, and in Figure 7.(b), it misidentifies the object in the girl’s hand. Both cases generate the corresponding answers based on the most prominent feature within the region. Using semi-supervised methods to create more region-level data may address this issue.

A.5 Discussion

In our exploration, we find GPT4RoI produces failure cases as shown in Section. A.4. To further improve the performance, we identify the following potential directions:

Model architecture. We find that 224 ×\times224 input image resolution struggles with understanding smaller regions. However, if we switch to a larger resolution, we must consider the potential burden on inference speed from global attention ViT architecture, while the more efficient CNN architecture or sliding window attention has no available pre-trained large-scale vision encoder like CLIP ViT-H/14.

More region-text pair data. The amount of available region-text pairs is notably smaller than that of image-text pairs, which makes it challenging to sufficiently align region-level features with language models. To tackle this issue, we may try to generate region-level pseudo labels by leveraging off-the-shelf detectors to generate bounding boxes for image-text data.

Region-level instructions. Although we have generated instructions for each task from existing open-source datasets, users in practical applications may ask various questions about an arbitrary number of regions, and the existing data may not contain satisfactory answers. To tackle this issue, we suggest generating a new batch of spatial instructions through manual labeling or by leveraging ChatGPT or GPT4.

Interaction mode. Currently, GPT4RoI only supports natural language and bounding box interaction. Incorporating more open-ended interaction modes such as point, scribble, or image-based search could further improve the user interaction experience.