LISA: Reasoning Segmentation via Large Language Model

Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, Jiaya Jia

Introduction

In daily life, users tend to issue direct commands like “Change the TV channel” to instruct a robot, rather than providing explicit step-by-step instructions such as “Go to the table first, find the TV remote, and then press the button to change the channel.” However, existing perception systems consistently rely on humans to explicitly indicate target objects or pre-define categories before executing visual recognition tasks. These systems lack the capacity to actively reason and comprehend users’ intentions based on implicit instructions. This self-reasoning ability is crucial in developing next-generation intelligent perception systems and holds substantial potential for industrial applications, particularly in robotics.

In this work, we introduce a new segmentation task — reasoning segmentation, which requires generating a binary segmentation mask based on an implicit query text involving complex reasoning. Notably, the query text is not limited to a straightforward reference (e.g., “the orange”), but a more complicated description involving complex reasoning or world knowledge (e.g., “the food with high Vitamin C”). To accomplish this task, the model must possess two key abilities: 1) reasoning complex and implicit text queries jointly with the image; 2) producing segmentation masks.

Inspired by the exceptional capacity of the Large Language Model (LLM) to reason and comprehend user intentions, we aim to leverage this capability to address the aforementioned first challenge. However, while several studies have integrated robust reasoning capabilities into multi-modal LLMs to accommodate visual input, the majority of these models primarily concentrate on text generation tasks and still fall short in performing vision-centric tasks that necessitate fine-grained output formats, such as segmentation masks.

In this work, we introduce LISA: a large Language Instructed Segmentation Assistant, a multi-modal LLM capable of producing segmentation masks. To equip the multi-modal LLM with segmentation abilities, we incorporate an additional token, i.e., , into the existing vocabulary. Upon generating the token, its hidden embedding is further decoded into the corresponding segmentation mask. By representing the segmentation mask as an embedding, LISA acquires segmentation capabilities and benefits from end-to-end training. Remarkably, LISA demonstrates robust zero-shot abilities. Training the model solely on standard semantic segmentation and referring segmentation datasets yields surprisingly effective performance on the complex reasoning segmentation task. Furthermore, we find that LISA’s performance can be significantly enhanced by fine-tuning on just 239 image-instruction reasoning segmentation pairs. As illustrated in Fig. 1, LISA can handle various scenarios, including: 1) complex reasoning; 2) world knowledge; 3) explanatory answers; and 4) multi-turn conversations.

In addition, to validate the effectiveness, we establish a benchmark for reasoning segmentation evaluation, called ReasonSeg. Comprising over one thousand image-instruction pairs, this benchmark offers persuasive evaluation metrics for the task. To align more closely with practical applications, we annotate the images from OpenImages (Kuznetsova et al., 2020) and ScanNetv2 (Dai et al., 2017) with implicit text queries that necessitate complex reasoning.

In summary, our contributions are as follows:

We introduce the reasoning segmentation task, which necessitates reasoning based on implicit human instructions. This task emphasizes the importance of self-reasoning ability, crucial for building a genuinely intelligent perception system.

We establish a reasoning segmentation benchmark, ReasonSeg, containing over one thousand image-instruction pairs. This benchmark is essential for evaluation and encourages the community to develop new techniques.

We present our model — LISA, which employs the embedding-as-mask paradigm to incorporate new segmentation capabilities. LISA demonstrates robust zero-shot ability on the reasoning segmentation task when trained on reasoning-free datasets and achieves further performance boost by fine-tuning on just 239 image-instruction pairs involving reasoning. We believe LISA will promote the development of perceptual intelligence and inspire new advancements in this direction.

Related Work

Semantic segmentation aims to assign a class label to every pixel in an image. Numerous studies (Shelhamer et al., 2017; Noh et al., 2015; Badrinarayanan et al., 2017; Ronneberger et al., 2015; Chen et al., 2018; Yu & Koltun, 2016; Liu et al., 2015; Zhao et al., 2017; 2018a; Yang et al., 2018; Fu et al., 2019; Huang et al., 2019; Zhao et al., 2018b; Zhu et al., 2019; Cheng et al., 2021; Lai et al., 2021; Tian et al., 2022; 2023) have proposed diverse designs (such as encoder-decoder, dilated convolution, pyramid pooling module, non-local operator, and more) to effectively encode semantic information. Research on instance segmentation (He et al., 2017; Zhang et al., 2021; Cheng et al., 2022) and panoptic segmentation (Kirillov et al., 2019; Xiong et al., 2019; Cheng et al., 2020; Li et al., 2021) has introduced various architectural innovations for instance-level segmentation, including DETR (Carion et al., 2020)-based structures, mask attention, and dynamic convolution. In recent years, typical segmentation tasks have made significant progress and become increasingly mature. Consequently, it is imperative to develop more intelligent interaction ways for image segmentation.

The referring segmentation task (Kazemzadeh et al., 2014; Nagaraja et al., 2016) enables interaction with human language, aiming to segment the target object based on a given explicit text description. Recently, Kirillov et al. (2023) introduced SAM, trained with billions of high-quality masks, supporting bounding boxes and points as prompts while demonstrating exceptional segmentation quality. X-Decoder (Zou et al., 2023a) bridges vision and language, unifying multiple tasks within a single model. SEEM (Zou et al., 2023b) further supports various human interaction methods, including text, audio, and scribble. However, these studies primarily focus on addressing multi-task compatibility and unification, neglecting the injection of new capabilities. In this work, we present LISA to tackle the reasoning segmentation task and enhance existing visual segmentors with self-reasoning abilities.

2 Multi-modal Large Language Model

Motivated by the remarkable reasoning abilities of LLMs, researchers are exploring ways to transfer these capabilities into the vision domain, developing multi-modal LLMs. Flamingo (Alayrac et al., 2022) employs a cross-attention structure to attend to visual contexts, enabling visual in-context learning. Models such as BLIP-2 (Li et al., 2023b) and mPLUG-OWL (Ye et al., 2023) propose encoding image features with a visual encoder, which are then fed into the LLM alongside text embeddings. Otter (Li et al., 2023a) further incorporates robust few-shot capabilities through in-context instruction tuning on the proposed MIMIC-IT dataset. LLaVA (Liu et al., 2023b) and MiniGPT-4 (Zhu et al., 2023) first conduct image-text feature alignment followed by instruction tuning. Koh et al. (2023) also investigates image retrieval for LLMs. Moreover, numerous works (Wu et al., 2023; Yang et al., 2023b; Shen et al., 2023; Liu et al., 2023c; Yang et al., 2023a) utilize prompt engineering, connecting independent modules via API calls, but without the benefits of end-to-end training. Recently, there have been studies examining the intersection between multi-modal LLMs and vision tasks. VisionLLM (Wang et al., 2023) offers a flexible interaction interface for multiple vision-centric tasks through instruction tuning but fails to fully exploit LLMs for complex reasoning. Kosmos-2 (Peng et al., 2023) constructs large-scale data of grounded image-text pairs, infusing grounding capabilities into LLMs. DetGPT (Pi et al., 2023) bridges the fixed multi-modal LLM and open-vocabulary detector, enabling detection to be performed based on users’ instructions. GPT4RoI (Zhang et al., 2023) introduces spatial boxes as input and trains the model on region-text pairs. In contrast, our work aims to 1) efficiently inject segmentation capabilities into multi-modal LLMs and 2) unlock self-reasoning abilities for current perception systems.

Reasoning Segmentation

The reasoning segmentation task is to output a binary segmentation mask M\bm{M}, given an input image ximg\bm{x}_{img} and an implicit query text instruction xtxt\bm{x}_{txt}. The task shares a similar formulation with the referring segmentation task (Kazemzadeh et al., 2014), but is far more challenging. The key distinction lies in the complexity of the query text in reasoning segmentation. Instead of a straightforward phrase (e.g., ”the trash can”), the query text may include more intricate expressions (e.g., ”something that the garbage should be put into”) or longer sentences (e.g., ”After cooking, consuming food, and preparing for food, where can we throw away the rest of the food and scraps?”) that involve complex reasoning or world knowledge.

2 Benchmark

Given the lack of quantitative evaluation, it is imperative to establish a benchmark for the reasoning segmentation task. To ensure reliable assessment, we have collected a diverse set of images from OpenImages (Kuznetsova et al., 2020) and ScanNetv2 (Dai et al., 2017), annotating them with implicit text instructions and high-quality target masks. To cover different scenarios, our text instructions consist of two types: 1) short phrases; 2) long sentences; as illustrated in Figure 2. The resulting ReasonSeg benchmark comprises a total of 1218 image-instruction pairs. This dataset is further partitioned into three splits: train, val, and test, containing 239, 200, and 779 image-instruction pairs, respectively. As the primary purpose of the benchmark is evaluation, the validation and testing sets include a larger number of image-instruction samples.

Our method

In this section, we first introduce the model architecture in Sec. 4.1. After that, we elaborate on the training data preparation and training parameters in Sec. 4.2.

Most current multi-modal LLMs (such as LLaVA, Flamingo (Alayrac et al., 2022), BLIP-2 (Li et al., 2023b), Otter (Li et al., 2023a), etc.) support image and text as input and text as output, but they cannot directly output fine-grained segmentation masks. VisionLLM (Wang et al., 2023) offers a solution by parsing segmentation masks as sequences of polygons, enabling the representation of segmentation masks as plain text and allowing end-to-end training within the framework of existing multi-modal LLMs. However, end-to-end training with the polygon sequences introduces optimization challenges and may compromise generalization ability unless a massive amount of data and computational resources are employed. For instance, training a 7B model, VisionLLM requires 4×84\times 8 NVIDIA 80G A100 GPUs and 50 epochs, which is computationally prohibitive. In contrast, training LISA-7B requires only 10,000 training steps on 8 NVIDIA 24G 3090 GPUs.

To this end, we propose the embedding-as-mask paradigm to infuse new segmentation capabilities into the multi-modal LLM. The pipeline of our method is illustrated in Fig. 3. Specifically, we first expand the original LLM vocabulary with a new token, i.e., , which signifies the request for the segmentation output. Given a text instruction xtxt\bm{x}_{txt} along with the input image ximg\bm{x}_{img}, we feed them into the multi-modal LLM F\mathcal{F}, which in turn outputs a text response y^txt\hat{\bm{y}}_{txt}. It can be formulated as

When the LLM intends to generate a binary segmentation mask, the output y^txt\hat{\bm{y}}_{txt} should include a token. We then extract the last-layer embedding h^seg\hat{\bm{h}}_{seg} corresponding to the token and apply an MLP projection layer γ\gamma to obtain hseg\bm{h}_{seg}. Simultaneously, the vision backbone Fenc\mathcal{F}_{enc} extracts the visual embeddings f\bm{f} from the visual input ximg\bm{x}_{img}. Finally, hseg\bm{h}_{seg} and f\bm{f} are fed to the decoder Fdec\mathcal{F}_{dec} to produce the final segmentation mask M^\hat{\bm{M}}. The detailed structure of the decoder Fdec\mathcal{F}_{dec} follows Kirillov et al. (2023). The process can be formulated as

Training Objectives.

The model is trained end-to-end using the text generation loss Ltxt\mathcal{L}_{txt} and the segmentation mask loss Lmask\mathcal{L}_{mask}. The overall objective L\mathcal{L} is the weighted sum of these losses, determined by λtxt\lambda_{txt} and λmask\lambda_{mask}:

Specifically, Ltxt\mathcal{L}_{txt} is the auto-regressive cross-entropy loss for text generation, and Lmask\mathcal{L}_{mask} is the mask loss, which encourages the model to produce high-quality segmentation results. To compute Lmask\mathcal{L}_{mask}, we employ a combination of per-pixel binary cross-entropy (BCE) loss and DICE loss, with corresponding loss weights λbce\lambda_{bce} and λdice\lambda_{dice}. Given the ground-truth targets ytxt\bm{y}_{txt} and M\bm{M}, these losses can be formulated as:

2 Training

As illustrated in Fig. 4, our training data consists of three parts, all of which are derived from widely-used public datasets. The details are as follows:

Semantic Segmentation Dataset. Semantic segmentation datasets typically consist of images and the corresponding multi-class labels. During training, we randomly choose several categories for each image. To generate data that matches the format of visual question answering, we employ a question-answer template like “USER: Can you segment the {CLASS_NAME} in this image? ASSISTANT: It is .”, where {CLASS_NAME} is the chosen category, and denotes the placeholder for tokens of image patches. The corresponding binary segmentation mask is used as the ground truth to provide mask loss supervision. During training, we also use other templates to generate the QA data to ensure data diversity. We adopt ADE20K, COCO-Stuff, and LVIS-PACO part segmentation datasets.

Vanilla Referring Segmentation Dataset. Referring segmentation datasets provide an input image and an explicit short description of the target object. Thus, it is easy to convert them into question-answer pairs using a template like “USER: Can you segment {description} in this image? ASSISTANT: Sure, it is .”, where {description} is the given explicit description. For this part, we adopt refCOCO, refCOCO+, refCOCOg, and refCLEF datasets.

Visual Question Answering Dataset. To preserve the original Visual Question Answering (VQA) ability of the multi-modal LLM, we also include the VQA dataset during training. We directly use the LLaVA-Instruct-150k data (Liu et al., 2023b) generated by GPT-4.

Notably, the training set does not include any reasoning segmentation sample. Instead, it only contains samples where the target objects are explicitly indicated in the query texts. Surprisingly, even without complex reasoning training data, LISA demonstrates impressive zero-shot ability on the ReasonSeg benchmark. Moreover, we find that further performance boost could be yielded by finetuning the model on only 239 image-instruction reasoning segmentation pairs.

Trainable Parameters.

To preserve the generalization ability of the pre-trained multi-modal LLM F\mathcal{F} (i.e., LLaVA in our experiments), we leverage LoRA (Hu et al., 2021) to perform efficient fine-tuning, and completely freeze the vision backbone Fenc\mathcal{F}_{enc}. The decoder Fdec\mathcal{F}_{dec} is fully fine-tuned. Additionally, the word embeddings of the LLM and the projection layer of γ\gamma are also trainable.

Experiment

Unless otherwise specified, we use LLaVA-7B-v1-1 or LLaVA-13B-v1-1 as the multi-modal LLM F\mathcal{F}, and adopt the ViT-H SAM backbone as the vision backbone Fenc\mathcal{F}_{enc}. The projection layer of γ\gamma is an MLP with channels of .

Implementation Details.

We adopt 8 NVIDIA 24G 3090 GPUs for training. The training scripts are based on deepspeed (Rasley et al., 2020) engine. We use AdamW (Loshchilov & Hutter, 2017) optimizer with the learning rate and weight decay set to 0.0003 and 0, respectively. We also adopt WarmupDecayLR as the learning rate scheduler, where the warmup iterations are set to 100. The weights of the text generation loss λtxt_gen\lambda_{txt\_gen} and the mask loss λmask\lambda_{mask} are set to 1.01.0 and 1.01.0, respectively, and those of the bce loss λbce\lambda_{bce} and the dice loss λdice\lambda_{dice} are set to 2.02.0 and 0.50.5, respectively. Besides, the batch size per device is set to 2, and the gradient accumulation step is set to 10. During training, we select at most 3 categories for each image in semantic segmentation datasets.

Datasets.

As mentioned in Sec. 4.2, our training data is composed of three types of datasets: (1) For the semantic segmentation dataset, we use ADE20K (Zhou et al., 2017) and COCO-Stuff (Caesar et al., 2018). Besides, to enhance the segmentation results for some part of an object, we also use part semantic segmentation datasets, including PACO-LVIS (Ramanathan et al., 2023), PartImageNet (He et al., 2022), and PASCAL-Part (Chen et al., 2014); (2) For the referring segmentation dataset, we use refCLEF, refCOCO, refCOCO+ (Kazemzadeh et al., 2014), and refCOCOg (Mao et al., 2016). (3) For the visual question answering (VQA) dataset, we use LLaVA-Instruct-150k dataset (Liu et al., 2023b). In order to avoid data leakage, we exclude the COCO samples whose images are present in the refCOCO(+/g) validation sets during training. Furthermore, we surprisingly find that by finetuning the model on only 239 samples of ReasonSeg image-instruction pairs, the model’s performance can be further boosted.

Evaluation Metrics.

We follow most previous works on referring segmentation (Kazemzadeh et al., 2014; Mao et al., 2016) to adopt two metrics: gIoU and cIoU. gIoU is defined by the average of all per-image Intersection-over-Unions (IoUs), while cIoU is defined by the cumulative intersection over the cumulative union. Since cIoU is highly biased toward large-area objects and it fluctuates too much, gIoU is preferred.

2 Reasoning Segmentation Results

The reasoning segmentation results are shown in Table 1. It is worth noting that existing works fail to handle the task, but our model can accomplish the task involving complex reasoning with more than 20%20\% gIoU performance boost. As mentioned before, the reasoning segmentation task is essentially different from previous referring segmentation in that it requires the model to possess reasoning ability or access world knowledge. Only by truly understanding the query, can the model do well in the task. The existing works are limited to explicit referring and have no proper way to understand an implicit query, but our model exploits multi-modal LLM to reach the goal.

Another finding is that LISA-13B outperforms the 7B counterpart substantially, especially on the long-query scenarios, which indicates that the current performance bottleneck may still lie in understanding the query text, and a stronger multi-modal LLM might lead to even better results.

3 Vanilla Referring Segmentation Results

To show that our model is also competent in the vanilla referring segmentation task, we make a comparison with existing state-of-the-art methods in Table 2. We evaluate the methods on refCOCO, refCOCO+, refCOCOg validation and testing sets. Our model achieves state-of-the-art results across various referring segmentation benchmarks.

4 Ablation Study

In this section, we conduct an extensive ablation study to reveal the contribution of each component. Unless otherwise specified, we report the metrics of gIoU and cIoU of LISA-7B on the validation set.

We emphasize that vision backbones other than SAM are also applicable in our framework. To verify this fact, we conduct ablation in Table 5.4. No matter whether we finetune the model on ReasonSeg training set, SAM performs better than Mask2Former-Swin-L. We explain that SAM is trained with billions of high-quality masks, and thus yields a higher metric than Mask2Former that is trained on merely the COCO dataset (Lin et al., 2014). We also notice that even with Mask2Former, our framework achieves a decent performance on the reasoning segmentation task, significantly outperforming previous works such as X-Decoder (Zou et al., 2023a). This reveals the fact that the design choice of vision backbone is flexible and not limited to SAM.

SAM LoRA Fintuning.

We also investigate the effectiveness of applying LoRA on SAM backbone. In Table 5.4, we note that the performance of LoRA finetuned SAM backbone is inferior to that of the frozen one. A potential reason is that fine-tuning impairs the generalization ability of the origianl SAM model.

SAM Pre-trained Weight.

To demonstrate the contribution of SAM pre-trained weight, we make a comparison between Experiments 1 and 3 in Table 5.4. Without being initialized by SAM pre-trained weight, the vision backbone is trained from scratch. This causes the performance falling behind that of the baseline model substantially.

MLP vs. Linear Projection Layer.

In experiments 2 and 3 of Table 5.4, we notice that making γ\gamma an MLP yields little performance decrease in gIoU, but a relatively higher performance in cIoU.

Contribution of All Types of Training Data.

In Table 5, we show the contribution of each type of data to the performance. It is worth noting that in Exp. 4, we do not use any semantic segmentation dataset, and the performance drops a lot. We conjecture that semantic segmentation datasets provides a large amount of ground-truth binary masks for training, since a multi-class label can induce multiple binary masks. This shows that semantic segmentation datasets are crucial in training.

Instruction Rephrasing by GPT-3.5.

During finetuning on the reasoning segmentation image-instruction pairs, we rephrase the text instruction by GPT-3.5, and randomly choose one. The comparison between Experiments 3 and 4 in Table 5.4 shows that the performance is increased by 2.2% gIoU and 2.9% cIoU. This result verifies the effectiveness of such data augmentation.

5 Qualitative Results

As depicted in Fig. 5, we provide a visual comparison with existing related works, including the model for open-vocabulary semantic segmentation (OVSeg), referring segmentation (GRES), and the generalist models for segmentation (X-Decoder and SEEM). These models fail to handle the displayed cases with various errors, while our approach produces accurate and high-quality segmentation results.

Conclusion

In this work, we have proposed a new segmentation task—reasoning segmentation. This task is significantly more challenging than the vanilla referring segmentation task, as it requires the model to actively reason based on implicit user instructions. To enable effective evaluation, we have introduced a benchmark for this task, namely ReasonSeg. We hope this benchmark will be beneficial for the development of related technologies. Finally, we have presented our model — LISA. By employing the embedding-as-mask paradigm, it injects new segmentation capabilities into current multi-modal LLMs and performs surprisingly well on the reasoning segmentation task, even when trained on reasoning-free datasets. Consequently, it demonstrates the ability to chat with segmentation mask outputs in various scenarios. We believe our work will shed new light on the direction of combining LLMs and vision-centric tasks.

References

Appendix A Appendix