Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
Chuofan Ma, Yi Jiang, Jiannan Wu, Zehuan Yuan, Xiaojuan Qi
Introduction
Multimodal Large Language Models (MLLMs) have spread the sparks of artificial general intelligence from language to the visual domain . Owing to the foundational capabilities of Large Language Models (LLMs) , MLLMs excel in vision-language tasks that require advanced understanding and reasoning, such as image captioning and visual question answering. However, despite these achievements, current MLLMs typically fall short of localization capabilities, thus cannot ground understanding to the visual context. Such limitations constrains the model from fulfilling its potential in real-world applications like robotics, autonomous driving, and augmented reality.
In light of the gap, one stream of research attempts to augment the LLM to directly output quantized object coordinates for localization (Fig. 2(a)). While this method is simple in design, the substantial computational demands of LLMs make it challenging to process high-resolution image inputs, which are essential for accurate localization. Besides, the nature of sequence outputs in LLMs is not well-suited for dense prediction tasks such as segmentation. These concerns elicit another stream of research, which incorporates an external localization module (e.g., SAM ) to decode bounding boxes or masks (Fig. 2(b)). This approach circumvents aforementioned issues, but introduces additional latency in inference as it requires processing the image input twice with the MLLM and the localization module, respectively.
The above motivates us to explore a new paradigm for grounded MLLMs. Drawing inspiration from open-vocabulary object detection , we decompose the grounding task into two sub-problems: discovering the object (localization) and relating the object to texts (recognition). We notice that localization alone requires little semantic understanding but demands perceptual skills, which is typically out of the scope of an LLM’s expertise. This inspires us to decouple localization and recognition within MLLMs. But instead of using external modules, we propose exploiting the spatial understanding capability in the visual tokenizer of MLLMs for localization (Fig. 2(c)). This perceive-then-understand design also resembles human vision process.
Building upon this concept, we introduce GromaIn Latin, Groma refers to an instrument used for accurate measurement, which implies our focus on accurate localization for MLLMs. (Grounded Multimodal Assistant), an MLLM with localized and fine-grained visual perception abilities. Specifically, Groma incorporates region tokenization alongside standard image tokenization to identify and encode potential regions of interest (ROIs) into region tokens. During this process, location information is extracted from the image and associated with region tokens, with each region token anchored to the underlying ROI. This allows Groma to ground its textual output by simply referring to region tokens, alleviating the need for the LLM to meticulously regress object coordinates. Moreover, the tokenizer of Groma can also encode user-specified region inputs (i.e., bounding boxes) into region tokens, which are directly inserted into user instructions to initiate referential dialogue.
Compared to previous methods that augment LLMs for localization , Groma circumvents the heavy computation of LLMs when handling high-resolution input by settling localization to the image tokenization process. That is, Groma can use high-resolution images for tokenizer input and downsampled image tokens for LLM input, which saves computation without sacrificing localization accuracy. Besides, unlike methods adopting separate designs for modeling grounding outputs and referring inputs , Groma seamlessly unifies the two capabilities with the use of region tokens.
From the data perspective, to improve the localized understanding of Groma, we adopt an extensive collection of datasets with region-level annotations for training, which encompasses a range of region semantics from objects and relationships to detailed region descriptions. In addition, to remedy the lack of long-form grounded data, we construct a visually grounded chat dataset called Groma Instruct for instruction finetuning. Groma Instruct is the first grounded chat dataset constructed with both visual and textual prompts, leveraging the powerful GPT-4V for data generation.
Our comprehensive experiments demonstrate the superiority of the design of Groma, with results showing that it outperforms all comparable MLLMs on established referring and grounding benchmarks. We also showcase that Groma maintains strong image-level understanding and reasoning abilities on the conversational VQA benchmark. Moreover, to assess the ability to localize multiple, diverse, and variably-sized objects, we adapt the LVIS detection benchmark for object grounding evaluation. On this challenging benchmark, Groma surpasses alternative methods by a significant margin (over AR), highlighting its robust and precise localization capabilities.
Related Work
Large language models (LLMs) such as GPT series and LLaMA have recently undergone rapid development and sparked a revolution in the field of natural language processing. Such progress inspires the community to extend the foundational capabilities of LLMs to the visual domain, giving birth to multimodal large language models (MLLMs). The pioneering works of MLLMs typically follow a tripartite architecture, comprising a visual encoder, a vision-language connector, and a large language model. Specifically, BLIP-2 and Flamingo first propose the Q-Former/Resampler to bridge vision and language. LLaVA and MiniGPT4 streamline this vision-language connector to a linear layer, and introduce visual instruction tuning to enhance the instruction-following ability of MLLMs. Following works further showcase the immense potential of MLLMs by scaling up the visual components to the magnitude as LLMs. While these works have exhibited impressive visual understanding capabilities, they are predominantly constrained to image-level tasks, such as image captioning and image visual question answering. This necessitates the research into region-level MLLMs, which unlock more nuanced and granular visual-language interactions.
0.2 Region-level MLLMs.
In pursuit of fine-grained and grounded image understanding, recent studies further integrate region-level data into the training of MLLMs . In particular, to model box inputs and outputs, Kosmos-2 and Shikra directly quantize bounding boxes into discrete location tokens or numeric representation of positions. GPT4RoI and RegionGPT use a simple pooling operation to extract the features within boxes or masks as the region representations. While Ferret proposes a spatial-aware visual sampler to deal with free-form region inputs. Besides, to achieve more accurate localization, some works resort to off-the-shelf models for pixel-level grounding. For instance, LISA takes the segmentation token generated by the MLLM as the prompts for SAM to produce the segmentation masks. GLaMM and LLaVA-Ground further advance the concept and enable grounded conversation generation. Our work shares the same focus with the aforementioned methods on region-level understanding and grounding. Yet, we distinguish ourselves from existing studies by proposing a novel perspective in enhancing the localization ability of MLLMs.
Method
In this section, we present Groma, a grounded multimodal large language model capable of understanding user-defined region inputs and generating visually grounded outputs. We first illustrate the model architecture of Groma in Sec. 3.1. Then we introduce how to format region input and output in Sec. 3.2. Finally, we detail the learning pipelines Sec. 3.3.
As illustrated in Fig. 3, Groma primarily consists of (1) an image encoder for scene-level image tokenization, (2) a region proposer for discovering regions of interest, (3) a region encoder for region-level image tokenization, and (4) a large language model for modeling multimodal input and output. We detail each component in the following paragraphs.
Groma employs a pretrained DINOv2 model as the image encoder with the input image resolution set to . Compared with the commonly adopted CLIP visual encoder, DINOv2 is preferred in this work for its compatibility with high-resolution inputs and fine-grained features for localizationA performance comparison between CLIP and DINOv2 on the detection benchmark is available in our ablation study.. However, the use of higher-resolution images leads to extended sequences of visual input for the language model, e.g., 1024 tokens in this case. To save computations, we further concatenate every four neighbor patch tokens into a single token following MiniGPT-v2 . But slightly different from , we merge tokens adjacent in 2D instead of 1D, which yields better results empirically.
1.2 Region Proposer.
To obtain localized understanding of the image, Groma innovatively incorporates a region proposer into the image tokenization process. Specifically, the region proposer is implemented as a class-agnostic detector head using the Deformable DETR (DDETR) transformer . The original classification head of DDETR is replaced by a binary classifier to score region proposals based on their localization quality. Inspired by ViTDet , we extract feature maps from the last 4 layers of the image encoder, and rescale these feature maps to construct a hierarchical feature pyramid as the input to the region proposer. For each image, the region proposer generates 300 region proposals, which are then filtered by NMS and objectness scores before fed into the region encoder.
1.3 Region Encoder.
The region encoder translates region proposals (i.e., bounding boxes), coming from both user input and the region proposer, into region tokens. Akin to the previous step, we select feature maps from the last three layers of the image encoder to create a hierarchical feature pyramid. A multi-scale ROIAlign module as implemented in is utilized to crop and fuse these hierarchical features into unified region tokens. Compared with alternative ways to represent regional inputs, such as numerical representation of positions and discrete location tokens , the region token representation offers distinct benefits as it is semantically aligned with the underlying region, which renders it more intuitive for the language model to comprehend.
1.4 LLM.
We adopt pretrained Vicuna as the language model of Groma. In particular, we instantiate Groma with the 7B version of Vicuna. Besides, we follow LLaVA v1.5 to use an MLP layer to project the image tokens and region tokens into the feature space of the LLM.
2 Input and Output Formatting
Beyond textual only instructions and responses, Groma offers the flexibility to accept user-specified regions as input (referring) and generate visually grounded answers (grounding). Specifically, although different in task formulations, both referring and grounding are unified into one format with the use of region tokens.
Remember in the tokenization process, each region token is inherently anchored to a concrete location in the image, corresponding to its region proposal. This connection allows the language model to ground its text output to particular regions in the image by simply referring to the associated region tokens. However, as region tokens are continuous embeddings, they cannot be directly integrated into the codebook of the language model and referenced in the text output. To bridge the gap, we further introduce a set of proxy tokens “
User: Here is an image with region crops from it. Image: A dog a frisbee a fallen man and
2.2 Referring Input.
For a region pointed out by the user, we treat it the same as region proposals from the region proposer, i.e., encoding it into a region token and assigning a proxy token to it. This allows us to incorporate user-specified regions into our instructions by inserting corresponding region tokens. A simple example of referential dialogue in Groma is given below, where
3 Model Training
The training of Groma is partitioned into three stages: (i) detection pretraining for localization ability, (ii) alignment pretraining for image-level and region-level vision-language alignment, (iii) instruction finetuning for enhanced conversation capability. Tab. 1 enumerates the datasets used at different training stages. Additionally, we provide the instruction templates used to convert task-specified datasets to instruction following format in Appendix 0.A.
This training stage only involves the image encoder and the region proposer, which collectively constitute a DDETR-like detector. The image encoder is kept frozen during training. To endow the region proposer with localization capability, an extensive collection of detection datasets, including COCO , Objects365 , OpenImages , and V3Det , is utilized for large-scale pretraining. Notably, category information is omitted from the training process, with a primary focus on box supervision.
Considering traditional detection data are typically limited to object-level annotations, we complement the training with a two million subset of SA1B data filtered by GLEE . Original mask annotations of SA1B are transformed into bounding boxes for consistency. The inclusion of this enriched dataset encourages the region proposer to produce region proposals across a wide spectrum of granularities, encompassing not only object instances but also their constituent parts and various background stuff.
3.2 Alignment Pretraining.
To align vision and language feature space of Groma, we pretrain the model on a wide range of vision-language tasks. Specifically, for image-level alignment, we leverage ShareGPT-4V-PT for detailed image captioning. For region-level alignment, we engage COCO , RefCOCO , RefCOCO+ , RefCOCOg , and Grit-20m for referring expression comprehension (REC), Visual Genome for region captioning, and Flickr30k Entities for grounded caption generation. To maintain training efficiency, we focus finetuning efforts on the MLP projection layer and the region encoder, while other modules are kept frozen throughout the training.
3.3 Instruction Finetuning.
Based on alignment pretraining, we refine the training data to focus exclusively on high-quality datasets and proceed to unfreeze the language model for finetuning purposes. At this stage, LLaVA Instruct and ShareGPT-4V are incorporated to improve the conversational and instruction-following capabilities of GromaLLaVA Instruct contains three types of instruction data, namely conversation, detailed description, and complex reasoning. Since the detailed description part of LLaVA Instruct has severe hallucinations, we replace it with ShareGPT-4V as in .. Besides, we curate a high-quality grounded chat dataset, named Groma Instruct (see next section for more details), to facilitate synergy of chatting and grounding abilities of Groma.
3.4 Discussions.
A major difference between the training of Groma and current MLLMs is the integration of dedicated detection pretraining, which endows Groma with robust and precise localization ability. Thanks to the decoupled architecture of location and understanding within Groma, we circumvent the need to involve the LLM during detection pretraining. Such a strategic design allows Groma to benefit from pretraining on millions of bounding box annotations — a task that would be computationally prohibitive for classic MLLMs.
GPT4V-assisted Grounded Conversation Generation
Visual dialogue data have proven to be crucial in advancing the conversational capability of the MLLM as a visual chatbot. Previous methods mostly rely on coarse-grained image descriptions to derive free-form visual dialogues, which typically lack fine-grained region details and precise location information . For grounded MLLMs, such free-form dialogue data are shown to be insufficient to enable the model to generate long-form grounded responses - as the format of grounded responses significantly deviates from that of normal responses, it could be challenging for the grounded MLLM to generalize its grounding capability to long-form conversations.
To bridge the gap, we have meticulously curated a dataset containing 30k visually grounded conversations for instruction finetuning, named Groma Instruct. An illustrative example from Groma Instruct is showcased in Fig. 4. Specifically, we select images with dense region annotations from Visual Genome (VG), and take the following steps to construct grounded conversations with the assistance of advanced GPT-4V model:
First, we remove highly overlapped regions (bounding boxes) from VG annotations, normally leaving 3-10 regions of interest for each image. Then we adapt the visual prompting techniques from SoM to overlay a bright numeric marker at the center of each region. Using this marked image as input unleashes the grounding capabilities of GPT-4V - it can easily make references to specific image regions by addressing the corresponding numbers.
Besides visual input, we supply GPT-4V with rich region descriptions, image descriptions, and image-based Q&A pairs, coming from COCO and VG annotationsWe select VG images that also have a coco id. Thus, we can retrieve corresponding image captions from COCO Caption.. While such textual context is optional for GPT-4V input, we empirically find it useful to reduce hallucinations in generated contents and resolve potential ambiguities in visual promptsThere are cases where two regions highly overlap with each other and GPT-4V can hardly tell from the image which region maps to which numeric marker. For these cases, GPT-4V could rely on the numbered region descriptions to find out correspondences between regions and markers..
Inspired by prior studies on visual chat data construction , we further provide GPT-4V with manually designed grounded chat as context examples. This provokes the in-context-learning ability of GPT-4V to generate grounded conversations in a uniform format. We also take a post-processing stage to filter out conversations not following the pre-defined format.
Experiments
In this section, we first quantitatively access the abilities of Groma on grounding (Sec. 5.2), referring (Sec. 5.3), and image-based conversation (Sec. 5.4) tasks. Then we provide qualitative results to exemplify the strong capabilities of Groma on a wide range of region-level tasks (Sec. 5.5). Finally, we ablate the design and training of Groma in Sec. 5.6.
We adopt DINOv2-L/14 as the image encoder and Vicuna-7B v1.5 as the language model. The region proposer follows an encoder-decoder architecture with 6 encoder layers and 6 decoder layers. We further employ mixed query selection and look-forward-twice scheme as in to accelerate convergence. We set NMS threshold to 0.6 and filter out region proposals with objectness scores lower than 0.15. Subsequently, we select the top 100 region proposals if there are more than 100 proposals left after filtering. This results in no more than 356 visual tokens in total. For training, we sequentially proceed 12 epochs of detection pretraining, 2 epochs of alignment pretraining, and 1 epoch of instruction finetuning. More training details can be found in the Appendix 0.C.
2 Grounding Benchmark Results
We evaluate the localization capability of Groma on visual grounding tasks. Tab. 2 showcases our performance on three classic referring expression comprehension benchmarks: RefCOCO , RefCOCO+ , and RefCOCOg . Groma notably surpasses other generalist models of similar model size across all metrics. Even in comparison with Qwen-VL , which uses a stronger visual tokenizer and trains on more grounding data, Groma delivers superior accuracy on average. Moreover, as a generalist model, Groma shows competitive results with state-of-the-art specialist models . These findings underscore the strong capability of Groma in visual grounding.
However, we notice that traditional REC benchmarks only cover a narrow range of common objects in their referring expressions, which is insufficient to thoroughly evaluate the MLLM’s localization capability. Therefore, we further introduce LVIS-Ground, an object grounding benchmark converted from the LVIS detection data. LVIS-Ground contains 4299 images covering 1203 categories of objects, with on average 3.7 target objects per image. Complementary to REC benchmarks, LVIS-Ground focuses on testing the model’s ability to locate multiple, diverse, and variably-sized objects. For more details of LVIS-Ground, please refer to the Appendix 0.B.
Tab. 3 presents ours results on LVIS-Ground. Notably, Groma demonstrates clear advantages over other grounded MLLMs, especially on the AR@0.75 metric. This evidences that the specialized design and training indeed bring more accurate localization for Groma. Moreover, it is noteworthy that current MLLMs all fall short of small object localization (AR@s metric). We conjecture this is mainly because the training data (e.g., RefCOCO/g/+, Flickr30k) lack annotations for small objects. We also notice a common failure mode of these methods is that, most of the time they only predict one box per image. This is an expected behavior as the they heavily rely on REC data for grounding training, which only has one target object per query. These findings call for the necessity of diversifying grounding data used for training in future MLLMs.
3 Referring Benchmark Results
We evaluate Groma on the region captioning task to assess its fine-grained region understanding capability. To prompt the model to generate region-level descriptions, we use queries like “Please describe
4 Conversational VQA Benchmark Results
In addition to region-level tasks, we further evaluate Groma on the conversational style VQA benchmark, LLaVA Bench (COCO) , which contains three types of questions, namely conversation, detailed description, and complex reasoning. As shown in Tab. 5, Groma surpasses the strong baseline method LLaVA and achieves competitive performance among grounded MLLMs, especially in detailed image description. This demonstrates that Groma maintains decent image understanding and visual chatting abilities. For the underperformance in conversation and complex reasoning questions, we speculate this could be resulted from the DINOv2 features. Recent studies have shown that DINOv2 image tokenizer slightly underperforms CLIP tokenizer in image understanding tasks as DINOv2 features are not inherently aligned with text. But we believe such gap can be closed by scaling up vision-language alignment pretraining.
5 Qualitative Results
Fig. 5 presents a comparison between Groma and other grounded MLLMs on the grounded image captioning task. We choose an exemplar image that is inherently challenging with multiple and occluded instances to ground. Groma manifests exceptional grounding performance in this case with the highest recall and minimum hallucinations. In addition, we provide several visualization examples in Fig. 6 for a complementary understanding of Groma’s abilities on grounded chat and referential dialogue. We show that Groma is capable of generating long-form, grounded and logically rich answers, which can be mainly attributed to the introduction of Groma Instruct data in finetuning.
6 Ablation
To quantitatively assess the differences in localization capabilities between CLIP and DINOv2, we compare the two backbones on the COCO detection benchmark in Tab. 6. For this comparison, we equip each backbone with a DDETR detection head and finetune only the detection head on COCO dataset. It can be seen that under the same resolution, DINOv2 backbone significantly outperforms CLIP backbone by AP. Furthermore, by scaling the resolution to , DINOv2 backbone achieves a commendable performance of AP. The results consolidate our choice of DINOv2 backbone in Groma.
6.2 Frozen LLM.
In Tab. 7, we reveal that Groma retains robust localized understanding even without finetuning the LLM, i.e., it demonstrates a referring ability on par with GPT4ROI ( vs. ) and grounding ability comparable to Ferret ( vs. ). This finding suggests our design effectively decouples localization and understanding within Groma, such that it requires minimum ‘new knowledge’ from the LLM for localized understanding.
6.3 Token Merge.
To save computations, Groma by default concatenates every 4 image tokens into one as LLM inputs. Through control experiments in Tab. 8, we find that such downsampling has negligible impacts on the grounding performances (e.g., less than average accuracy drop on the REC benchmarks). The results evidence that the decoupled design is optimal in both efficiency and localization accuracy.
Limitations and Conclusions
In this paper, we introduce a novel paradigm, Groma, to unleash the localized perception capabilities of MLLMs. We make the pioneering attempt to embed localization into image tokenization. Our paradigm is based on a perception-then-understand mindset that separates localization from high-level understanding and reasoning. Without introducing external modules, our approach overcomes the resolution bottleneck of using LLMs as location decoders and unifies referring and visual grounding tasks. Extensive experiments showcase the superior performance of our approach in localized perception, as evidenced by its success in referring and visual grounding tasks.
However, the current implementation does not support free-form region inputs and pixel-level grounding. A promising direction to address such limitations is to re-implement the region encoder with a visual sampler as in and replace the box region proposer by a mask region proposer like Mask2Former . We leave this for future studies.
References
Appendix 0.A Task-Specified Instruction Templates
In complementary to discussion on training datasets in Sec. 3.3, we list a few instruction templates used to convert task-specified datasets to instruction following format in Tab. 9. Specifically, we convert the detection dataset COCO to multiple objects grounding data in a similar format as REC data.
Appendix 0.B LVIS-Ground Benchmark
Current MLLMs typically do not support detecting multiple categories of objects at the same time. Therefore, to customize the LVIS detection benchmark for MLLM evaluation, each time we only select one object class that is included in the image to ground. For instance, the grounding query can be formulated as “Locate all {object class name} in this image”. However, this ‘one-by-one’ evaluation strategy unavoidably leads to low efficiency. To save time and maintain class balance, we randomly sample at most 5 images for each object categorySome categories have fewer than 5 samples in the original LVIS validation set. from the LVIS validation set to construct LVIS-Ground.
There are often multiple ground-truth boxes for a query in LVIS-Ground. In such cases, traditional methods either adopt the ANY-Protocol or MERGED-BOXES-Protocol to evaluate performance . To be specific, the ANY-Protocol considers recall to be if the prediction matches any of the ground-truth boxes (e.g., with IoU 0.5), which fails to truly reflect the model’s capability in finding out all object instances. On the other hand, the MERGED-BOXES-Protocol merges all ground-truth boxes into a smallest enclosing box as the ultimate ground-truth box. However, this protocol ignores the atomicity of individual boxes, and is not well-suited for instance-level prediction evaluation.
To better evaluate recall for multiple ground-truths, we propose a new protocol termed AS-MANY-Protocol. This protocol selects the top-k predicted boxes (where k is the number of ground-truth boxes) and measures recall over all ground-truth boxes. For example, if there are 3 out of 5 ground-truth boxes hit by the top-5 predicted boxes, the recall is . Besides, we follow common practice in detection to calculate average recall over 10 IoU thresholds (ranging from 0.5 to 0.95) as the primary metric on LVIS-Ground.
Appendix 0.C More Implementation Details
Table 10 lists the detailed hyper-parameter configuration used for Groma training. It takes roughly days to finish stage training on 8 A100 GPUs. For some large-scale datasets, we merely sample a subset from them during training. The total number of training samples in one epoch is given in Tab. 10.