SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, Jiaming Han, Siyuan Huang, Yichi Zhang, Xuming He, Hongsheng Li, Yu Qiao
Introduction
Since the era of big data, large language models (LLMs) have attained tremendous strides (OpenAI, 2023a; b; Brown et al., 2020; Touvron et al., 2023a; Zhang et al., 2022), showcasing unprecedented application scenarios and generalization capabilities. To further expand their capacity ceiling, visual images are also introduced as inputs to develop powerful multi-modal large language models (MLLMs) (Zhang et al., 2023a; Li et al., 2023d; Liu et al., 2023d; Zhu et al., 2023; Zhao et al., 2023). These methods can not only generate well-organized language responses inherited from LLMs, but also unlock the multi-modal understanding capability for a wide range of applications, such as providing detailed image captions, answering visual questions, localizing different objects on the image, etc.
Existing MLLMs explored various strategies to endow LLMs with visual instruction-following capacities. 1) Freezing the LLMs during pre-training, and only learning a projection network for vision-language alignment, e.g., a simple MLP layer of LLaMA-Adapter V2 (Gao et al., 2023b) and an attention-based visual abstractor of mPLUG-Owl (Ye et al., 2023). 2) Constructing training data of new tasks to endow MLLMs with new visual understanding abilities, e.g., referential dialogues of Kosmos-2 (Peng et al., 2023b) and region-level grounding of Shikra (Chen et al., 2023b). 3) Employing advanced image encoders for extracting visual embeddings, e.g., the CLIP encoder (Radford et al., 2021) in LLaVA (Liu et al., 2023c) and the Q-Former (Li et al., 2023d) in MiniGPT-4 (Zhu et al., 2023).
In this paper, we propose a versatile MLLM, SPHINX, with a mixing of four significant aspects: model weights, tuning tasks, visual embeddings, and high-resolution sub-images. The main characteristics and findings of our approach is illustrated as follows:
Unfreezing LLMs for pre-training. Although the frozen LLM can effectively preserve its long-sentence generation capability, it constrains the potential of better cross-modal alignment via further pre-training on vision-language data. Therefore, we unfreeze the entire LLM, and combine the vision-language datasets (Schuhmann et al., 2021) for cross-modal alignment and RefinedWeb (Penedo et al., 2023) for language-specific tuning. This pre-training strategy not only enables LLMs to learn more cross-modal knowledge, but also alleviates the forgetting issue to generate detailed language responses.
Mixed model weights. Vision-language data from particular domains might contain special semantics, e.g., synthetic captions (Schuhmann et al., 2022) compared to real-world ones (Schuhmann et al., 2021). Considering that directly mixing such data might confuse the MLLM, we introduce a weight-mixing strategy to efficiently combine such domain-specific knowledge. Based on the MLLM pre-trained on real-world data, we fine-tune it on the synthetic data, and then linearly combine the finetuned LLM’s weights with the real-world ones. In this way, the two types of models would not be affected by contradictory data and our final SPHINX can effectively integrate knowledge from both synthetic and real-world domains.
Mixed tuning tasks. Different from existing task-specific MLLM models (Ye et al., 2023; Peng et al., 2023b; Chen et al., 2023b; Liu et al., 2023d; Gao et al., 2023b), we integrate a diverse set of visual instruction tasks to tune the pre-trained model, aiming to acquire a wide range of capabilities. Our mixing of tasks includes basic visual question answering (VQA), region-level referring expression comprehension/generation (REC/REG), multi-object detection and relation reasoning, text-oriented chart/document VQA, human pose estimation, etc. By such a comprehensive multi-task training paradigm, our SPHINX is a well-performing generalist model for visual instruction following.
Mixed visual embeddings. To take the advantage of different encoders, we propose to mix the visual embeddings from various vision backbones (Oquab et al., 2023; Li et al., 2023d; Radford et al., 2021) with different network architectures (CNN vs. ViT), pre-training paradigms (supervised vs. self-supervised), and information granularity (global vs. local). By mixing the different image tokens channel-wisely and sequence-wisely, SPHINX obtains stronger visual representations and leads to better vision-language alignment efficacy.
On top of this, we further investigate another challenging issue within existing MLLMs, i.e., the limited resolution of input images. As the pre-trained image encoders normally adopt a relatively low image resolution, e.g., 224224, it severely hinders fine-grained visual comprehension and reasoning for MLLMs. However, simply upsampling the images for encoders would harm the pre-trained positional prior, and, more importantly, lead to expensive computational overhead (the complexity increases quadratically to image size in self-attention mechanisms). Therefore, we propose to endow SPHINX with a longer sequence of visual embeddings of mixing different scales and high-resolution sub-images.
Mixed scales and high-resolution sub-images. we first spatially divide the input high-resolution image into multiple sub-images, and also downsample it into a low-resolution one. Then, we feed all the images concurrently into the mixed visual encoders, and concatenate the extracted multiple token groups to represent the entire high-resolution visual features. By mixing visual embeddings of different scales and sub-images, our SPHINX can adaptively explore more fine-grained visual semantics from the high resolution and multi-scale image representations, while maintaining encoding efficiency.
Note that, as the different sub-images of high-resolution images do not interact with each other in the visual encoder, they are forced to interchange information within the attention layers of LLMs, which motivates LLMs to process visual conditions more thoroughly and deeply. By the proposed three-fold mixer along with a longer visual token sequence, SPHINX fine-tunes LLMs, e.g., LLaMA-2 (Touvron et al., 2023b), to be a powerful MLLM with superior visual instruction-following capacity. As shown by the examples in Figure 1, our model excels in a variety of vision tasks, e.g., detecting different objects with remarkable precision and parsing their relations, or accurately interpreting the content within complicated figures. Importantly, as shown in Figure 2, SPHINX can achieve impressive fine-grained visual perception for high-resolution images, which exhibits state-of-the-art performance on extensive evaluation benchmarks, e.g., MMBench (Liu et al., 2023f), MME (Fu et al., 2023a), and POPE (Li et al., 2023e).
Related Work
The field of Natural Language Processing (NLP) has witnessed significant progress over the years, particularly with the advent of LLMs. With Transformer (Vaswani et al., 2017) as the fundamental architecture, LLMs (OpenAI, 2023a; Radford et al., 2019; OpenAI, 2023b) have demonstrated unprecedented performance in modeling intricate language patterns over extensive contexts. Therein, BERT (Devlin et al., 2018) showcases the benefits of pre-training on vast text corpora and fine-tuning on specific tasks, setting new standards on various benchmarks. OpenAI’s GPT series (Radford & Narasimhan, 2018; Radford et al., 2019; OpenAI, 2023a; b), especially GPT-3 (Brown et al., 2020), harness the power of massive model scaling, with billions and even trillions of parameters. To obtain better instruction following ability, InstructGPT (Ouyang et al., 2022) and ChatGPT (OpenAI, 2023a) are presented to exhibit exceptional fluency and versatility in open-domain conversation tasks, ranging from text generation to question answering. Recently, the instruction tuning based on LLaMA (Touvron et al., 2023a) and LLaMA-2 (Touvron et al., 2023b) has gained great popularity as open-source LLMs in the community. Therein, Alpaca (Taori et al., 2023) and LLaMA-Adapter (Zhang et al., 2023a) respectively adopt full and parameter-efficient fine-tuning to acquire favorable instruction-following LLMs. Vicuna (Chiang et al., 2023) and GPT-4-LLM (Peng et al., 2023a) further showcase the improvement brought by higher-quality instruction datasets. Other efforts also extend LLMs for match problem solving (Wang et al., 2023a; Zhou et al., 2023), visual model system (Wu et al., 2023; Yang et al., 2023), and open-world recognition (Zhang et al., 2023b; Zhu et al., 2022). In this paper, we develop our SPHINX based on the superior language understanding of LLaMA-2 (Touvron et al., 2023b) and instruction tuning experience of LLaMA-Adapter series (Zhang et al., 2023a; Gao et al., 2023b), which introduce a three-fold mixer to extend the capability ceiling of instruction-following LLMs for multi-modal input.
In addition to language instruction following, many efforts have been made to inject multi-modal conditions into LLMs for wider application scenarios. As prior attempts, VisualGPT (Chen et al., 2022) and BLIP series (Li et al., 2023d; 2022; Dai et al., 2023) indicate the potential of aligning LLMs with visual input for image captioning and question answering. Flamingo (Alayrac et al., 2022) and Kosmos-1 (Huang et al., 2023) further exhibit promising multi-modal understanding performance for image-text interleaved contexts. With large-scale pre-training and model sizes, GPT-4 (OpenAI, 2023b) and Bard (Google, 2023) both showcase remarkable proficiency in vision-language understanding and reasoning over diverse multi-modal tasks. In parallel, a bunch of works have been proposed to align LLaMA with vision modality for advanced visual instruction-following capabilities. LLaVA (Liu et al., 2023d) and MiniGPT-4 (Zhu et al., 2023) utilize a simple projection layer to connect vision encoders (Li et al., 2023d; Radford et al., 2021) with LLMs. LLaMA-Adapter V2 (Gao et al., 2023a) introduces zero-initialized attention mechanisms for efficient visual instruction tuning, and mPLUG-Owl (Ye et al., 2023) adopts delicately designed intermediate networks for cross-modal alignment. For more modality input, ImageBind-LLM (Han et al., 2023) and PandaGPT (Su et al., 2023) further incorporate audio and video conditions guided by ImageBind (Girdhar et al., 2023). Besides, recent MLLMs are also extended to region-level parsing (Chen et al., 2023b; Peng et al., 2023b), in-context learning (Li et al., 2023a; b), arbitrary image resolutions (Bavishi et al., 2023), text-to-image generation (Wen et al., 2023; Dong et al., 2023), and 3D question answering (Xu et al., 2023; Guo et al., 2023; Hong et al., 2023). Different from previous works, our SPHINX aims for image-conditioned MLLM, and proposes a three-fold mixer, i.e., model weights, tuning tasks, and visual embeddings, attaining superior generalization capacity for multi-modal learning.
SPHINX
In this section, we introduce a versatile MLLM, SPHINX, with the joint mixing of model weights, tuning tasks, visual embeddings, and high-resolution sub-image tokens in Section 3.1 and Section 3.2. Finally, in Section 3.3, we introduce several extended applications of SPHINX.
The overall mixing paradigm of SPHINX is shown in Figure 3. We adopt a two-stage training paradigm: the first pre-training stage for vision-language alignment, and the second fine-tuning stage for visual instruction-following learning. During the two stages, we apply the proposed mixing of model weights and tuning tasks, respectively. The model is composed of an LLM, e.g., LLaMA-2 (Touvron et al., 2023b), a mixing of vision encoders, and two linear projection layers.
Existing MLLMs (Zhu et al., 2023; Li et al., 2023d; Dai et al., 2023) generally freeze the entire LLM during the pre-training by image-caption data, and only train intermediate projection layers for vision-language alignment. This strategy can prevent LLMs from over-fitting to generating only short sentences, since the pre-training caption data mostly contain concise descriptions of images. However, the frozen weights largely constrain the cross-modal learning potential of LLMs with large-scale vision-language data. Therefore, we propose to unfreeze the entire LLM along with learnable linear projection layers, for more sufficient vision-language adaption. On the other hand, the vision encoders are kept frozen for high-quality image representations. To particularly preserve the long-sentence generation ability of LLM, we supplement the existing pre-training vision-language data with additional text corpora data Penedo et al. (2023) for language-only tuning. More specifically, in every iteration, we sample one text and several image-caption data respectively from language and vision-language datasets.
Some vision-language data from particular domains contain distinct semantic knowledge, such as the synthetic captions of LAION-COCO (Schuhmann et al., 2022) compared to real-world descriptions of LAION-400M (Schuhmann et al., 2021). We propose a weight mixing strategy of domain-specifically tuned LLMs to integrate respective knowledge from real-world and synthetic data. We first utilize the most common domain data (LAION-400M (Schuhmann et al., 2021)) for pre-training, which endows the MLLM with fundamental visual understanding capabilities. Then, we regard such a pre-trained model as the initial checkpoint to further fine-tune the LLM on synthetic domains, e.g., LAION-COCO (Schuhmann et al., 2022). Finally, to take advantage of the best data domains, we directly conduct a weighted mixing of two LLMs’ weights for semantic aggregation. In detail, we denote the parameters of the fundamental LLM as , and the fine-tuned parameters by synthetic data as . The mixing process is formulated as
where denotes the mixing coefficient, and represents the mixed LLM weights with aggregated semantics. Compared to fusing different domain data for joint pre-training, our weight mix strategy can encourage every MLLM to better learn domain-unique knowledge, and exhibit flexible scalability for any new data domains.
After pre-training and model weight mixing, the MLLM has achieved satisfactory alignment between vision and language data. To further enhance the instruction-following capacity, we collect instruction data from a wide range of multi-modal tasks, and jointly fine-tune the model to learn a vision generalist, instead of a specialist for specific scenarios. Previous open-source MLLMs can only perform simple visual question answering (VQA) and single large object referring. In contrast, we enable SPHINX to be jointly fine-tuned with a wide range of tasks, and design a set of task-specific instructions to avoid inter-task conflict. The mixed tasks include general VQA, region-level referring expression comprehension/generation (REC/REG), multi-object detection and relation reasoning, text-oriented chart/document VQA, and human pose estimation. For example, we adopt “Detect all objects shown in the image” for general object detection, and “Detect all texts and provide their bounding box coordinates” for document layout detection. Please refer to Table 1 for detailed instructions on different benchmarks. Thanks to the superior reasoning capacity of LLM and proper designs of task prompts, SPHINX, for the first time, showcases multi-purpose capabilities of visual understanding and perception, excelling in various application scenarios.
To capture robust visual representations from different aspects, we propose to ensemble a variety of vision backbones for image encoding. The visual backbones with different characteristics are chosen as follows. 1) Different network architectures. As CNN (He et al., 2016a) and ViT (Dosovitskiy et al., 2020) mainly aggregate different types of visual appearances, i.e., neighboring dependencies and long-range interactions, we adopt CLIP (Radford et al., 2021) models respectively with ConvNeXt (Woo et al., 2023) and ViT image encoders. 2) Different pre-training paradigms. Supervised training can impose explicit semantic information from textual captions or category labels, while self-supervised learning enforces the model to explore implicit pretext task signals. Thus, we further employ the ViT self-supervised by DINOv2 (Oquab et al., 2023) as well as the text-supervised vision encoders, CLIP. 3) Different information granularity. The aforementioned visual encoders all produce visual tokens in the patch level. To better capture global features, we also adopt Q-Former (Li et al., 2023d) to summarize visual embeddings via querying from the global context. After all the aforementioned encoding, we first channel-wisely concatenate the patch level visual tokens. Then, by using two projection layers for dimension alignment, we spatial-wisely concatenate the representations between those of Q-Former and the other patch-level features. The obtained image tokens are directly placed in front of language instructions, which provide visual context for the language instructions.
2 The Mixing of Scales and High-Resolution Sub-images
With the above-mentioned joint mixing strategy, SPHINX already showcases superior performance for diverse visual perception and reasoning tasks. However, one key challenge remains, i.e., the limited resolution of the input images. To tackle the problem, we further propose to utilize the mixed visual tokens of high-resolution sub-images, as shown in Figure 4.
State-of-the-art open-source MLLMs (Li et al., 2023d; Liu et al., 2023d; Gao et al., 2023b; Chen et al., 2023b; Peng et al., 2023b; Chen et al., 2023a) works adopt frozen image encoders during all training stages, in order to preserve the pre-trained visual semantics. Therefore, the image resolution of MLLMs is usually set as 224224, severely hindering their efficacy for fine-grained visual perception, especially region-level grounding and description. However, directly processing the upsampled image is not optimal for two reasons. First, to align the image size, the pre-trained positional encoding vectors in ViT are also required to be upsampled correspondingly, which would harm the prior spatial cues. Second, the computation complexity of ViT increases quadratically to the input image size. Thus, naively upsampling the image leads to extensive inference time and GPU memory consumption.
In our SPHINX, we extend the mixing of visual embeddings to more scales and high-resolution sub-images, allowing for efficient high-resolution image encoding. For an input high-resolution image, e.g., 448448, we construct five corresponding images of 224224, and feed them as independent images into our mixed vision encoders. Specifically, we first downsample the input image to 224224 as an abstract representation, and also downsample the input image to 448448 and crop four sub-images of 224224 from the four corners of the 448448 image, which preserve the detailed visual information. In this way, we enable MLLMs to not only capture fine-grained visual appearances with 224224 positional encodings, but also achieve favorable computation efficiency. Afterwards, the five groups of image tokens are encoded and concatenated as a long sequence for feeding into LLM, where the first one group encodes global semantics, and the other four record fine-grained local features. Importantly, as the image tokens of different patches do not have interaction through the vision encoders, they are forced to interact within the LLM to obtain complete visual information. Such a strategy, in turn, motivates LLMs to parse the relations within visual conditions for better cross-modal learning. From this perspective, our SPHINX can be regarded as a new paradigm for similar to ViT (Dosovitskiy et al., 2020), where the mixed vision encoders serve as a patch embedding layer, and the LLM plays the role for patch interaction as a vision decoder. On visual understanding tasks requiring higher resolutions, SPHINX achieves significant improvement with the mixed visual representations of scales and high-resolution sub-images.
3 Extensions to Wider Applications
In this section, we respectively introduce some extended applications derived from SPHINX.
In addition to multi-purpose visual instruction-following, we can also integrate SPHINX with other visual foundation models to tackle more challenging vision tasks. Figure 5 and 6 respectively show two applications for language-referred segmentation and image editing.
Given that our MLLM is able to output accurate detection boxes with user-provided descriptions or semantic categories, we can cascade the Segment Anything Model (SAM) (Kirillov et al., 2023) for language-referred instance or semantic segmentation. In detail, we regard the predicted bounding boxes from SPHINX as box prompts, and feed them into SAM for segmenting corresponding instances. In this way, we effectively incorporate the semantic reasoning capability of LLMs and the class-agnostic segmentation of SAM.
Based on the segmentation results from SAM, we refer to Inpaint Anything (Yu et al., 2023a) to integrate image inpainting models (LaMa (Suvorov et al., 2021)) and text-to-image generative models (Stable Diffusion (Rombach et al., 2021)) for high-quality image inpainting and editing. Specifically, we first detect and segment the user-indicated objects via SPHINX and SAM as illustrated in the previous paragraph. Then, we feed the segmentation mask into LaMa (Suvorov et al., 2021) for removing the corresponding objects with contextual data. After this, the user can prompt Stable Diffusion (Rombach et al., 2021) to further generate new visual content to replace the original ones. This setting integrates our SPHINX, SAM, LaMa, and Stable Diffusion to achieve language-driven image inpainting and editing.
3.2 Fine-tuning SPHINX for Visual Recognition
Empowered by the joint mixing of weights, tasks and visual embeddings, our SPHINX can comprehend robust and diverse visual category semantics. We propose to regard SPHINX as a universal initialization for traditional visual recognition tasks. For instance, given a classification task of ImageNet-1K (Russakovsky et al., 2015), we transform the task into a single-turn conversation format of “Classify the image.” as the instruction and use “This is a [CLASS]” as the response. By performing supervised fine-tuning on the text-converted dataset, we observe fast training convergence on ImageNet-1K. Surprisingly, with only one epoch, SPHINX can achieve 70.8% classification accuracy without any data augmentation. This convergence speed is much faster than traditional approaches, such as ResNet (He et al., 2016b) and ViT (Dosovitskiy et al., 2020) that normally take around 300 training epochs and require strong data augmentation.
Experiments
As mentioned in Section 3.1, our training pipeline consists of two stages. In stage 1, or the Pre-training stage, we start from a text-only LLM, and build the multi-modal capabilities from scratch with large-scale noisy datasets. In stage 2, or the fine-tuning stage, we extract the strong capabilities learned in stage 1 on practical tasks by further training with diverse and high-quality instruct-following datasets. The construct of the datasets and the training configuration for both stages are detailed as follows.
We use two image captioning datasets LAION-400M (Schuhmann et al., 2021) and LAION-COCO (Schuhmann et al., 2022) for multi-modal alignment. As we full-fine-tune the language model backbone for long steps, we also jointly train with a text-only dataset RefinedWeb (Penedo et al., 2023) to avoid harming its text reasoning capability due to catastrophic forgetting.
We fine-tune the weight of the large language model and the visual projections in the pre-training stage, among which the weight of large language model is initialized from off-the-shelf open-source weights such as LLaMA-2 (Touvron et al., 2023b) and the visual projections are initialized randomly. The visual encoders themselves are kept frozen with their originally pre-trained weights throughout the training. We use the AdamW optimizer (Kingma & Ba, 2014) with , a cosine annealing learning rate schedule for steps from to with the first steps being a linear warm-up from to , and a constant weight decay of . For the joint training on both images and texts, we form each batch with image-text pairs from LAION-400M or LAION-COCO and text tokens from RefinedWeb. Since captions in LAION-400M and LAION-COCO are based on web-crawled data and generally do not contain much fine-grained information, we only utilize one global view of each image, i.e., the low resolution of 224224, for faster training. We do not apply any form of language prompts during pre-training. The pre-training time is around 125 hours on 32 A100 GPUs with a 7B language model and about twice the time with a 13B language model.
In the multi-task fine-tuning phase, our objective is to equip the MLLM with the versatile needs of downstream tasks. Building upon insights from prior research (Liu et al., 2023d; Dai et al., 2023; Chen et al., 2023b; Zhu et al., 2023; Liu et al., 2023b), we include instruction following data such as LLaVA (Liu et al., 2023d) and ShareGPT (ShareGPT, 2023), exposing the model to tasks requiring explicit directives. For general Vision Question Answering (VQA), we leverage datasets like VQAV2 (Agrawal et al., 2015) and GQA (Hudson & Manning, 2019). Expanding the scope to out-of-domain knowledge, we integrate datasets like OKVQA (Marino et al., 2019) and A-OKVQA (Schwenk et al., 2022), providing the model with information beyond the training data. Optical Character Recognition (OCR) datasets, such as OCRVQA (Mishra et al., 2019) and TextCaps (Sidorov et al., 2020) are utilized to increase the text understanding ability of SPHINX. We introduce abundant general object detection and pose estimation datasets, such as COCO (Lin et al., 2014) and LVIS (Gupta et al., 2019) to inspire the model’s capabilities of localization, classification, and human pose estimation. To address grounding tasks, we incorporate RefCOCO (Kazemzadeh et al., 2014) and VG (Krishna et al., 2017) datasets, training the model to handle referring object localization. Additionally, Grounding Caption datasets, such as those from Flickr30k (Plummer et al., 2015), further refine the understanding of descriptions in the context of image regions. Despite the diversity of data sources, we streamline the training by converting all datasets into a multi-turn conversation format. This not only reduces training costs but also enhances overall efficiency.
The trained and frozen network components are identical as the pre-training stage. The optimizer settings are similar to the pre-training stage, except that we use a batch size of 128, a maximum learning rate of , a minimum learning rate of 0, and a linear warmup for epoch during fine-tuning. Training data are sampled from the mixture of datasets following their natural frequencies, i.e., the chance of a dataset being sampled from is proportional to its original size. We follow the image preprocessing steps of (Chen et al., 2023b; Liu et al., 2023b), i.e., padding the image along the shorter edge to make it a square before resizing, for better handling of images with extreme aspect ratios. The fine-tuning takes about 38 hours with 16 A100 GPUs with a 13B language model. The maximum training sequence length is set to 3072.
2 Quantitative evaluation
In this section, we provide a comprehensive evaluation of SPHINX and showcase results across multiple benchmarks. Our evaluation encompasses both quantitative metrics and qualitative assessments, providing a holistic understanding of our VLM model’s performance.
We show in Figure 7 the effectiveness of introducing a text-only dataset (i.e., RefinedWeb) to jointly train with image captioning in the pre-training stage. We design an experiment using only vision-language data and without using RefinedWeb. We observe that the text-only loss grows if the model is not trained with RefinedWeb, showing that our joint-training scheme is effective in preserving the text-modeling capability while adapting for cross-modal understanding.
In our model evaluation, we prioritize aligning with each benchmark’s desired output format. To achieve this, we employ distinct prompts tailored to benchmarks that necessitate long answers, short answers, and multiple-choice responses. The detailed information is provided in Table 1. This approach ensures that our model is capable of handling diverse scenarios.
We denote the fundamental variant of our MLLM as SPHINX, which takes as input a low-resolution image of 224224, and produces 289 visual tokens (257 from the mixed CLIP (Radford et al., 2021) and DINOv2 (Oquab et al., 2023), and 32 from Q-Former (Li et al., 2023d)). Then, we denote our high-resolution variant as SPHINX-1k and SPHINX-2k. SPHINX-1k processes the image resolution of 448448 by evenly dividing four sub-images with 1,445 visual tokens, i.e., five groups of 289 tokens (one group for downsampled image and four groups for sub-images). SPHINX-2k further processes a higher resolution of 762762 with evenly divided nine sub-images of 2,890 visual tokens, i.e., ten groups of 289 tokens.
We test our model on recently proposed MLLM benchmarks to comprehensively evaluation of the model’s characteristic such as MME (Fu et al., 2023b), Seedbench (Li et al., 2023c), POPE (Li et al., 2023e), LLaVA-Bench (In-the-Wild) (Liu et al., 2023d), MM-Vet (Yu et al., 2023b), MathVista (Lu et al., 2023), MMbench (Liu et al., 2023g), CCbench (Contributors, 2023), Tiny LVLM (Shao et al., 2023) and Touchstone (Bai et al., 2023b). We show the result in Table 2. We observe that the SPHINX surpasses previous state-of-the-art MLLM performances on 6 out of 10 benchmarks. We compare our model with strong baselines including BLIP-2 (Li et al., 2023d), InstructBLIP (Dai et al., 2023), Shikra (Chen et al., 2023b), Qwen (Bai et al., 2023a), Fuyu (Bavishi et al., 2023) and LLaVA1.5 (Liu et al., 2023b). The gap between SPHINX and SPHINX-1k on POPE suggests that the introduction of high-resolution sub-images can significantly improve visual hallucination problems.
Furthermore, we evaluate general VQA benchmarks, such as VQAV2 (Agrawal et al., 2015), OKVQA (Marino et al., 2019), GQA (Hudson & Manning, 2019), vizwiz (Gurari et al., 2018), ScienceQA (Lu et al., 2022), visual spatial reasoning (VSR) (Liu et al., 2023a), IconQA (Lu et al., 2021). Additionally, we conduct experiments on Text-oriented VQA such as TextVQA (Singh et al., 2019), OCR-VQA (Mishra et al., 2019). We provide the results in Table 3. SPHINX achieves comparative results across all benchmarks. We observe that SPHINX-1k and SPHINX-2k significantly outperform SPHINX in VQAv2 datasets and text-oriented VQA that demand fine-grained visual information, showcasing the effectiveness of our visual mixed-up approach for achieving high resolution without relying on a visual encoder trained specifically on high-resolution images. Although the performances of SPHINX on text-oriented VQA surpass strong baselines, such as BLIP-2 and InstructBLIP, it is still below Qwen-VL-7B due to the lack of text-related pre-training data. In the future, we will introduce more text-related pre-training datasets.
Table 4 evaluates SPHINX on REC benchmarks with RefCOCO (Kazemzadeh et al., 2014), RefCOCO+ (Mao et al., 2015), and RefCOCOg (Mao et al., 2015) datasets. SPHINX outperforms most state-of-the-art models, including specialist model G-DINO-L Liu et al. (2023e) and other visual-language generalist models. Compared to a recent strong baseline Qwen-VL-7B (Bai et al., 2023a), which also leverages the large language model for visual understanding, our model still achieves better results across all splits by a large margin. Moreover, SPHINX-1k and SPHINX-2k enable the use of high-resolution input images, leading to consecutive improvement over SPHINX and narrowing down the gap to the strong specialist model UNINEXT, which adopts a larger input image size. These results demonstrate the competitive capability of SPHINX for visual grounding.
3 Demonstrations
In this section, we present the qualitative outcomes of SPHINX, showcasing its capabilities in SAM-assisted segmentation, general object detection, human pose estimation, document layout detection, anomaly detection, and etc. Surprisingly, SPHINX also exhibits improved performance on the chain of thoughts and obtains emergent cross-task abilities.
We integrate SPHINX with SAM to enhance segmentation capabilities. This integration involves detecting bounding boxes for the target objects and subsequently providing the bounding box coordinates to SAM for the generation of segmentation masks. The results, depicted in Figure 8, showcase a notable performance improvement achieved through the collaboration of SPHINX and SAM. Surprisingly, We observe that the predicted masks for small objects are extremely accurate such as the cell phone in the last row. The synergistic application of SPHINX and SAM underscores the considerable potential inherent in our methodology.
In Figure 9, the performance of SPHINX ’s detection capabilities is showcased. The upper row displays the synchronized jumping of five teenagers, each assuming distinct poses. Notably, SPHINX accurately predicts the pose with key points for each individual, leaving no participant overlooked. The middle row illustrates the SPHINX ’s reasoning ability to focus on a specified region. We observe that SPHINX successfully recognize the desired objects and detailed answer to the question. The bottom row indicates SPHINX ’s superior diagram understanding ability, which produces accurate layout detection and content comprehension.
The enhanced visual reasoning capabilities of our model with object detection are showcased in Figure 10. Notably, SPHINX leverages the object detection feedback by initially instructing SPHINX to generate object detection results and then requesting it to answer questions based on localization outcomes. The model will prioritize selecting the most relevant objects for coordinate feedback based on the query content, rather than all detected objects. This underscores the idea that in multi-task training, the synergy between different tasks can significantly enhance overall performance. Furthermore, the model exhibits commendable Contextual Understanding (COT) by effectively integrating information from diverse elements in the image, resulting in more powerful reasoning ability.
We highlight SPHINX’s proficiency in understanding user hints. As depicted in Figure 10, initially requesting the model to predict all dogs in the image leads to the misidentification of other objects. However, upon offering additional hints about the desired object, SPHINX demonstrates an improved comprehension of instructions and accurately predicts all dogs in the image.
The original referring object comprehension and pose estimation are two different tasks, where the former detects object bounding boxes according to textual descriptions, and the latter outputs human keypoints from given bounding boxes. Interestingly, as shown in Figure 11 (Top), by our mixing of the two tuning tasks, our SPHINX acquires the emergent capacity for referring pose estimation, i.e., generating human keypoints directly from textual descriptions. Such an observation indicates that our SPHINX fully comprehend the semantics across different vision-language tasks, and implicitly connect them via superior reasoning power.
It is important for industrial monitoring and healthcare to detect rare events or outliers that may indicate abnormal or suspicious behavior. As shown in Figure 11 (Bottom), our SPHINX also excels in anomaly detection. Although we do not explicitly involve related training data, our MLLM still demonstrates superior localization accuracy for unsharp defects. This indicates wide potentials of SPHINX in real-world applications.
Endowed with diverse multi-task pre-training, SPHINX can perform multi-level dense captioning by iterative promoting itself. Given an input image, prompting SPHINX with “Detect all objects shown in the image” can localize the position of all objects. Then, we iteratively prompt each detected region with “Please provide a short description for this region : [x1, y1, x2, y2]” to extract a simple property on the localized region. To get a deeper understanding on the detected regions, we crop all images based on the detection results. Each cropped view is fed independently into SPHINX with two prompts, namely, “Provide a one-sentence caption for the provided image.” and “Generate a detailed description about the image.”. By doing so, we can detect all objects shown in the image and densely label all boxes with property, simple caption, and detailed caption. The multi-level dense captioning results are illustrated in Figure 12.
Conclusion
In this paper, we propose SPHINX, a versatile multi-modal large language model (MLLM) with multi-purpose visual instruction-following capabilities. In our MLLM, we introduce a joint mixing of three different aspects: model weights of pre-trained LLMs by real-world and synthetic data, tuning tasks for diverse visual perception and reasoning tasks, and visual embeddings from different types of vision backbones. On top of this, we further devise to endow SPHINX with the capacity to process high-resolution images by mixing different visual scales and sub-images, which exhibits superior fine-grained visual understanding performance. Via our proposed three-fold mixing strategy, SPHINX achieves impressive performance over a wide range of multi-modality evaluation benchmarks, and can serve as a strong vision generalist to tackle object detection, region-level captioning, and human pose estimation, etc. Our MLLM can also be integrated with other visual foundation models for wider functionalities, e.g., SAM (Kirillov et al., 2023) for language-referred segmentation and Stable Diffusion (Rombach et al., 2021) for image editing. Our future work will focus on incorporating a wider range of vision-language tasks into SPHINX for all-purpose capabilities.