Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xiaoshuai Sun, Rongrong Ji

Introduction

Driven by the remarkable success of large language models (LLMs) (Touvron et al., 2023; Chen et al., 2020), research on multi-modal large language models (MLLMs) also receives an influx of interest in the machine learning community (Liu et al., 2023b; Luo et al., 2023a; Alayrac et al., 2022; Chen et al., 2022, 2023b). Numerous efforts have been recently devoted to extending LLMs to more modalities, achieving breakthroughs on various vision-language tasks (Goyal et al., 2017; Singh et al., 2019; Hudson & Manning, 2019). Despite advances, existing MLLMs still fall short of granular visual recognition. For instance, the powerful GPT4-V also suffers from hallucinations when identifying small and occluded objects (Tong et al., 2024). This shortcoming inevitably limits the practical use of MLLMs.

To compensate for this shortcoming, practitioners often resort to scaling up model size and increasing per-training data size (Alayrac et al., 2022; Li et al., 2023b; Bai et al., 2023). For instance, InstructBLIP (Dai et al., 2023) adopts over 129M image-text pairs for vision-language (VL) alignments, and shows that a larger visual encoder is beneficial for MLLMs. Motivated by this, Qwen-VL (Bai et al., 2023) further increases the parameters of visual encoder to 1.9 billion and uses 1.5 billion pre-training data. Despite progress, this paradigm is prohibitively expensive, which often consumes about thousands of GPU hours.

Orthogonal to these works, we study the visual shortcoming of MLLMs from the perspective of input image resolutions. As revealed in previous VL research (Jiang et al., 2020; Tong et al., 2024; Luo et al., 2023b), increasing the resolution of input images is a straightforward solution to improve visual recognition, which becomes more important for MLLMs that involve visual chain-of-thought (Rose et al., 2023). As shown in Fig. 1, increasing the resolution of LLaVA-1.5 (Liu et al., 2023a) from 384 ×\times 384 to 672 ×\times 672 can bring obvious performance gains (+4.6%) on TextVQA (Singh et al., 2019). However, the use of high-resolution images will greatly exacerbate the already high computational cost of MLLMs. For instance, 448×448448\times 448 resolution will increase the computation complexity of LLaVA by about 1.4 times compared with the default 336×336336\times 336. In addition, due to the complex structure of MLLMs, the training will become unstable as the resolution is greatly increased, e.g., a sharp drop at 1,022×1,0221,022\times 1,022 resolution, as shown in Fig. 1. We assume that the length of visual sequences greatly exceeds the pre-trained context length, leading to training instability.

In this paper, we propose a novel and efficient method for the high-resolution image adaptation of MLLMs, namely mixture-of-resolution adaptation (MRA). As shown in Fig. 1, MRA adopts an innovative dual visual pathway design to process the input images of high- and low-resolutions simultaneously. Specifically, one pathway aims to encode global information of low-resolution images, while the other one serves to capture fine-grained semantics from high-resolution images. Meanwhile, these two pathways are closely interacted via the novel mixture-of-resolution adapters (MR-Adapters), which embeds the high-resolution visual information into the low-resolution modeling. In this way, we can use a much fewer number of visual tokens to represent the input images from macro- to micro-views. With the careful design of dual-pathway structure, MRA can easily increase the image resolution up to 1,536 ×\times 1,536 pixels while maintaining high efficiency.

To validate MRA, we apply it to a recent MLLLM called LLaVA (Liu et al., 2023b, a), and term the new model as LLaVA-HR. We conduct extensive experiments on 11 vision-language (VL) tasks, including common VL tasks like VQA2.0 (Goyal et al., 2017) and emerging benchmarks such as POPE (Li et al., 2023c). Experimental results show that LLaVA-HR outperforms existing MLLMs on 8 of 11 VL tasks, e.g., +9.6% over LLaVA-1.5 on TextVQA. More importantly, the training and inference of LLaVA-HR are cost-effective. The pre-training and instruction tuning of LLaVA-HR (7B, 1,024 ×\times 1,024) only take a total of 20.7 hours on 8 A800 GPUs, which is hundreds of times cheaper than InstructBLIP (Dai et al., 2023) and Qwen-VL (Bai et al., 2023). With the same resolution, its inference speed is 3 times faster than LLaVA-1.5 (Liu et al., 2023a).

In summary, our contributions are three folds:

We reveal the significance of image resolution for MLLMs and propose a novel and efficient adaptation scheme, termed mixture-of-resolution adaption (MRA), which adopts a novel dual visual pathway design to obtain the benefits of high-resolution visual information while keeping training and inference efficient.

We propose a novel mixture-of-resolution adapter (MR-Adapter) for MRA, which can embed the high-resolution information into the low-resolution visual pathway to improve visual descriptive power.

Based on MRA, we propose a powerful MLLM, coined LLaVA-HR, which outperforms existing MLLMs on 8 of 11 VL tasks and spends much cheaper training expenditure than most MLLMs.

Related Work

Driven by the great successes of large language models (LLMs) (Gilardi et al., 2023; Touvron et al., 2023; Chen et al., 2020), growing interest has been aroused in building end-to-end multimodal large language models (MLLMs) (Liu et al., 2023b; Zhu et al., 2023; Luo et al., 2023a; Fuyu-8B, 2023; Peng et al., 2023; Liu et al., 2023c). In particular, most existing MLLMs adopt a modular structure (Luo et al., 2023a; Liu et al., 2023b), which utilizes an intermediate network to project the visual features into the word embedding space of the LLM. Then, the LLM is used to accomplish various VL tasks in an autoregressive manner. Based on the modular structure, existing MLLMs can be distinguished by the designs of the intermediate network. Popular MLLMs represented by LLaVA (Liu et al., 2023b) often adopt a linear projection layer or an MLP layer to connect the visual encoder and the LLM (Liu et al., 2023b, a; Chen et al., 2023a, b; Peng et al., 2023). The other works employ sampler-based modules to bridge the gap between the visual encoder and the LLM (Bai et al., 2023; Alayrac et al., 2022; Li et al., 2023b; Dai et al., 2023). These sampler-based modules can effectively reduce the number of visual tokens, but often requires a large-scale pre-training to achieve a promising performance (Bai et al., 2023; Li et al., 2023b). Despite the effectiveness, most existing MLLMs still adopt a low visual resolution, e.g., 336 ×\times 336, which greatly limits their performance in fine-grained tasks.

2 Visual Representations for MLLMs

The pursuit of better visual representations has been a popular research trend in the VL community (Lu et al., 2019; Jiang et al., 2020; Radford et al., 2021; Ren et al., 2024). Early endeavors mainly explore the object-level features for VL models (Lu et al., 2019; Zhang et al., 2021). Driven by the large-scale image-text pre-training, grid features from CLIP (Radford et al., 2021) have demonstrated the great efficiency and generalization in MLLMs (Liu et al., 2023b; Chen et al., 2022; Alayrac et al., 2022). Based on grid features, existing researchers mainly improve visual representations by scaling up the visual encoder. For example, PaLI (Chen et al., 2022) increases the parameters of visual encoder to 3 billions and shows the significant performance boost of MLLMs. In contrast to these works, we improve the visual representations for MLLMs from the perspective of image resolution, and propose a novel and efficient solution, namely mixture-of-resolution adaptation.

Preliminary

We first recap the structure of multimodal large language models (MLLMs), which consists of an image encoder FI(⋅)\mathcal{F_{I}(\cdot)}, an intermediate network FP(⋅)\mathcal{F_{P}(\cdot)} and an LLM FL(⋅)\mathcal{F_{L}(\cdot)}.

In some MLLMs (Liu et al., 2023b, a), FP(⋅)\mathcal{F_{P}}(\cdot) is often a stack of simple linear layers, which are used to directly project the visual tokens onto the semantic space of LLMs. Although simple and effective, this strategy inevitably leads to a longer visual sequence as the resolution increases, e.g., 5,329 tokens for 1,022 ×\times 1,022 resolution in LLaVA-1.5. In practice, processing such a large number of tokens is computationally expensive in MLLMs. To further reduce the number of visual tokens, recent advances adopt the sampler-based module for FP(⋅)\mathcal{F_{P}}(\cdot) , e.g., QFormer (Li et al., 2023b), which aggregates visual features into several tokens that LLM can directly handle. Nevertheless, these methods often require large-scale pre-training to achieve VL alignments (Bai et al., 2023; Li et al., 2023b).

Based on the above analyses, we conclude that the main difficulty of high-resolution image adaptation lies in the rapidly growing visual sequence. This issue motivates us to further explore how to efficiently encode richer visual information with fewer visual tokens.

Mixture-of-Resolution Adaptation

To address the above issues, we propose a novel and efficient method for MLLMs, termed mixture-of-resolution adaptation (MRA), of which structure is depicted in Fig. 3. The core idea of MRA is to embed high-resolution information into the low-resolution one via a dual pathway design. In this case, MRA can keep a smaller number of visual tokens while encoding richer visual information.

2 Dual Visual Pathways

As shown in Fig. 3, dual visual pathways are the key design of MRA, and their benefits are maximized from two aspects.

Visual functionality. Firstly, the dual visual pathways process images from macro- and micro-views, which is inspired by the visual system of human being (Merigan & Maunsell, 1993; Robertson & Lamb, 1991). Particularly, Robertson & Lamb (1991) find that the visual system processes local and global semantics via different pathways. Based on this finding, we adopt a similar mechanism to our MRA. Specifically, one visual pathway aims to capture fine-grained semantics from high-resolution images i.e., processing images from local view. In contrast, the other pathway is designed to encode global information from low-resolution images, achieving a larger receptive field.

Visual alignment. Due to different resolutions, these two pathways often produce visual features of different shapes, impeding their quick alignments (Yu et al., 2019). To overcome this limitation, we adopt different downsampling rates for the low- and high-resolution pathways, respectively. Thus, their output features can keep the same spatial shape.

Based on the above observations, we design the dual visual pathways with a convolutional network (CNN) (Liu et al., 2022) and a vision transformer (ViT) (Dosovitskiy et al., 2020). Specifically, CNN is equipped with a downsampling stride of 32 to process high-resolution images. ViT encodes low-resolution images with a downsampling stride of 14. Notably, such designs also ensure the efficiency of MLLMs, where the high-resolution images are processed by the efficient CNN, and the number of visual tokens is also kept small via the large downsampling stride.

3 Mixture-of-Resolution Adapter

As shown in Fig. 3, high-resolution information can be fused with the features in each block of ViT. In this case, the low-resolution features of ViT also contain rich semantics, improving the visual descriptive power of MLLMs.

4 The Deployment on MLLM

We apply MRA to a popular MLLM called LLaVA-1.5 (Liu et al., 2023a), and construct a new model, namely LLaVA-HR. Its training consists of two stages, i.e., low-resolution pre-training and high-resolution instruction tuning.

Stage 1: Low-Resolution Pre-training. Similar to LLaVA (Liu et al., 2023b) and LLaVA-1.5 (Liu et al., 2023a), this stage aims to optimize the projector to align the visual features with the word embeddings of LLM. Therefore, the image encoder and the LLM are frozen during pre-training. Besides, we adopt low resolutions for two pathways. In this stage, the MR-Adapter is not inserted, and output features of dual pathways are directly combined.

Stage 2: High-Resolution Instruction Tuning. During instruction tuning, we greatly increase the resolution of the high-resolution pathway, e.g., from 384×\times 384 to 1,024×\times 1,024. And the low-resolution one is also accordingly adjusted to ensure the visual alignment of two pathways, e.g., from 336×\times 336 to 448×\times 448. Meanwhile, the MR-Adapter is then applied to connect two visual pathways. Different from the first training stage, the entire MLLM will be fully optimized to better accommodate high-resolution images.

Experiments

Multimodal benchmarks for MLLM. We evaluate LLaVA-HR on four emerging multimodal benchmarks for MLLMs, including MME (Fu et al., 2023), POPE (Li et al., 2023c), SEED (Li et al., 2023a) and MM-VET (Yu et al., 2023). In particular, MME and MM-VET evaluate the multimodal perception and cognition abilities of MLLMs. SEED extends the modalities of evaluation to images and videos. POPE aims to evaluate the visual hallucinations of MLLMs. The metrics used in our paper follow their default settings. For MME, we follow LLaVA-1.5 to report the perception score.

Common vision-language benchmarks. We also evaluate LLaVA-HR on seven VL datasets, including VQAv2 (Goyal et al., 2017), GQA (Hudson & Manning, 2019), OKVQA (Marino et al., 2019), OCRVQA (Mishra et al., 2019), ScienceQA (Lu et al., 2022), VizWiz (Gurari et al., 2018) and TextVQA. In particular, ScienceQA (Lu et al., 2022), VizWiz (Gurari et al., 2018) and TextVQA are three zero-shot tasks, and their samples are not appeared in our training data. We report the accuracy on the test set of OCRVQA, the test set of VizWiz, and the val set of OKVQA. We organize samples of these tasks in instruction formats of LLaVA-1.5 (Liu et al., 2023a).

2 Implementation Details

In LLaVA-HR, we use CLIP-ViT-L (Radford et al., 2021; Ilharco et al., 2021) and CLIP-ConvNeXt-L (Liu et al., 2022) as the dual visual paths to encode low- and high-resolution images, respectively. In LLaVA-HR-X, the CLIP-ConvNeXt-L is replaced with the stronger CLIP-ConvNeXt-XXL. The MR-Adapter is applied into the last three stages of ViT. Following LLaVA-1.5, we first pre-train LLaVA-HR on LCS-558K (Liu et al., 2023b), which contains 558k image-text pairs. During the pre-training stage, both the visual encoder and the LLM are frozen, and only the MLP projector is fine-tuned. AdamW (Kingma & Ba, 2014) is used as the optimizer, and the learning rate and batch size are set to 1e-3 and 256, respectively. Visual resolutions are set to 336×\times336 and 384×\times384 for the ViT and the CNN, respectively. During instruction tuning, we follow LLaVA-1.5 to use 665k VL instruction data. At this stage, the entire model is updated with a learning rate of 2e-5. Besides, we increase the resolution of ViT and CNN to 448×\times448 and 1,024×\times1,024, respectively. The training epoch is set to 1 for pre-training and instruction tuning.

3 Experimental Results

Comparison with baselines. In Tab. 1, we compare the performance and efficiency of LLaVA-HR with LLaVA-1.5 (Liu et al., 2023a) with different image resolutions. From this table, we observe that increasing image resolution obviously improves the performance of two models on four tasks, e.g., +4.8% of LLaVA-1.5 on TextVQA. However, the performance of LLaVA-1.5 drops significantly at the resolution of 1,024×\times1,024. To explain, the number of visual tokens greatly exceeds the pre-trained context length of the LLM, which easily causes the instability during training. In contrast, the performance of LLaVA-HR is consistently improved from 384 ×\times 384 resolution to 1,024 ×\times 1,024 resolution. Besides, the total gain of LLaVA-HR is more obvious than that of LLaVA-1.5 (Liu et al., 2023a), e.g., +8.33% of LLaVA-HR vs. +4.82% of LLaVA-1.5, greatly confirming the effectiveness of MRA.

In Tab. 2, we further compare four common baselines with the similar resolution, i.e., ∼\sim760×\times760. “ViT+MLP” is the default setting of LLaVA-1.5 as the reference. “Conv+MLP” replaces the visual backbone with ConvNeXt (Liu et al., 2022), which uses a larger downsampling rate to reduce the number of visual tokens. “ViT+Resampler” and “ViT+Pooling+MLP” refer to the two pooling strategies for reducing the number of visual tokens. As can be seen, all compared methods are inferior to LLaVA-HR. In particular, using a convolutional network as the visual backbone greatly improves efficiency, but its performance still lags behind LLaVA-HR by a large margin, e.g., -108.9 on MME (Fu et al., 2023). Similarly, “ViT+Resampler” and “ViT+Pooling+MLP” also sacrifice performance for efficiency. Overall, these comparisons further confirm the designs of MRA.

Despite effectiveness, the expenditure of LLaVA-HR is also cost-effective. In particular, increasing resolution from 384 ×\times 384 to 1,024 ×\times 1,024 slows down the training and inference of LLaVA-1.5 by 344.8% and 325%, respectively. However, these costs are reduced to only 17.6% and 20.8% in LLaVA-HR. Despite better performance, the training and inference speeds of LLaVA-HR are three times faster than LLaVA-1.5. Besides, the costs of GPU memory also remain cheap for LLaVA-HR. For example, adapting the resolution of 1,536 ×\times 1,536 for LLaVA-HR only consumes 52G GPU memory, but the same settings for LLaVA-1.5 will cause GPU memory overflow. These results greatly confirm the efficiency of our MRA and LLaVA-HR.

Ablation studies. In Tab. 3, we conduct comprehensive ablation studies for MRA on four VL benchmarks. Firstly, we validate the different designs of the dual visual pathways. From these results, we find that removing one pathway will lead to significant performance drops, e.g., -1.5% on VQAv2. Besides, scaling up the high-resolution encoder brings more gains than that of the low-resolution one, e.g., +2.1% vs. +0.9% on TextVQA. We assume that the stronger high-resolution image encoder can better capture the fine-grained visual information. Then, we ablate different fusion directions and strategies in MRA. Specifically, changing the fusion direction obviously degenerates the performance, e.g., -61.3 on MME. Finally, we ablate the designs of the mixture-of-resolution adapter. Specifically, the best choices of mapping modules for the low- and high-resolution pathways are convolution blocks and MLP blocks, respectively. Besides, the choices of gating function also affect performance and the tanh function perform the best. These ablations further confirm the designs of MR-Adapter.

Comparison with existing MLLMs. In Tab. 4 - 5, we compare LLaVA-HR with existing MLLMs on 11 VL tasks. On the four MLLM benchmarks, we observe comprehensive advantages of LLaVA-HR against existing MLLMs. In particular, LLaVA-HR achieves 1554.9 scores in MME benchmark, outperforming LLaVA-1.5 by +23.6. On POPE, a benchmark including video evaluations, LLaVA-HR-X still outperforms existing MLLMs by a large margin, i.e., +3.7% gains. Besides, LLaVA-HR achieves the best performance on the benchmark for visual hallucinations, i.e., POPE, suggesting that its visual hallucinations are greatly alleviated. Notably, Fuyu-8b (Fuyu-8B, 2023) is capable of high-resolution images, but its performance is much inferior to LLaVA-HR, e.g., 728.6 vs. 1554.9 on MME.

Tab. 5 gives the performance comparison on common VL tasks. On in-domain tasks, LLaVA-HR achieves the best results on three tasks, e.g., 82.6 on VQAv2 and 61.5 on OKVQA. On OCRVQA, Qwen-VL-Chat collects more in-domain data for training, so it performs better than LLaVA-HR. Under the zero-shot setting, we can observe more significant advantages of LLaVA-HR on the fine-grained tasks, e.g., VizWiz and TextVQA. Most notably, even Qwen-VL-Chat is pre-trained with 24.8M OCR samples, it still performs worse than LLaVA-HR-X on TextVQA. These results suggest the significance of high resolution for these tasks. In contrast, most images of ScienceQA are synthetic and of low resolution, so the advantages of LLaVA-HR are not obvious. Overall, these results greatly confirm the effectiveness and generalization of LLaVA-HR and our MRA.

3.2 Qualitative Experiments

In Fig 2 (a), we compare the predictions of LLaVA-HR with different resolutions. The visualizations show that higher image resolution obviously improves the capability of MLLMs on fine-grained tasks. For example, LLaVA-HR with a resolution of 1,024 ×\times 1,024 can well capture granular visual content, e.g., the tiny boat in the first example. Besides, high image resolution also enables LLaVA-HR a stronger ability of text recognition. For instance, the small and blurred phrase of “wo ich wohne” in the second example are correctly identified by the high-resolution LLaVA-HR. These results greatly confirm the significance of high image resolution in addressing visual shortcoming. In Fig 2 (b), we further compare the predictions of LLaVA-HR-X, LLaVA-1.5 (Liu et al., 2023a) and GPT4-V (OpenAI, 2023) in visual information extraction. Notably, LLaVA-HR-X shows a comparable ability with GPT4-V on this challenging task. As shown in Fig 2 (b), LLaVA-HR-X and GPT4-V can correctly extract almost all visual content of the driver license and organize it in JSON format. Compared to GPT4-V, LLaVA-HR-X also correctly identifies the hair color of the person, which requires fine-grained visual reasoning. In contrast, LLaVA-1.5 can only recognize simple visual content like “class” and “SEX”, and fail to extract most visual information. These results further validate the effectiveness of MRA in addressing visual shortcoming of MLLMs.

Conclusion

In this paper, we study the visual shortcoming of MLLMs from the perspective of image resolution, and propose a novel and efficient method for high-resolution adaptations of MLLMs, namely mixture-of-resolution adaptation (MRA). MRA adopts dual visual pathways to process images of both high and low resolutions, where high-resolution information is embeded into the low-resolution modeling via the novel mixture-of-resolution adapters (MR-Adapters). We apply MRA to a popular MLLM called LLaVA-1.5, and construct a new high-resolution MLLM, termed LLaVA-HR. Experimental results not only validate the effectiveness of LLaVA-HR in addressing visual shortcoming, but also confirm its remarkable efficiency against existing MLLMs.

This work was supported by National Key R&D Program of China (No.2022ZD0118201) , the National Science Fund for Distinguished Young Scholars (No.62025603), the National Natural Science Foundation of China (No. U21B2037, No. U22B2051, No. 62176222, No. 62176223, No. 62176226, No. 62072386, No. 62072387, No. 62072389, No. 62002305 and No. 62272401), the Natural Science Foundation of Fujian Province of China (No.2021J01002, No.2022J06001), and the China Fundamental Research Funds for the Central Universities (Grant No. 20720220068).

References