MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Zhuofan Zong, Bingqi Ma, Dazhong Shen, Guanglu Song, Hao Shao, Dongzhi Jiang, Hongsheng Li, Yu Liu

Introduction

Significant achievements in multimodal large language models (MLLMs) have been witnessed due to their remarkable proficiency in solving open-world tasks. MLLMs acquire visual perception capacity while inheriting sophisticated reasoning abilities and knowledge from large language models (LLMs). The core idea behind MLLMs is projecting the vision encoder representation into an LLM through a projector, facilitating a general-purpose multimodal understanding.

General multimodal understanding requires comprehending complex image contexts across various tasks and scenarios. The CLIP vision encoder, pre-trained on large-scale image-text pairs with a contrastive loss, is widely considered as a flexible and popular choice among the latest leading MLLMs. However, training data and optimization target of the vision encoder determine its inconsistent performance across tasks and scenarios, which will bias the generalization of multimodal large language models. For instance, MLLMs with a single CLIP vision encoder usually perform poorly on fine-grained tasks such as grounding and optical character recognition (OCR) . Several works have attempted to incorporate extra state-of-the-art vision encoder experts to cope with the challenge. For example, both SPHINX and MoF integrate vision self-supervised learning features of DINOv2 with MLLMs to enhance their visual grounding capabilities. Vary introduces a new vision encoder expert for improved fine-grained document and chart parsing ability. Intuitively, it is necessary to explore the utilization of more task-specific vision encoder experts in MLLMs to promote model generalization across various domains.

We aim to start the exploration through empirical analysis of readily available vision experts. In particular, we focus on the multimodal capabilities of seven distinct state-of-the-art vision encoders based on LLaVA-1.5-7B . The results in Tab. 1 reveal that MLLMs with these task-specific vision encoders achieve optimal performance in their respective area. Concurrently, we note that the plain fusion (concatenation) of vision encoder experts adopted in previous works (e.g., SPHINX) would not bring consistent improvement compared with the single task-specific vision expert in its proficient task. The inherent bias of each expert introduces biased information and leads to performance degradation in the plain fusion paradigm. For example, DINOv2 serves as an expert in visual grounding but performs poorly at text-oriented tasks. Representation of DINOv2 would be regarded as biased information in text-related scenarios so incorporating DINOv2 for these tasks would inevitably cause performance decrease. Consequently, a flexible method of vision encoder ensemble that dynamically activates and weights context-relevant task-specific vision experts can fully unleash the capacity of these models while avoiding model bias.

In this paper, we propose MoVA, a powerful MLLM, adaptively routing and fusing task-specific vision experts with a coarse-to-fine mechanism. Inspired by the powerful tool-use capabilities of LLM, the coarse-grained context-aware expert routing aims to employ LLM to select vision experts with strong relevance to the user’s image and instruction from the expert model pool. To improve the efficiency and effectiveness of context-aware expert routing, we integrate expert-routing low-rank adaptation (LoRA) into the LLM component of MoVA. The fine-grained expert fusion facilitates better extraction and integration of expert representations based on multimodal context. Specifically, the expert knowledge extractor in the mixture-of-vision-expert adapter (MoV-Adapter) will extract diverse task-specific knowledge from various vision experts through mixture-of-expert (MoE) cross-attention layers. The dynamic gating network can allocate precise expert-wise soft weights for the integration of extracted task-specific knowledge. Under the coarse-to-fine paradigm, we provide a flexible and effective manner of leveraging representation from experts based on multimodal context and model expertise, further enhancing the model generalization ability.

We conduct comprehensive experiments on various benchmarks to evaluate the effectiveness of MoVA, including MLLM benchmarks, visual question answering (VQA), visual grounding, image segmentation, and biomedical understanding. Without any bells and whistles, MoVA can achieve significant performance gains over current state-of-the-art methods.

To sum up, the key contributions of this paper are as follows:

(1) By analyzing the performance of individual vision encoders versus the plain fusion of multiple encoders across various tasks, we reveal that the inherent bias of each vision encoder can diminish its generalization ability across other irrelevant domains.

(2) We propose MoVA, a powerful MLLM composed of coarse-grained context-aware expert routing and fine-grained expert fusion with MoV-Adapter. Based on multimodal context and model expertise, MoVA fully leverages representation from multiple context-relevant vision encoder experts flexibly and effectively while avoiding biased information brought by irrelevant experts.

(3) We demonstrate the effectiveness of each component in MoVA by elaborate ablation studies. Without any bells and whistles, MoVA can achieve significant performance gains over current state-of-the-art methods in a wide range of challenging benchmarks.

Related Work

Large language models (LLMs) have achieved remarkable progress in Natural Language Processing (NLP). The emergence of GPT-3 demonstrates that models will manifest profound capabilities in few-shot learning and zero-short learning with increasing model parameters, training data, and training computation . The powerful conversational and comprehension ability of ChatGPT and GPT4 is also attributed to the generalization of LLMs. Meanwhile, many institutions are involved in the research on LLM pretraining and fine-tuning, bringing a series of open-source LLMs, including LLaMA , Vicuna , Baichuan , Yi , Qwen , ChatGLM , InternLM , etc. Apart from the traditional dense causal transformer paradigm, mixture-of-expert (MoE) is also a popular LLM architecture design. Switch Transformer leverages sparse experts to scale up LLMs into trillion parameters, where a router will choose the most appropriate experts from the expert pool based on each input token. Considering only part of the experts will be involved in model training and inference, LLM with MoE can benefit from both large parameter complexity and low computation cost. Mistral 8<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>78<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>7B model outperforms the LLaMA2-70B model on multiple benchmarks, further verifying the effectiveness of the MoE.

2 Multimodal Large Language Models

Recent multimodal large language models (MLLMs) usually leverage the alignment from visual features to the linguistic feature space to achieve superior vision-language understanding capabilities based on off-the-shelf LLMs and vision encoders. CLIP vision encoder , which is trained in contrastive learning from billions of diverse image-text pairs , is widely used among these works. For example, LLaVA adopts an MLP projector to align visual tokens from the frozen CLIP vision encoder to the embedding layer of LLM. However, The representation from CLIP exhibits strong discriminative abilities in classification and recognition but only has limited performance on downstream tasks like location and relation understanding . To break through this bottleneck, some works turn to unlock the CLIP vision encoder and further fine-tune the parameter with training data for downstream tasks. For instance, Qwen-VL collected massive training data for grounding and OCR to jointly optimize the CLIP vision encoder and LLM. Recent works propose to involve an extra frozen vision encoder to enhance the performance of MLLMs. SPHINX is one of the pioneers, where grounding capabilities have been significantly improved with the assistance of the DINOv2 . Vary introduces an extra encoder training on large-scale charts and document data to improve the performance on related downstream tasks.

MoVA Methodology

MoVA comprises five key components: 1) a pre-trained large language model (LLM) that generates accurate responses given the image tokens and instructions; 2) a base vision encoder; 3) vision experts that generate task-specific vision latent features; 4) an auxiliary expert-routing low-rank adaption (LoRA) module that helps LLM select appropriate experts based on images and instructions; 5) mixture-of-vision-expert adapter (MoV-Adapter) that performs fine-grained expert fusion based on the multimodal context.

As illustrated in Fig. 1, MoVA consists of two stages: coarse-grained context-ware expert routing and fine-grained expert fusion with MoV-Adapter. First, our coarse-grained context-ware expert routing leverages the tool-use capabilities of LLM, routing the most appropriate experts from NN expert candidates via LLM to help the model answer the user’s question. We incorporate the expert-routing LoRA module into the LLM to improve the efficiency and effectiveness of expert routing. This expert-routing LoRA module is trained with expert routing annotations and can better align the LLM and the routing task. In the second stage, we turn to enhance the visual representation with a novel MoV-Adapter module in a fine-grained manner. More specifically, we leverage the cross-attention mechanism to extract the task-specific knowledge of representations from chosen experts. Meanwhile, the dynamic gating network in MoV-Adapter can allocate soft weights to the extracted knowledge of each expert according to the input image and instruction. Then the extracted knowledge can be effectively integrated into the foundational representation of the base vision encoder. Finally, the enhanced visual representation with instruction tokens is fed to the LLM to generate an accurate response. In Sec. 3.2 and Sec. 3.3, we will focus on our core contributions, the context-aware expert routing strategy, and the expert fusion with MoV-Adapter. In Sec. 3.4, we will introduce the training process.

The vision encoders in MoVA consist of a base encoder and multiple task-specific vision encoder experts. We choose the pre-trained CLIP ViT-L-336px as the base encoder. Our vision experts include several state-of-the-art task-specific encoders: DINOv2, Co-DETR, SAM, Pix2Struct, Deplot, Vary, and BiomedCLIP. The corresponding expertise is presented in Tab. 1. For example, both Pix2Struct and Vary will be used when the user asks the MLLM to scan the document image. MoVA is flexible and easy to generalize to all decoder-only LLMs. We mainly consider Vicuna-7B and Yi-34B as our language model in this work.

2 Coarse-grained Context-aware Expert Routing

The context-aware expert routing strategy aims to employ the impressive reasoning capacity of LLM to select vision experts with strong relevance to the user’s image and instruction from the expert model pool.

Specifically, we perform the context-aware expert routing in three steps during inference. First, the input image, user questions, and descriptions of expert models are converted into appropriate instructions that prompt the MLLM to perform expert selection. An example of the prompt instruction input and selection output is shown in Tab. 2. Such a routing task does not require high-resolution input images, hence we directly downsample the base encoder’s visual feature to obtain a coarse image embedding (e.g., 6464 image tokens). Considering the difference between routing instructions and conventional multimodal data, avoiding task conflict is necessary. We integrate additional lightweight expert-routing LoRA layers into the LLM for accurate expert routing and further disentangle the routing task from original multimodal tasks (e.g., conversation about natural scenes). The downsampled image tokens and instruction tokens are then fed to the LLM as inputs. Finally, the LLM generates the output text and we parse it to determine which vision expert should be selected for fine-grained knowledge extraction in the second stage. For instance, as depicted in Tab. 2, the LLM directly outputs the option’s letter of DINOv2 and Pix2Struct, thus we only utilize them for the subsequent extraction. During training, we do not perform context-aware expert routing and replace the routing outputs with our routing annotations to improve efficiency.

2.2 Routing Data Construction.

Compared with other MLLMs, MoVA requires additional routing annotations. We first introduce the formal definition of the data structure for an unambiguous understanding of the routing data. The data structure for expert routing introduces additional routing annotation R\mathcal{R} to the conventional multimodal data (I,Q,A)(\mathcal{I},\mathcal{Q},\mathcal{A}). Here, I\mathcal{I} represents the image, Q\mathcal{Q} and A\mathcal{A} refer to the question-answer pair, and R\mathcal{R} refers to the expert set which contains the most appropriate ones to solve this question. Then the construction process for routing data can be formulated as (I,Q,A)→R(\mathcal{I},\mathcal{Q},\mathcal{A})\rightarrow\mathcal{R}, with the primary objective being to derive vision experts that optimally align with the sample (I,Q,A)(\mathcal{I},\mathcal{Q},\mathcal{A}). Intuitively, the language modeling loss can serve as an effective metric for evaluating how a data sample aligns with the vision expert. Specifically, we can reuse the LLaVA-1.5-7B models with various vision encoders presented in Sec. 1 to perform loss computation. Here, we denote the model with the base encoder as M0\mathcal{M}_{0} and the model with jj-th expert among NN experts as Mj\mathcal{M}_{j}. For the ii-th sample (Ii,Qi,Ai)(\mathcal{I}_{i},\mathcal{Q}_{i},\mathcal{A}_{i}), we send it to models {Mj∣j∈{0,1,…,N}}\{\mathcal{M}_{j}|j\in\{0,1,\ldots,N\}\} and calculate the language modeling loss {Li,j∣j∈{0,1,…,N}}\{\mathcal{L}_{i,j}|j\in\{0,1,\ldots,N\}\}. The jj-th expert is regarded as a useful expert for the ii-th sample only if Li,j<Li,0\mathcal{L}_{i,j}<\mathcal{L}_{i,0} and will be added to the routing set Ri\mathcal{R}_{i}. Note that we only keep up to 3 vision experts to avoid the increasing computation costs brought by too many additional experts. All the routing annotations of our training data are generated offline. We can directly parse and input these offline results to the subsequent expert fusion component and LLM during training.

3 Fine-grained Expert Fusion with MoV-Adapter

3.2 Expert Knowledge Extractor.

For the ii-th MoV-Adapter block and the jj-th cross-attention layer, we take input feature Xi\mathbf{X}^{i} as query, and the aligned expert feature F^j\hat{\mathcal{F}}_{j} as the key and value:

3.3 Dynamic Gating Network.

where Pji∈(0,1)\mathbf{P}^{i}_{j}\in(0,1) is the soft weight for the jj-th expert in the ii-th block.

3.4 Transformer Block.

The transformer block in the adapter block follows the vanilla design, consisting of a self-attention layer and an FFN layer. Taking the fused visual representation X^i\hat{\mathbf{X}}^{i}, its output will serve as the input feature Xi+1\mathbf{X}^{i+1} for the next adapter block.

4 Training Paradigm

As depicted in Fig. 3, the training process of MoVA consists of three stages: MoV-Adapter pretraining, supervised finetuning, and expert-routing LoRA training.

To improve multimodal generalization, we first construct 15M visual instruction samples across diverse public datasets for different downstream tasks as the training data:

Image Caption: DataComp-1B Only 4M image-text pairs are randomly selected for the efficiency, ShareGPT4V-PT , and ALLaVA-4V .

Visual Grounding and Localization: Objects365 , RefCOCO , VisualGenome , PointQA , and Flickr30K .

Chart Understanding: MMC-Instruction , Chart2Text , DVQA , and SciGraphQA .

Text Recognition and Document Parsing: LLaVAR-PT and 3M English document images from Common Crawl https://commoncrawl.org.

Biomedical Image Understanding: LLaVA-Med .

Note that for each dataset, we add the annotations of coarse-grained expert routing via the method proposed in Sec. 3.2.2. During the pretraining phase, we only optimize the MoV-Adapter along with the base vision encoder while preserving the capabilities of the initial large language model. Meanwhile, we leverage the routing annotations to choose experts and ignore representations from irrelevant ones during training.

4.2 Supervised Finetuning.

We utilize high-quality visual instruction tuning data that build upon LLaVA-665K for finetuning. Additionally, we integrate several visual question answering datasets across various domains, such as DocVQA , ChartQA , InfographicVQA , AI2D , ST-VQA , TextVQA , SynthDoG-en , Geometry3K , PGPS9K , Geo170K , VQA-RAD , and SLAKE .

We also encompass equivalent comprehensive captions generated by the advanced GPT4-V for improved world knowledge. In the supervised fine-tuning stage, task-specific vision experts are frozen and we jointly optimize the base vision encoder, MoV-Adapter, and LLM. The objective of supervised fine-tuning is to align the enhanced visual representation and the embedding of LLM, boosting its visual instruction-following capabilities. The coarse-grained routing annotations are also directly used in this training phase.

4.3 Expert-routing LoRA Training.

We introduce the expert-routing LoRA layers into the LLM and only train these LoRA layers in the final stage. We use the same instruction tuning data as the second stage for routing task tuning.

Experiments

As mentioned in Sec. 3.4, our training pipeline consists of three stages. In the pretraining stage, we use the AdamW optimizer with an initial learning rate of 2<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>10−42<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>10^{-4}, a batch size of 1024, and train the model for 1 epoch. We jointly finetune the weights of the base vision encoder, MoV-Adapter, and LLM with a batch size of 128 and an initial learning rate of 2<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>10−52<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>10^{-5} during supervised fine-tuning. In the last stage, only the LoRA layers are trained and we keep the same hyperparameter setting as the supervised fine-tuning phase. We use 3 transformer blocks (L<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>=</mo></mrow><annotationencoding="application/x−tex">=</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3669em;"></span><spanclass="mrel">=</span></span></span></span></span>3L<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>=</mo></mrow><annotation encoding="application/x-tex">=</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3669em;"></span><span class="mrel">=</span></span></span></span></span>3) in the MoV-Adapter and its hidden dimension is 1024, which is consistent with the base vision encoder CLIP. The input resolution of the base vision encoder is set as 672×\times672. Two residual blocks with an average pooling are employed in the MoV-Adapter to reduce the number of output image tokens from 2304 to 576. More details about vision experts are released in the Appendix.

For the proxy setting performed in Tab. 1, we follow the default setting of LLaVA-1.5 but incorporate several additional datasets, including DocVQA , ChartQA , RefCOCO referring segmentation data , LLaVA-Med , VQA-RAD , and SLAKE .

2 MLLM Benchmarks

We empirically analyze the multimodal capacity and generalization ability of MoVA on a wide range of challenging MLLM benchmarks in Tab. 3. Specifically, this comprehensive assessment is conducted on MME , MMBench , QBench , MathVista , MathVerse , and POPE .

Compared to other open-source MLLMs with similar model complexity, MoVA with Vicuna-7B achieves the best performance across 7 MLLM benchmarks while offering a more favorable balance between training efficiency and performance. For instance, MoVA-7B surpasses the recent state-of-the-art LLaVA-NeXT-7B with a dynamic high resolution design, processing only 20% image tokens.

Furthermore, we adopt Hermes-Yi-34B as the LLM to validate the scaling property of MoVA. As depicted in Tab. 3, the performance of MoVA-34B is on par with popular proprietary MLLMs (e.g., Gemini-Pro ) and outperforms Qwen-VL-Plus on 5 MLLM benchmarks. For example, MoVA establishes new records on MMBench and MMBench-CN, even surpassing the GPT-4V by a clear margin. These results suggest that the ensemble of vision experts with adaptive expert routing can serve as an effective dimension for MLLM model scaling.

3 Visual Question Answering

The evaluation results on VQA benchmarks are presented in Tab. 4. In this section, we divide these benchmarks into general VQA benchmarks and text-oriented VQA benchmarks .

Thanks to the dynamic and efficient task-specific knowledge extraction, MoVA achieves state-of-the-art performances across diverse VQA benchmarks. For general VQA benchmarks, MoVA-7B outperforms InternVL-Chat equipped with InternViT-6B on VQAv2 and GQA by 4.2% and 1.9%, respectively. Besides, MoVA shows its proficiency in text recognition in various scenarios, including scene text, chart, document, and diagram. For instance, MoVA-7B catches up to the current state-of-the-art generalist CogAgent with 18 billion parameters on these text-oriented benchmarks with smaller model size. The MoVA model with 38B parameters even surpasses the well-established specialist model PALI-X-55B by clear margins. The outstanding performances on distinct VQA benchmarks demonstrate MoVA’s robust generalization capabilities across diverse domains.

4 Visual Grounding

We conduct experiments on Referring Expression Comprehension (REC) benchmarks to evaluate the visual grounding ability of MoVA. The results are presented in Tab. 5. Compared with the previous leading generalist CogVLM with 17B parameters, MoVA-7B attains higher scores on 6 of 8 splits while reducing the model size by 40%. Besides, the performance of MoVA-7B is on par with the state-of-the-art specialist models that are elaborately designed for grounding tasks. For example, MoVA-7B achieves a score of 90.22% on RefCOCO+ val, which is 2.46% higher than the score of UNINEXT-H . Our largest model MoVA-34B further pushes the performance bound of visual grounding on these benchmarks. These impressive results demonstrate MoVA’s remarkable visual grounding capacity.

5 Medical Visual Question Answering

This experiment is conducted on popular medical VQA benchmarks VQA-RAD and SLAKE. We directly leverage the medical VQA evaluation metric adopted by LLaVA-Med. Each sample of VQA-RAD and SLAKE is observed only once during the training process of MoVA and LLaVA-1.5. For a fair comparison, we compare MoVA with the LLaVA-Med variant that is finetuned with only 1 epoch on the benchmark. The performance of the LLaVA-Med specialist that is fully finetuned on downstream tasks is also reported. As presented in Tab. 7, MoVA-7B consistently yields higher scores than LLaVA-Med and LLaVA-1.5 on both medical VQA benchmarks, exhibiting its medical visual chat ability.

6 Image Segmentation

In this experiment, we aim to investigate if task-specific knowledge can improve MoVA on the segmentation task. Therefore, we introduce a simple design to extend MoVA to segmentation tasks. Unlike segmentation generalists that adopt an additional pixel decoder with high-resolution images for high-quality mask generation, we just formulate the referring segmentation task as sequential polygon generation . We finetune MoVA and the baseline with a SAM-Huge backbone on the RefCOCO referring segmentation datasets. MoVA achieves 57.1% gIoU on the testA benchmark, which is 2.6% higher than the 54.5% of baseline. This result indicates that MoVA is capable of exploiting task-specific knowledge to solve segmentation tasks.

7 Ablation Study

As presented in Tab. 7, we perform an ablation to thoroughly delve into the effect of each component. First, we try to replace the context-aware routing with random routing. Without task-relevant vision experts, the performance drops by a large margin, especially on the text-oriented ChartQA and DocVQA benchmarks. Removing context-aware routing to leverage all vision experts also brings similar results. It proves that both these modifications introduce biased information from irrelevant vision experts due to the removal of context-aware routing. Then, we ablate the effectiveness of the MoV-Adapter by replacing it with simple linear layers. The removal of fine-grained expert feature fusion downgrades performance across all datasets. These results delineate that each component in MoVA can consistently yield significant gains.

7.2 Number of activated experts.

In the context-aware routing phase, the number of activated experts KK is dynamic. We compare such a data-dependent design with other variations of constant KK in this experiment. As presented in Tab. 9, the overall performance of dynamic KK consistently outperforms other models with constant KK. This reveals this dynamic implementation can fully exploit the task-specific knowledge of relevant experts while avoiding the incorporation of biased information.

7.3 Adapter Design.

In this section, we conduct ablation studies on the design of the MoV-Adapter. The number of adapter blocks directly affects the model’s complexity and efficiency. As presented in Table 9, we compared the impact of using 2, 3, and 4 adapter blocks on the model’s performance. We observed that the baseline with 3 blocks can achieve better performance than other settings with 2 blocks or 4 blocks. Therefore, we empirically set LL to 3 by default. Then, we substituted our multimodal gating for uniform gating to investigate its effectiveness. Each of the experts is assigned the same soft weight in the uniform gating. We find uniform gating brings consistent performance drops in the test benchmarks. It indicates that the lack of the dynamic soft-weight technique harms the overall performance since it fails to perform precise knowledge extraction.

8 Visualization

In Fig. 4, we present a few samples in three different scenarios: visual grounding, chart understanding, and biomedical multimodal processing. For each example, we provide the question, image, routing information for the expert models, and their corresponding weights. The coarse-grained routing effectively identifies relevant experts for these cases. The fine-grained routing also accurately assigns specific weights to each identified expert, ensuring precise expert allocation.

Conclusion

In this paper, we reveal that the inherent bias of each vision encoder can diminish its generalization ability across other irrelevant domains by analyzing the performance of individual vision encoders versus the plain fusion of multiple encoders across various tasks. To deal with the problem, we propose MoVA, a powerful MLLM composed of coarse-grained context-aware expert routing and fine-grained expert fusion with MoV-Adapter. Based on multimodal context and model expertise, MoVA fully leverages representation from multiple context-relevant vision encoder experts flexibly and effectively while avoiding biased information brought by irrelevant experts. We demonstrate the effectiveness of each component in MoVA by elaborate ablation studies. Without any bells and whistles, MoVA can achieve significant performance gains over current state-of-the-art methods in a wide range of challenging benchmarks.

References