A Survey on Multimodal Large Language Models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, Enhong Chen
Introduction
Recent years have seen the remarkable progress of large language models . By scaling up data size and model size, these LLMs raise amazing emergent abilities, typically including In-Context Learning (ICL) , instruction following , and Chain of Thought (CoT) . Although LLMs have demonstrated surprising zero/few-shot reasoning performance on most Natural Language Processing (NLP) tasks, they are inherently “blind” to vision since they can only understand discrete text. Concurrently, large vision foundation models make rapid progress in perception , and the traditional combination with text pays more attention to modality alignment and task unity , developing slowly in reasoning.
In light of this complementarity, unimodal LLMs and vision models run towards each other at the same time, ultimately leading to the new field of MLLM. Formally, it refers to the LLM-based model with the ability to receive and reason with multimodal information. From the perspective of developing Artificial General Intelligence (AGI), MLLM may take a step forward from LLM for the following reasons: (1) MLLM is more in line with the way humans perceive the world. Our humans naturally receive multisensory inputs that are often complementary and cooperative. Therefore, multimodal information is expected to make MLLM more intelligent. (2) MLLM offers a more user-friendly interface. Thanks to the support of multimodal input, users can interact and communicate with the intelligent assistant in a more flexible way. (3) MLLM is a more well-rounded task-solvers. While LLMs can typically perform NLP tasks, MLLMs can generally support a larger spectrum of tasks.
GPT-4 ignites a research frenzy over MLLM because of the amazing examples it shows. However, GPT-4 does not open the multimodal interface, and no information about the model has been made public up until now. In spite of this, many efforts have been made by the research community to develop capable and open-sourced MLLMs, and some surprising practical capabilities have been exhibited, such as writing website codes based on images , understanding the deep meaning of a meme , and OCR-free math reasoning . We write this survey to provide researchers with a grasp of the basic idea, main method, and current progress of MLLMs. Note that we mainly focus on visual and language modalities, but also include works involving other modalities. Specifically, we divide the existing MLLMs into four types with corresponding summarizations and, meanwhile, open a GitHub page that would be updated in real-time. To the best of our knowledge, this is the first survey on MLLM.
Overview
This paper categorizes recent representative MLLMs into four main genres: Multimodal Instruction Tuning (M-IT), Multimodal In-Context Learning (M-ICL), Multimodal Chain-of-Thought (M-CoT), and LLM-Aided Visual Reasoning (LAVR). The first three constitute the fundamentals of MLLMs, while the last one is a multimodal system with LLM as the core. Note that the three techniques are relatively independent and can be utilized in combination. Therefore, our illustration of a concept may also involve others.
We organize the survey according to the four main categories and introduce them sequentially. We start with a detailed introduction of M-IT (§3.1) to reveal how LLMs can be adapted for multimodality in terms of two aspects: architecture and data. Then we introduce M-ICL (§3.2), an effective technique commonly used at the inference stage to boost few-shot performance. Another important technique is the M-CoT (§3.3), which is typically used in complex reasoning tasks. Afterward, we further summarize several roles that LLMs mainly take in LAVR (§3.4), which frequently involves the three techniques. Finally, we finish our survey with a summary and potential research directions.
Method
Instruction refers to the description of tasks. Instruction tuning is a technique that involves finetuning pre-trained LLMs on a collection of instruction-formatted datasets . Tuning in this way, LLMs can generalize to unseen tasks by following new instructions, thus boosting zero-shot performance. This simple yet effective idea has sparked the success of subsequent works in the realm of NLP, such as ChatGPT , InstructGPT , FLAN , and OPT-IML .
The comparisons between instruction tuning and related typical learning paradigms are illustrated in Fig. 1. The supervised finetuning approach usually requires many task-specific data to train a task-specific model. The prompting approach reduces the reliance on large-scale data and can fulfill a specialized task via prompt engineering. In such a case, though the few-shot performance has been improved, the zero-shot performance is still quite average . Differently, instruction tuning learns how to generalize to unseen tasks, rather than fitting specific tasks like the two counterparts. Moreover, instruction tuning is highly related to multi-task prompting .
Contrastively, traditional multimodal models are still confined to the first two tuning paradigms, lacking the zero-shot ability. Therefore, many recent works have explored extending the success of instruction tuning in LLMs to multimodality. In order to extend from unimodality to multimodality, the corresponding adaptations are necessary for both the data and the model. For the data, researchers usually acquire M-IT datasets by adapting existing benchmark datasets or by self-instruction . Regarding the model, a common approach is to inject the information of foreign modalities into LLMs and treat them as strong reasoners. Relevant works either directly align foreign embeddings to the LLMs or resort to expert models to translate foreign modalities into natural languages that LLMs can ingest . Formulated in this way, these works transform LLMs into multimodal chatbots and multimodal universal task solvers through multimodal instruction tuning.
In the following parts of this section, we first offer the foundational knowledge (§3.1.2). Before transitioning to the delineation of M-IT, we additionally introduce a common process prior to M-IT, i.e., alignment pre-training (§3.1.3). Then we structure the remaining content as illustrated in Fig. 2: We first introduce how the M-IT data are collected (§3.1.4), followed by a detailed discussion of the model adaption for MLLMs, i.e., various ways of bridging the gap between different modalities (§3.1.5). Finally, we introduce the evaluation methods to assess instruction-tuned MLLMs (§3.1.6).
1.2 Preliminaries
This section briefly illustrates the general structure of multimodal instruction samples and the common process of M-IT.
A multimodal instruction sample often includes an instruction and an input-output pair. The instruction is typically a natural language sentence describing the task, such as, “Describe the image in detail.” The input can be an image-text pair like the Visual Question-Answering (VQA) task or only an image like the image captioning task . The output is the answer to the instruction conditioned on the input. The instruction template is flexible and subject to manual designs , as exemplified in Table 1. Note that the instruction samples can also be generalized to multi-round instructions, where the multimodal inputs are shared .
Formally, a multimodal instruction sample can be denoted in a triplet form, i.e., , where represent the instruction, the multimodal input, and the ground truth response, respectively. The MLLM predicts an answer given the instruction and the multimodal input:
Here, denotes the predicted answer, and are the parameters of the model. The training objective is typically the original auto-regressive objective used to train the LLMs , based on which the MLLM is forced to predict the next token of the response. The objective can be expressed as:
where is the length of the ground-truth response.
1.3 Modality Alignment
It is common to perform large-scale (compared to instruction-tuning) pre-training on paired data to encourage alignment between different modalities , which is prior to the M-IT. The alignment datasets are typically image-text pairs or Automatic Speech Recognition (ASR) datasets, which all contain text. More specifically, the image-text pairs describe images in the form of natural language sentences, while the ASR datasets comprise transcriptions of speech. A common approach for alignment pre-training is to keep pre-trained modules (e.g. visual encoders and LLMs) frozen and train a learnable interface , which is illustrated in the following section.
1.4 Data
The collection of multimodal instruction-following data is a key to M-IT. The collection methods can be broadly categorized into benchmark adaptation, self-instruction , and hybrid composition. We illustrate these three methods sequentially.
Benchmark datasets are rich sources of high-quality data. Hence, abundant works have utilized existing benchmark datasets to construct instruction-formatted datasets. Take the transformation of VQA datasets for an example, the original sample is an input-out pair where the input comprises an image and a natural language question, and the output is the textual answer to the question conditioned on the image. The input-output pairs of these datasets could naturally comprise the multimodal input and response of the instruction sample (see §3.1.2). The instructions, i.e., the descriptions of the tasks, can either derive from manual design or from semi-automatic generation aided by GPT. Specifically, some works hand-craft a pool of candidate instructions and sample one of them during training. We offer an example of instruction templates for the VQA datasets as shown in Table 2. The other works manually design some seed instructions and use these instructions to prompt GPT to generate more .
Note that since the answers of existing VQA and caption datasets are usually concise, directly using these datasets for instruction tuning may limit the output length of MLLM. There are two common strategies to tackle this problem. The first one is to modify instructions. For example, ChatBridge explicitly declares short and brief for short-answer data, as well as a sentence and single sentence for caption data. Similarly, InstructBLIP inserts short and briefly into instruction templates for public datasets that inherently prefer short responses. The second one is to extend the length of existing answers . For example, M3IT proposes to rephrase the original answer by prompting ChatGPT with the original question, answer, and context.
Although existing benchmark datasets can contribute a rich source of data, they usually do not well meet human needs in real-world scenarios, such as multiple rounds of conversations. To tackle this issue, some works collect samples through self-instruction , which bootstraps LLMs to generate textual instruction-following data using a few hand-annotated samples. Specifically, some instruction-following samples are hand-crafted as seed examples, after which ChatGPT/GPT-4 is prompted to generate more instruction samples with the seed samples as guidance. LLaVA extends the approach to the multimodal field by translating images into texts of captions and bounding boxes, and prompting GPT-4 to generate new data in the context of seed examples. In this way, an M-IT dataset is constructed, called LLaVA-Instruct-150k. Following this idea, subsequent works such as MiniGPT-4 , ChatBridge , GPT4Tools , and DetGPT develop different M-IT datasets catering for different needs.
Apart from the M-IT data, language-only user-assistant conversation data can also be used to improve conversational proficiencies and instruction-following abilities . LaVIN directly constructs a minibatch by randomly sampling from both language-only and M-IT data. MultiInstruct probes different strategies for training with a fusion of single modal and multimodal data, including mixed instruction tuning (combine both types of data and randomly shuffle), sequential instruction tuning (text data followed by multimodal data), and Adapter-based sequential instruction tuning. The empirical results show that mixed instruction tuning is at least not worse than solely tuning on multimodal data.
1.5 Modality Bridging
Since LLMs can only perceive text, bridging the gap between natural language and other modalities is necessary. However, it would be costly to train a large multimodal model in an end-to-end manner. Moreover, doing so would take the risk of catastrophic forgetting . Thus, a more practical way is to introduce a learnable interface between the pre-trained visual encoder and LLM. The other approach is to translate images into languages with the help of expert models, and then send the language to LLM.
The learnable interface is responsible for connecting different modalities when freezing the parameters of the pre-trained models. The challenge lies in how to efficiently translate visual content into text that LLM can understand. A common and feasible solution is to leverage a group of learnable query tokens to extract information in a query-based manner , which first has been implemented in Flamingo and BLIP-2 , and subsequently inherited by a variety of work . Furthermore, some methods use a projection-based interface to close the modality gap . For example, LLavA adopts a simple linear layer to embed image features and MedVInT-TE uses a two-layer multilayer perceptron as a bridge. There are also works that explore a parameter-efficient tuning manner. LLaMA-Adapter introduces a lightweight adapter module in Transformer during training. LaVIN designs a mixture-of-modality adapter to dynamically decide the weights of multimodal embeddings.
Apart from the learnable interface, using expert models, such as an image captioning model, is also a feasible way to bridge the modality gap . Differently, the idea behind the expert models is to convert multimodal inputs into languages without training. In this way, LLMs can understand multimodality by the converted languages indirectly. For example, VideoChat-Text uses pre-trained vision models to extract visual information such as actions and enriches the descriptions using a speech recognition model. Though using expert models is straightforward, it may not be as flexible as adopting a learnable interface. The conversion of foreign modalities into text would typically cause information loss. As VideoChat-Text points out, transforming videos into textual descriptions distorts spatial-temporal relationships.
1.6 Evaluation
There are various metrics to evaluate the performance of the model after M-IT, which can be broadly categorized into two types according to the question genres, including closed-set and open-set.
Closed-set questions refer to a type of questions where the possible answer options are predefined and limited to a finite set. The evaluation is usually performed on benchmark-adapted datasets. In this case, the responses can be naturally judged by benchmark metrics . For example, InstructBLIP reports the accuracy on ScienceQA , as well as the CIDEr score on NoCaps and Flickr30K . The evaluation settings are typically zero-shot or finetuning . The first setting often selects a wide range of datasets covering different general tasks and splits them into held-in and held-out datasets. After tuning on the former, zero-shot performance is evaluated on the latter with unseen datasets or even unseen tasks. In contrast, the second setting is often observed in the evaluation of domain-specific downstream tasks. For example, LLaVA and LLaMA-Adapter report finetuned performance on ScienceQA . LLaVA-Med reports results on biomedical VQA .
The above evaluation methods are usually limited to a small range of selected tasks or datasets, lacking a comprehensive quantitative comparison. To this end, some efforts have endeavored to develop new benchmarks specially designed for MLLMs . For example, Fu et al. construct a comprehensive evaluation benchmark MME that includes a total of 14 perception and cognition tasks. All instruction-answer pairs in MME are manually designed to avoid data leakage. 10 advanced MLLMs are evaluated with detailed leaderboards and analyses. LAMM-Benchmark is proposed to evaluate MLLMs quantitatively on a variety of 2D/3D vision tasks. Video-ChatGPT proposes a quantitative evaluation framework for video-based conversational models, which incorporates two types of assessments, i.e., evaluation of video-based generative performance and zero-shot question-answering.
In contrast to the closed-set questions, the responses to open-set questions can be more flexible, where MLLMs usually play a chatbot role. Because the content of the chat can be arbitrary, it would be trickier to judge than the closed-ended output. The criterion can be classified into manual scoring, GPT scoring, and case study. Manual scoring requires humans to assess the generated responses. This kind of approach often involves hand-crafted questions that are designed to assess specific dimensions. For example, mPLUG-Owl collects a visually related evaluation set to judge capabilities like natural image understanding, diagram and flowchart understanding. Similarly, GPT4Tools builds two sets for the finetuning and zero-shot performance respectively, and evaluates the responses in terms of thought, action, arguments, and the whole.
Since manual assessment is labor intensive, some researchers have explored rating with GPT, namely GPT scoring. This approach is often used to evaluate performance on multimodal dialogue. LLaVA proposes to score the responses via GPT-4 in terms of different aspects, such as helpfulness and accuracy. Specifically, 30 images are sampled from the COCO validation set, each associated with a short question, a detailed question, and a complex reasoning question via self-instruction on GPT-4. The answers generated by both MLLM and GPT-4 are sent to GPT-4 for comparison. Subsequent works follow this idea and prompt ChatGPT or GPT-4 to rate results or judge which one is better .
A main issue of GPT-4 based scoring is that currently, its multimodal interface is not publicly available. As a result, GPT-4 can only generate responses based on image-related text content, such as captions or bounding box coordinates, without accessing the image . It thus may be questionable to set GPT-4 as the performance upper bound in this case. An alternative approach is to compare different capabilities of MLLMs through case studies. For example, mPLUG-Owl uses a visually related joke understanding case to compare against GPT-4 and MM-REACT . Similarly, Video-LLaMA offers some cases to demonstrate several capabilities, such as audio-visual co-perception and common-knowledge concept recognition.
Some other methods focus on a specific aspect of MLLMs. For instance, MultiInstruct proposes a metric called sensitivity that assesses the model’ robustness to varied instructions. Li et al. delve into the object hallucination problem and propose a query method POPE to assess performance in this regard. Zhao et al. consider safety issues and propose to evaluate the robustness of MLLMs to adversarial attacks.
2 Multimodal In-Context Learning
ICL is one of the important emergent abilities of LLMs. There are two good traits of ICL: (1) Different from traditional supervised learning paradigms that learn implicit patterns from abundant data, the crux of ICL is to learn from analogy . Specifically, in the ICL setting, LLMs learn from a few examples along with an optional instruction and extrapolate to new questions, thereby solving complex and unseen tasks in a few-shot manner . (2) ICL is usually implemented in a training-free manner and thus can be flexibly integrated into different frameworks at the inference stage. A closely related technique to ICL is instruction-tuning (see §3.1), which is shown empirically to enhance the ICL ability .
In the context of MLLM, ICL has been extended to more modalities, leading to Multimodal ICL (M-ICL). Building upon the setting in (§3.1.2), at inference time, M-ICL can be implemented by adding a demonstration set, i.e., a set of in-context samples, to the original sample. In this case, the template can be extended as illustrated in Table 3. Note that we list two in-context examples for illustration, but the number and the ordering of examples can be flexibly adjusted. In fact, models are commonly sensitive to the arrangement of demonstrations .
In terms of applications in multimodality, M-ICL is mainly used in two scenarios: (1) solving various visual reasoning tasks and (2) teaching LLMs to use external tools . The former usually involves learning from a few task-specific examples and generalizing to a new but similar question. From the information provided in instructions and demonstrations, LLMs get a sense of what the task is doing and what the output template is, and finally generate expected answers. In contrast, examples of tool usage are often text-only and more fine-grained. They typically comprise a chain of steps that could be sequentially executed to fulfill the task. Thus, the second scenario is closely related to CoT (see §3.3).
3 Multimodal Chain of Thought
As the pioneer work points out, CoT is “a series of intermediate reasoning steps”, which has been proven to be effective in complex reasoning tasks . The main idea of CoT is to prompt LLMs to output not only the final answer but also the reasoning process that leads to the answer, resembling the cognitive process of humans.
Inspired by the success in NLP, multiple works have been proposed to extend the unimodal CoT to Multimodal CoT (M-CoT). We summarize these works as shown in Fig. 3. To begin with, similar to the situation in M-IT (see §3.1), the modality gap needs to be filled (§3.3.1). Then, we introduce different paradigms for acquiring the ability of M-CoT (§3.3.2). Finally, we delineate more specific aspects of M-CoT, including the configuration (§3.3.3) and the formulation of chains (§3.3.4).
To transfer the success from NLP to multimodal, modality bridging is the first issue to address. There are broadly two ways to achieve this: through the fusion of features or through transforming visual input into textual descriptions. Similar to the case in §3.1.5, we classify them as learnable interface and expert model respectively and discuss them in sequence.
This approach involves adopting a learnable interface to map visual embedding to the word embedding space. The mapped embeddings can then be taken as a prompt, which is sent to LLMs with other languages to elicit M-CoT reasoning. For example, CoT-PT chains multiple Meta-Nets for prompt tuning to simulate a reasoning chain, where each Meta-Net embeds visual features into a step-specific bias to the prompt. Multimodal-CoT adopts a two-stage framework with a shared Transformer-based structure , where visual and textual features interact through cross-attention.
Introducing an expert model to translate visual input to textual descriptions is an alternative modality bridging way. For example, ScienceQA adopts an image captioning model and feeds the concatenation of the image captions and original language input to LLMs. Though simple and straightforward, this approach may suffer from information loss in the captioning process .
3.2 Learning Paradigms
The learning paradigm is also an aspect worth investigating. There are broadly three ways to acquire the M-CoT ability, i.e., through finetuning and training-free few/zero-shot learning. The sample size requirement for the three ways is in descending order.
Intuitively, the finetuning approach often involves curating specific datasets for M-CoT learning. For example, ScienceQA constructs a scientific question-answering dataset with lectures and explanations, which can serve as sources of learning CoT reasoning, and finetunes on this proposed dataset. Multimodal-CoT also uses the ScienceQA benchmark but generates the output in a two-step fashion, i.e., the rationale (chain of reasoning steps) and the final answer based on the rationale. CoT-PT learns an implicit chain of reasoning through a combination of prompt tuning and step-specific visual bias.
Compared with finetuning, few/zero-shot learning is more computationally efficient. The main difference between them is that the few-shot learning typically requires hand-crafting some in-context examples so that the model can learn to reason step by step more easily. In contrast, the zero-shot learning does not require any specific example for CoT learning. In this case, by prompting designed instructions like “Let’s think frame by frame” or “What happened between these two keyframes” , models learn to leverage the embedded knowledge and the reasoning ability without explicit guidance. Similarly, some works prompt models with descriptions of the task and tool usage to decompose complex tasks into sub-tasks.
3.3 Chain Configuration
Chain configuration is an important aspect of reasoning and can be categorized into adaptive and pre-defined formations. The former configuration requires LLMs to decide on their own when to halt the reasoning chains , while the latter setting stops the chains with a pre-defined length .
3.4 Generation Patterns
How the chain is constructed is a question worth studying. We summarize the current works into (1) an infilling-based pattern and (2) a predicting-based pattern. Specifically, the infilling-based pattern demands deducing steps between surrounding context (previous and following steps) to fill the logical gaps . In contrast, the predicting-based pattern requires extending the reasoning chains given conditions such as instructions and previous reasoning history . The two types of patterns share a requirement that the generated steps should be consistent and correct.
4 LLM-Aided Visual Reasoning
Inspired by the success of tool-augmented LLMs , some researches have explored the possibilities of invoking external tools or vision foundation models for visual reasoning tasks. Taking LLMs as helpers with different roles, these works build task-specific or general-purpose visual reasoning systems.
Compared with conventional visual reasoning models , these works manifest several good traits: (1) Strong generalization abilities. Equipped with rich open-world knowledge learned from large-scale pretraining, these systems can easily generalize to unseen objects or concepts with remarkable zero/few-shot performance . (2) Emergent abilities. Aided by strong reasoning abilities and abundant knowledge of LLMs, these systems are able to perform complex tasks. For example, given an image, MM-REACT can interpret the meaning beneath the surface, such as explaining why a meme is funny. (3) Better interactivity and control. Traditional models typically allow a limited set of control mechanisms and often entail expensive curated datasets . In contrast, LLM-based systems have the ability to make fine control in a user-friendly interface (e.g. click and natural language queries) .
The following parts of this section are organized as displayed in Fig. 4: we start with introducing different training paradigms employed in the construction of LLM-Aided Visual Reasoning systems (§3.4.2). Subsequently, we delve into the primary roles that LLMs play within these systems (§3.4.3). Finally, we wrap up our discussion with various types of performance evaluation.
4.2 Training Paradigms
According to training paradigms, LLM-Aided Visual Reasoning systems can be divided into two types, i.e., training-free and finetuning.
With abundant prior knowledge stored in pre-trained LLMs, an intuitive and simple way is to freeze pre-trained models and directly prompt LLMs to fulfill various needs. According to the setting, the reasoning systems can be further categorized into few-shot models and zero-shot models. The few-shot models entail a few hand-crafted in-context samples (see §3.2) to guide LLMs to generate a program or a sequence of execution steps. These programs or execution steps serve as instructions for corresponding foundation models or external tools/modules. The zero-shot models take a step further by directly utilizing LLMs’ linguistics/semantics knowledge or reasoning abilities. For example, PointCLIP V2 prompts GPT-3 to generate descriptions with 3D-related semantics for better alignment with corresponding images. In CAT , LLMs are instructed to refine the captions according to user queries.
To activate the planning abilities with respect to tool usage and to improve the instruction-following abilities of the system, GPT4Tools introduces the instruction-tuning approach (see §3.1). A new tool-related instruction dataset is collected and used to finetune the model.
4.3 Functions
In order to further inspect what roles LLMs exactly play in LLM-Aided Visual Reasoning systems, existing related works are divided into three types:
The first two roles, i.e., the controller and the decision maker, are related to CoT (see §3.3). It is frequently used because complex tasks need to be broken down into intermediate simpler steps. When LLMs act as a controller, the systems often finish the task in a single round, while multi-round is more common in the case of the decision maker. We delineate how LLMs serve these roles in the following parts.
In this case, LLMs act as a central controller that (1) breaks down a complex task into simpler sub-tasks/steps and (2) assigns these tasks to appropriate tools/modules. The first step is often finished by leveraging the CoT ability of LLMs. Specifically, LLMs are prompted explicitly to output task planning or, more directly, the modules to call . For example, VISPROG prompts GPT-3 to output a visual program, where each program line invokes a module to perform a sub-task. In addition, LLMs are required to output argument names for the module input. To handle these complex requirements, some hand-crafted in-context (see §3.1) examples are used as references . This is closely related to the optimization of reasoning chains (see §3.3), or more specifically, the least-to-most prompting technique. In this way, complex problems are broken down into sub-problems that are solved sequentially.
In this case, complex tasks are solved in a multi-round manner, often in an iterative way . Decision Makers often fulfill the following responsibilities: (1) Summarize the current context and the history information, and decide if the information available at the current step is sufficient to answer the question or complete the task; (2) Organize and summarize the answer to present it in a user-friendly way.
When LLM is used as a Semantics Refiner, researchers mainly utilize their rich linguistics and semantics knowledge. Specifically, LLMs are often instructed to integrate information into consistent and fluent natural language sentences or generate texts according to different specific needs .
4.4 Evaluation
There are two ways to evaluate the performance of LLM-Aided Visual Reasoning systems, namely benchmark metrics and manual assessment .
A straightforward evaluation way is to test the system on existing benchmark datasets since the metrics can directly reflect how well the model finishes the task. For example, Chameleon is evaluated on complex reasoning benchmarks, including ScienceQA and TabMWP . IdealGPT reports the accuracy on VCR and SNLI-VE .
Some works adopt manual ratings to evaluate specific aspects of models. For example, ChatCaptioner asks human annotators to judge the richness and correctness of captions generated by different models. GPT4Tools calculates successful rates of thought, action, argument, and the overall successful rate to measure the model’s capability in assigning tool usage. VISPROG manually calculates the accuracy when assessing the model on language-guided image editing tasks.
Challenges and Future Directions
The development of MLLMs is still in a rudimentary stage and thus leaves much room for improvement, which we summarize below:
Current MLLMs are still limited in perception capabilities, leading to incomplete or wrong visual information acquisition . This may be due to the compromise between information capacity and computation burden. More specifically, Q-Former only uses 32 learnable tokens to represent an image, which might induce information loss. Nonetheless, scaling up the token size would inevitably bring a larger computation burden to LLMs, whose input length is usually limited. A potential method is to introduce large vision foundation models like SAM to compress visual information more efficiently .
The reasoning chain of MLLMs may be fragile. For example, Fu et al. find that in a math calculation case, although MLLM calculates the correct result, it still delivers a wrong answer due to the broken of reasoning. This indicates that the reasoning ability of a unimodal LLM may not be equal to that of the LLM after receiving visual information. The topic of improving multimodal reasoning is worth investigating.
The instruction-following ability of MLLMs needs upgrading. After M-IT, some MLLMs fail to generate the expected answer (“yes” or “no”) despite an explicit instruction, “Please answer yes or no” . This suggests that instruction tuning may need to cover more tasks to improve generalization.
The object hallucination issue is widespread , which largely affects the reliability of MLLMs. This may be ascribed to insufficient alignment pre-training . Thus, a possible solution is to perform a more fine-grained alignment between visual and textual modalities. The fine granularity refers to the local features of images, which can be obtained by SAM , and the corresponding local textual descriptions.
Parameter-efficient training is needed. Both the existing two modality bridging manners i.e., the learnable interface and the expert model, are preliminary explorations to reduce the computation burden. More efficient training methods may unlock more power in MLLMs with limited computational resources.
Conclusion
In this paper, we perform a survey of the existing MLLM literature and offer a broad view of its main directions, including three common techniques (M-IT, M-ICL, and M-CoT) and a general framework to build task-solving systems (LAVR). Moreover, we underscore the current research gaps to be filled and point out some promising research directions. We hope this survey can offer readers a clear picture of the current progress of MLLM and inspire more work.