What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong
Introduction
Large Language Models (LLMs) have progressed rapidly in recent years and achieved impressive performance in language understanding and generalization. With instruction fine-tuning , LLMs can be further improved to follow open-ended instructions from non-expert users and serve as dialog-based assistants in our daily lives. Leveraging powerful LLMs, recent studies have examined methods for adapting LLMs to multimodal inputs (e.g., images [18, 19, 20, liu2023visual, 21, 22, 23, 24], videos , and audio ) and outputs (e.g., vision tasks , and robotic manipulation skills ). Notably, GPT4 has astounded the world with its impressively stable zero-shot versatile yet practical capabilities, such as generating descriptions, stories, poetry, advertisements, and codes given images, which were rarely observed in previous vision language models .
However, it still remains a mystery that: How does GPT4 obtain its impressive smartness? Though actively investigated recently, the existing models are usually different in network structure, training data, training recipes, prompts, and evaluation benchmarks, which makes it extremely hard to tell which factors are crucial in achieving a high-performance multi-modal language model. In addition, suitable quantitative benchmarks for evaluating and comparing such models are lacking, making it difficult to attribute and quantify the progress in open-sourced multi-modal LLMs.
Therefore, in this paper, we conduct a systematic study on training GPT4-style models to address the aforementioned issues. According to the existing literature, we identify three possible keys to achieving high performance for multi-modal LLMs: network structures, training data, and diversified instructions. Regarding network structures, we explore different LLM adaptation strategies, including the widely utilized cross-attention-based structure and the recently popular decoder-only structure with a multi-modal adapter [liu2023visual, 23, 22]. Besides, we investigate different backbones including LLaMA-7B and Vicuna-7B to assess whether language instruction fine-tuning affects the final multi-modal performance. As for training data, we experiment with several large-scale datasets (e.g. COYO700M , DataComp1B , and BlipCapFilt ) consisting of image-text pairs to observe the effects of different data combinations. For instructions, we manually label at least three prompts for each task and generate more with GPT4 to figure out the influence of the diversity of language prompts. In total, there are 500 prompts for over 50 tasks. In summary, we implement 20 variants with controlled settings and conduct extensive experiments to draw reliable conclusions both quantitatively and qualitatively.
For benchmarking, we argue that the evaluation of multi-modal LLMs is essentially different from typical visual-language methods. The primary challenge when evaluating a GPT4-style model is balancing text generation capability and multi-modal understanding accuracy. To address this, we present a new benchmark incorporating both video and image data to evaluate both the multi-modal understanding and text generation performances. Using our proposed benchmark, we evaluate a large bunch of open-source methods and provide a comprehensive review. Concretely, we adopt two protocols for quantitative evaluation. First, we collect an Open-ended Visual Question Answering (Open-VQA) test set, including questions on objects, OCR, counting, reasoning, action recognition, chronological ordering, and more. Different from standard VQA , the ground-truth answer in Open-VQA is open-ended. To evaluate the performance on Open-VQA, we prompt GPT4 to make it a discriminator, yielding a 95% agreement with human evaluation. This benchmark is used to evaluate the accuracy of all models. Additionally, we adopt the OwlEval test set proposed by mPLUG-owl to assess the text generation ability given images. Though OwlEval is a tiny set containing only 82 questions based on 50 images, it covers a diverse range of tasks such as generating descriptions, stories, poems, advertisements, codes, and other sophisticated yet practical analyses of given images. In this part, we recruit human annotators to rank different models.
Based on extensive analysis of our controlled experiments, our findings can be summarized as follows:
Prefix-tuning with trainable adaptors has shown better performances to adapt LLMs to multi-modal inputs compared to cross attention (e.g. Flamingo ).
Data quality is more important than quantity. We find that models trained on large-scale image text pairs like COYO700M and DataComp1B are not better to generate languages than models trained on a much smaller but high-quality dataset, since they can contaminate the output distribution.
Diversified prompts are crucial to the improvement of the instruction-following ability and, consequently, final performance.
For the multi-modal adaptation of LLMs, it is crucial to carefully balance the multi-modal understanding and text generation abilities. Multi-modal adaptation based on instruction-finetuned models like Vicuna can improve the instruction-following abilities.
Through our study, we present Lynx, a simple prefix-tuning GPT4-style model, with a two-stage training recipe. For the first stage, we use 120M image-text pairs to align visual and linguistic embeddings. For the second stage, we finetune our model with 20 multi-modal tasks with image or video inputs and NLP instruction data to learn to follow instructions. We transform all multi-modal datasets into the instruction-following format with manually written prompts and more GPT4-generated ones to keep the consistency of all training data. The resulting model performs the most accurate multi-modal understanding while exhibiting the best multi-modal generation ability compared to existing open-sourced models.
Lynx
Lynx is a GPT4-style large language model that can take images and videos as inputs. Built on top of Vicuna, it is further trained with additional trainable adapters on high-quality image-text pairs and visual language tasks. In this section, we will introduce our Lynx in detail, including the problem formulation (2.1), architecture (2.2), pretraining (2.3), and instruction finetuning (2.4).
A GPT4-style large language model is defined as a decoder-only transformer that takes both visual and instructional tokens as inputs and generates responses in text auto-regressively. Formally, the input includes vision tokens and instruction tokens , where and represent the number of vision tokens and instruction tokens. The vision tokens and instruction tokens in our model are directly concatenated to form the input of the decoder-only model. Conditioned on the multi-modal inputs, the model predicts the response in an auto-regressive manner, i.e., each word is predicted conditioned on all input tokens and previous predictions. Therefore, the sentence is predicted by the following equation:
In large language models , the network is usually trained on numerous text corpus to learn the causal relationships among tokens. Similarly, our model is also trained on the collected visual-language tasks to learn the next-word distribution. Notably, compared to the contrastive pretraining , pretraining with next-word prediction requires data with fluent texts that can represent the “natural” causal dependency between the predicted word and the past context very well . We will introduce the details of data collection and selection in Section 2.3 and 2.4 in detail.
2 Details of Model Architecture
Our model takes simultaneously vision and language as inputs to generate text responses following the input instructions. The overall structure of our model is shown in Fig.1. Concretely, vision inputs are first processed by a vision encoder to get a sequence of vision tokens . After that, are fused with instruction tokens for multi-modal tasks. In our model, we directly concatenate the projected vision tokens and instruction tokens as the input of LLMs, which can then be processed by the decoder-only LLMs naturally. We call this structure “prefix-finetuning” (PT) in contrast to the cross-attention-based models like Flamingo . Moreover, we find that by adding a small trainable adapter after some layers in the frozen LLMs, the performance could be further improved with low training costs. To generate responses, the left-to-right causal decoder auto-regressively predicts the next token by taking all previous tokens as inputs until encountering the
Adapter
The trainable adapters are inserted into the LLMs after every blocks. In our experiments, . As shown in Figure 2(b), the adapter linearly projects each token into a lower-dimensional space and then re-projects it back. Concretely, in Lynx, the hidden state for each token is 4096-d. The adapter first imposes layer normalization onto the hidden states. Then a linear layer is used to downsample the dimension of each token state from 4096 to 2048, based on which SiLU is set as the non-linear activation function, which keeps consistent with LLaMA . Finally, the other linear layer is used to re-map the 2048-d hidden state back to 4096-d.
Vision Encoder
To extract vision features of images and video frames, we apply EVA-1B as our vision encoder . It maps an image to a sequence of visual tokens. The downsample rate is 14, meaning that an image with resolution will be represented by a sequence of tokens. To improve the efficiency of training and inference, we adapt the resampler mechanism that reduces the dimensions of vision inputs by injecting the long vision token sequence into a short and learnable query sequence :
where is the input image, is the raw tokens directly given by the vision encoder, is the condensed token sequence consisting of 32 tokens regardless of the number of raw tokens from the vision encoder.
3 Pretraining
During pretraining, we utilize more than 120M image-text pairs to train the newly added layers so as to build connections of different modalities. Our pretraining follows the typical next-word prediction training with the cross entropy loss. To accelerate pretraining, we first pre-train our model on images of 224224 resolution. Nevertheless, we found that only pretraining on a low resolution is not enough for some downstream tasks like table reading and OCR. Therefore, after 100k steps of pretraining on low-res images, we continue to increase the input resolution to 420420 and train the model for another 10k steps.
Training data during this phase mainly consists of BlipCapFilt 115M , CC12M , CC3M , and SBU . Besides, we also add high-quality labeled data during pretraining that have been also used in the instruction finetuning phase, like captioning, visual question answering, and classification. Details of all pretraining datasets are listed in Table 9. Our model is trained on a total of 14B tokens from all these datasets during the pretraining stage and 3B tokens during the instruction-finetuning stage.
4 Instruction Fintuning
To finetune our model with diversified instructions, we collect an instruction finetuning multi-modal dataset based on the public ones. Our dataset consists of 50+ text-only, image-text, and video-text tasks mainly belonging to 5 categories: Text-only Instruction-Following, Image/Video Visual Question Answering, Image/Video Captioning, Classification, and Image-conditioned Dialog for Complex Reasoning and Instruction Following. We also provide the corresponding instructions for all of these tasks (see Appendix Table 9 for details). To do so, we manually labeled at least 3 different prompts for each of these tasks, and then invoke GPT4 to automatically generate more based on the following “meta prompt”, i.e., the prompt used to generate prompts for different tasks:
Here are some instructions that define a visual-language task. Continue to write 15 instructions with the same meaning: 1) PROMPT1; 2) PROMPT2; 3) PROMPT3;
Besides, we also collect some available public (visual-)text instruction data (also listed in Table 9) to further improve the ability of our model to follow open-ended instructions, including the instruction data used in FlanT5 , Alpaca , Mini-GPT4 , LLAVA , and Baize .
We follow the same causal prediction loss as in pretraining, i.e., the cross entropy loss to predict the next word based on all previous tokens. Nevertheless, we observed that different weight combinations of the instruction data have a crucial influence on the final performance. Empirically, we finally impose the weight strategy presented in Table 9.
Experiment
In this section, we aim to answer the following questions according to empirical studies:
a) How can we evaluate the performance of a GPT4-style model? (Section 3.1)
b) Compared to existing models, what are the advantages of our Lynx? (Section 3.2)
c) What matters to train a high-performance GPT4-style model? (Section 3.3)
d) What is the performance of Lynx in open-world zero-shot scenarios? (Section Appendix F)
The evaluation of GPT4-style generative language models is challenging because the quality of natural languages is inherently subjective and highly depends on specific cases. Existing models like PaLM-E , PaLI , BLIP2 , or InstructBLIP turn to the evaluation on visual-language benchmarks like image caption or visual question answering , i.e., fine-tuning multi-modal LLMs on a single downstream task on which the evaluation is conducted. Nevertheless, though it may achieve better performance, over-finetuning on such benchmarks will damage the generation ability of large language models, which conflicts with the primary motivation to use large language models. Moreover, such benchmarks, especially the (semi-)automatically generated ones like TDIUC , always contain a high ratio of easy or noisy examples, making them less suitable. On the contrary, other methods like MiniGPT4 or LLaVA only showcase their performance in some challenging yet practical scenarios without quantitative results due to the lack of quantitative benchmarks for such generative multi-modal language models. Therefore, in this section, we propose to evaluate the GPT4-style models in the following two aspects:
A cleaned subset of visual-language benchmark, which should be challenging and compatible with generative models, with prompted GPT4 to get the quantitative results.
An open-world challenging yet practical test set to evaluate the performance on realistic scenarios where GPT4-style models are needed, with humans to evaluate the user experience.
To do so, we manually collect an Open-VQA test set consisting of 450 samples with image or video input, which contains diverse questions on objects, OCR, counting, reasoning, action recognition, chronological ordering, etc., from VQA 2.0 , OCRVQA , Place365 , MSVD , MSRVTT , and Something-Something-V2 (SthV2) . Though Place365 is a classification task and SthV2 is a video captioning task, we write proper prompts to make them both VQA tasks. Besides, we carefully examine the data and modify the questions and ground-truth answers if necessary to make them reliably correct and challenging enough to be a benchmark for GPT4-style models. Randomly sampled examples are given in Fig. 3(a). Different from the traditional VQA benchmark, Open-VQA supports open-ended answers. To achieve so, we prompt GPT4 to make it the referee, which achieves a consistency of more than 95% compared with humansWe evaluate the consistency on 100 samples from a randomly selected subset with our model.. The prompt for GPT4 used in this phase is as follows:
Given the question “QUESTION”, does the answer “PREDICTION” imply the answer “GROUND_TRUTH”? Answer with Yes or No.
Moreover, general-purpose language generation with image inputs is also important to multi-modal LLMs. Therefore, we also adopt the OwlEval test set proposed by mPLUG-owl , which contains 82 questions based on 50 images, where 21 from MiniGPT-4 , 13 from MM-REACT , 9 from BLIP2 , 3 from GPT4 , and 4 collected by mPLUG-owl itself. The test set includes diversified and practical cases such as dense image captioning, dialogue writing, story writing, poem writing, teaching, programming, etc.
We give some examples in Fig.3(b). However, OwlEval is proposed together with mPLUG-owl. Hence, directly using it as the benchmark is possibly unfair to other models. To make the comparison fair, we pad each image in the OwlEval with 8 pixels as shown in Fig.3(b) before feeding them into the models. We recruit human annotators to evaluate the performance. Scores range from 1 to 5. If two models are considered to be equally good or bad, they will have the same score. For each data, the annotator will assign a score for each model. We only allow at most 2 models that are equally good or bad, and for each annotator, the total number of ties should be no more than 10 for the whole set. During the evaluation, the correctness has the highest priority, then should be the richness of the generated content.
Finally, we also compare our method with others on the newly proposed MME benchmark , which includes 14 different subtasks that evaluate the perception and cognition ability of multi-modal large language models.
2 Quantitative Experiments
We first evaluate our model as well as several existing open-sourced multi-modal LLMs on the Open-VQA benchmark. Results are shown in Table 8. We can conclude that our model has achieved the best performance both in the image and video understanding tasks. Notably, InstructBLIP also achieves high performance in most cases, even better than our model in OCR, color recognition, and action recognition tasks. However, we observe that it always outputs one word for the question as shown in Fig.5 and 6, which is less preferred by most of the users (see Fig.4). We also showcase some of the examples in Fig. 5. More cases including video VQA examples can be found in Fig. 10 and 11 in the appendix. We can see that our model can give the correct answer in most cases as well as a concise reason that supports the answer, which makes it more user-friendly.
OwlEval benchmark
We evaluate the performances of general-purpose natural language generation on OwlEval test set. From the human evaluation results in Fig.4, we can see that our model has the best language generation performance while keeping high performance on the Open-VQA benchmark. BLIP2 and InstructBLIP , though achieved high performance on the Open-VQA benchmark, are not preferred by human users due to their extremely short outputs, i.e., in most cases, they only output one word or phrase as the answer without any explanation. In contrast, MiniGPT4 and mPLUG-Owl are trained less to fit the Open-VQA benchmark and keep more language generation ability. Hence, they are preferred over the BLIP models, though they may make more factual errors. We also show some results on the OwlEval in Fig. 6.
In general, we observe that if a model has lower accuracy on the Open-VQA benchmark, it tends to make factual errors inconsistent with the given image during text generation. Nevertheless, models with higher performance on the Open-VQA benchmark usually tend to lose language generation ability, e.g., generate short sentences. We attribute this conclusion to the under-training or over-training on visual-language tasks. To be specific, existing training data from visual-language tasks always includes short outputs. By training on these data, the model can learn to align the visual and linguistic concepts, yet lose the language generation ability inherited from the large language model. From the high performance of our model, we can see that one possible way to train a high-performance model with better language generation ability is to carefully select and clean the data, as well as design the proper sampling ratios. Nevertheless, the key to balance language generation and correctness is a high-quality visual-language dataset that contains clean and rich expressions, which should be explored in our future work.
MME benchmark
We also compare Lynx with available existing open-source models on the MME benchmark . Results are shown in Figure 7 and Appendix B. We can see that our model is a state-of-the-art model in 7 out of 14 subtasks, especially for the perception tasks including Color, Celebrity, Scene, Landmark, Position, Count, and Existence. Yet, from the figure, we can also see that our model seems not to perform well on cognition tasks including Code Reasoning, Text Translation, and Numerical. Notably, cognition benchmarks including Code Reasoning, Text Translation, and Numerical in MME only contain 20 examples, which may cause high variance in the evaluation of different checkpoints.
3 Ablation Study
We conduct an in-depth ablation study to investigate the impact of different components or training recipes on multi-modal understanding and language generation performances. In this section, we follow the same evaluation method proposed in Section 3.1.
As shown in Table 3, our experiments show that in the aspect of correctness, instruction-finetuned backbone (e.g.Vicuna) performs slightly better on our Open-VQA benchmark (like LLaVA) as shown in Table 3 and 4, but slightly worse on the OwlEval benchmark (Figure 4). However, Vicuna-based model does indeed follow the instruction better. For example, the average answer length given the instruction “give a short answer” is 15.81, compared to 20.15 from the LLaMA-based model. One can also refer to Figure 9(a) for examples of the comparison in terms of their instruction-following ability.
Impact of Diversified Prompts
It has been proved to be important to train LLMs on instruction data so as to make them follow instructions correctly . Therefore, we ablate our model with diversified prompts written by both users and GPT4. The results in Table 3 and 4 show that our prompts help to balance different abilities. Moreover, we also find that by using diversified prompts, our model can follow the open-ended instructions better than the ones trained without these prompts (Table 10). This observation accords with the text-only models. The human evaluation results in Figure 8(b) also accord with our observations. Diversified tasks and prompts will help to improve the generalization of the model to new tasks and instructions.
Impact of Training Data
We investigate the impact of data quantity and quality by training our model with or without the large-scale yet noisy image-text pairs (COYO700M and DataComp1B ). During our experiments, we find training data in both pretraining and finetuning largely influence the model performance. Different from traditional visual-language pretraining , we find that multi-modal LLMs do not benefit from large-scale but noisy image-text pairs because many of the texts in such datasets are not fluent or natural language expressions. For the generative pretraining in our model, they largely damage the language generation ability as shown in Figure 9(b). As a result, pretraining on such large-scale datasets achieves no better results than only training on a much smaller but cleaner dataset as evaluated by the human users as shown in Figure 8(c).
Prefix-Tuning vs. Cross-Attn
We follow Flamingo , concretely Open-Flamingo , to implement the cross-attention method. Following its original settings, we only use multi-modal instruction data for pre-training. For the finetuning stage, we experiment with two variants, with or without trainable LLM, i.e., with or without the use of text instruction data. As shown in Table 3 and 4, both of them perform worse than our prefix-tuning with adapters. Though the models can generate fluent and relevant responses, their outputs usually do not give correct answers to the questions. We also verified our conclusion with human annotators, as shown in Figure 8(d). Results show that human users give lower preference to the cross-attention models. Overall, cross-attention models could require more hyper-parameter searching to achieve better performances, and we leave it to further work.
Impact of Larger Image Resolution
We increase image resolution in the first stage with only 10K step training. After that, we freeze the vision encoder and thus the expense of increasing image resolution is affordable. For rigor, we also conducted an experiment to verify the impact of image resolutions on the model performance. The experiment results in Table 3 and 4 show that the training on 420x420 resolution achieves better performance than the models only trained on 224x224.
Related Work
Large language models (LLMs) have been widely investigated in recent years due to their good generality on zero-shot tasks, including GPT3 , PaLM , BLOOM , Chinchilla , T5 , LLaMA , OPT , GLM , etc. After being pre-trained on massive text corpora, such models can perform surprisingly well on downstream tasks without further finetuning. In particular, the simple yet efficient structure of decoder-only models like GPT-3 can easily scale up to hundreds of billions of parameters and show an elegant scaling law with the increase of model size and data amounts . Moreover, recent advances in instruction finetuning have also shown that large-scale language models can be finetuned with limited amounts of instruction data to follow open-ended instructions in natural language. This not only improves their performance on downstream tasks substantially but also makes it a user-friendly assistant in our daily life .
Centralized Multi-modal Interactive System.
Inspired by the achievements in LLMs, it is straightforward to ask a question: Is it possible to design a model that accepts multi-modal inputs while being able to chat with humans in natural language? Therefore, recent works investigate actively to design of such multi-modal interactive models. One of the most intuitive ideas, such as Visual ChatGPT , MM-REACT , HuggingGPT , InternGPT , SayCan , InnerMonologue , integrates various existing individual models or tools such as OCR, object detection, image captioning, visual question answering, text-to-image generation, or robot manipulation policies by a centralized controller. In such a system, the LLM works as a “manager” that directly accepts instructions from users and selects the most appropriate tools to respond to requests while the integrated individual models are “workers” responsible for a specific kind of task. Typically, such models are powerful to address problems that are already well-defined. Yet, they, to some extent, lack zero-shot ability when encountering open-ended instructions which cannot be handled by any of their workers.
End-to-end Multi-modal Large Language Models.
By contrast, inspired by the recent advances of LLMs, it has also been shown feasible and promising to directly train the neural networks that directly accept multi-modal inputs and output responses end-to-end. To achieve so, one intuitive idea is to adapt the LLMs to multi-modal inputs by adding some additional trainable parameters and finetuning them on multi-modal data. For example, Flamingos is one of the early works to explore this idea. Firstly, it takes a vision encoder (like NFNet in their original version, or recent CLIP ViT ) to extract visual embeddings. Then, it applies multi-layer cross-attention to fuse the multi-modal inputs for the final prediction. Recent works directly concatenate vision embeddings to the inputs of LLMs and finetune LLMs end-to-end. To do so, they usually add an additional projection layer to map the vision embeddings to the same dimension as the language embeddings, and then directly feed them into LLMs for further training. Different methods may take different training strategies. BLIP2 designs a Q-Former, which is the only trainable part, to align the dimensions of vision and language tokens. PaLM-E , which is built upon PaLM , is trained totally end-to-end with no fixed layers using a mix of multi-modal datasets including WebLI 10B dataset . Mini-GPT4 freezes all weights of the vision encoder and the LLM while only finetuning the weights of the projection layer. LLAVA fixes the vision encoder while keeping the LLMs trainable during the instruction finetuning stage. mPLUG-owl tunes the vision encoder and keeps LLMs fixed to align the vision and language embeddings in the first stage while further tuning the LLMs and keeping the vision encoder fixed in the second instruction-finetuning stage. KOSMOS-1 does not rely on any pretrained LLMs and is trained from scratch on large amounts of mixed data including image-text pairs (COYO700M , LAION2B , etc.), text corpora (Common Crawl, the Pile , etc.), and interleaved image-text data. These models are all powerful and show promising results to develop multi-modal large language models.
Discussions and Limitations
As shown in our experiments, prefix-tuning with adaptors show good performance on open-ended instruction-following tasks after training in billions of multi-modal tokens. By contrast, cross-attention models are not that efficient to achieve good performance, though more hyper-parameter searching could improve its performances and we leave it in future work.
Multi-modal LLMs are not as instruction-following as LLMs.
In our experiments, we find that current multi-modal LLMs are not as good at the instruction following as language models. For example, InstructBLIP tends to generate short responses regardless of the input instructions, while other models tend to generate long sentences without considering the instruction like “Give a short answer” or “Answer in one word”. We assume that this is from the lacking of high-quality and diversified multi-modal instruction data.
The quality of training data is critical to model performance.
As concluded in Section 3.3, based on the experimentation on different pretraining data, we find that a small number of high-quality data with fluent texts can perform even slightly better than the large-scale noisy datasets. We attribute this to the difference between generative pretraining and contrastive pretraining, since generative pretraining is directly learning the conditional distribution of words but not the similarity between texts and images. Therefore, to train a high-performance multi-modal LLM, despite the quantity of data, it is crucial to prepare a high-quality dataset that satisfies: 1) it includes high-quality and fluent texts; 2) it aligns the texts and images well.
Tasks and prompts are crucial for zero-shot abilities.
As shown in Section 3.3, diversified prompts have a great impact on the final performance. The essential observation behind this is that the zero-shot generality of multi-modal language models depends on the diversity of tasks involved during training. The model can generalize to more and more unseen instructions as it sees more and more types of tasks. This accords with the observation in text-only models .
Balancing the correctness and language generation ability is important.
In our experiments, we find that if the model is under-trained on downstream tasks such as VQA, it will suffer from the problem of hallucination and keep making mistakes. While if the model is over-trained on downstream tasks, it will not be able to follow the user’s instructions to generate long answers. Therefore, it would be important to carefully balance the training data to train it so as to correctly read images and videos while keeping its generation ability.
2 Limitations
It is hard to evaluate a multi-modal large language model since its evaluation is essentially different from traditional visual-language models. Though we take the first step to quantitatively evaluate both the multi-modal understanding accuracy and language generation ability, it is still an open problem: how can we establish a comprehensive and automatic benchmark to evaluate existing multi-modal large language models?
Training Data
Though we have successfully collected and cleaned a mixed dataset to train our Lynx, we still put a lot of effort to balance different abilities (e.g. correctness and language generation, long and short answers). Moreover, there are still no available image-text datasets that contain long texts which are ideal for pretraining. Besides, restricted by the computational resources that we can use, we do not conduct extensive experiments to find the optimal data combination strategy (e.g. sampling ratios, tasks, and prompts), which has been left for future work.
Multi-lingual
Our model is built upon LLaMA , which is mainly trained on English corpus. Therefore, our model is not that good at multi-lingual responses. Though it can understand and sometimes output other languages (like shown in Figure 16), it is still unexplored how to build a high-performance multi-lingual and multi-modal large language model.
Safety
Currently, we do not conduct safety checks and restrict the outputs of our model. Therefore, the model may output contents that are not appropriate and even toxic, depending on and restricted by the data used for training. The authors do not support the use of harmful language generation using our codes and models, like any usage on ethical, political, and racism issues.
Conclusions
In this paper, we present Lynx, a multi-modal GPT4-style large language model that can take as input images/videos and responses with open-ended natural languages. Through extensive empirical study, we show that our model outperforms other existing open-source models both in multi-modal understanding and language generation. We also explore different factors that can affect the performance of a multi-modal large language model and conclude that: 1) for network structure, prefix-tuning is better than cross-attention to fuse different modalities; 2) instruction following is closely related to the number of tasks and prompts used for training; 3) the generative pretraining is much more sensitive the quality of training data than previous pretraining methods such as contrastive training; 4) balancing the correctness and language generation is important for multi-modal large language models.
For future work, it is promising to scale up the model to a larger size (e.g. 30B and 65B LLaMA ), as well as a larger and more diversified set of instructional tasks. Moreover, a large-scale and high-quality multi-modal dataset is also needed to train such models. Therefore, it is worth the effort to collect such a dataset, which will be a great contribution to this area. Multi-lingual ability and safety are also undoubtedly crucial for realistic applications.
Acknowledgements
We would like to acknowledge Hang Li at ByteDance for his generous assistance in insightful comments in technical discussions. Additionally, we extend our appreciation to the colleagues at ByteDance for their efforts and support of this project. We are also thankful to the LLaMA and Vicuna teams for granting us access to their models.
References
Appendix A Experimental Details
We use the DeepSpeed to accelerate training, and set the BFloat16 as the default model precision. We report the detailed model training hyperparameters in Table 5.
A.2 Hyper-parameters for Generation
During the deployment of all models, we find that for most of them, the performance would be better if we apply a description-first strategy. That is, before sending the request from the user, by default, we feed a fixed prompt “Describe the image in detail” first in the “0th” round of the conversation. After that, the user’s instructions will be sequentially processed. Nevertheless, we found that the quality of generated texts by MiniGPT4 using this description-first strategy is worse than the ones directly generated. Therefore, for MiniGPT4 , we generated the response with its default settings. Similarly, for mPLUG-owl , we follow the default parameters presented at http://vlarena.opengvlab.com/. Detailed settings can be found in 6 for different tasks.
Appendix B MME Performance
Appendix C Case Study
C.2 Video VQA Cases
Appendix D Training Data
D.2 Prompt Examples
Appendix E OwlEval Cases
Appendix F Open Demonstrations
F.2 Multi-lingual Response
F.3 Instruction-following Ability
We also demonstrate the instruction-following ability of different models. We can see that both Lynx and mPLUG-owl can follow instructions correctly to some extent. Yet, InstructBLIP is not sensitive to different instructions.