OneLLM: One Framework to Align All Modalities with Language
Jiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang, Kaipeng Zhang, Dahua Lin, Yu Qiao, Peng Gao, Xiangyu Yue
Introduction
Large Language Models (LLMs) are getting increasingly popular in the research community and industry due to their powerful language understanding and reasoning capabilities. Notably, LLMs such as GPT4 have reached performance nearly on par with humans in various academic exams. The progress in LLMs has also inspired researchers to employ LLMs as an interface for multimodal tasks, such as vision-language learning , audio and speech recognition , video understanding , etc.
Among these tasks, vision-language learning is the most active field, with more than 50 vision LLMs proposed in the recent half-year alone . Typically, a vision LLM comprises a visual encoder, an LLM, and a projection module connecting the two components. The vision LLM is first trained on massive paired image-text data for vision-language alignment and then fine-tuned on visual instruction datasets, enabling it to complete various instructions tied to visual inputs. Beyond vision, significant efforts have been invested in developing other modality-specific LLMs, such as audio , video , and point clouds . These models generally mirror the architectural framework and training methodology of vision LLMs, and rely on the solid foundation of pretrained modality-specific encoders and well-curated instruction-tuning datasets for their effectiveness.
There are also several attempts to integrate multiple modalities into one MLLM . As an extension of vision LLM, most previous works align each modality with the LLM using modality-specific encoders and projection modules (middle of Fig. 1). For instance, X-LLM and ChatBridge connect pretrained image, video, and audio encoders with LLMs using separate Q-Former or Perceiver models. However, these modality-specific encoders usually differ in architecture and considerable effort is required to unify them into a single framework. Furthermore, pretrained encoders that deliver reliable performance are usually restricted to widely used modalities such as image, audio, and video. This limitation poses a constraint on MLLMs’ ability to expand to more modalities. Thus, a crucial challenge for MLLMs is how to build a unified and scalable encoder capable of handling a wide range of modalities.
We get inspiration from recent works on transferring pretrained transformers to downstream modalities . Lu et al. proved that a frozen language-pretrained transformer can achieve strong performance on downstream modalities such as image classification. Meta-Transformer demonstrated that a frozen visual encoder can achieve competitive results across 12 different data modalities. The insights from the works mentioned above suggest that pretrained encoders for each modality may not be necessary. Instead, a well-pretrained transformer may serve as a universal cross-modal encoder.
In this paper, we present OneLLM, an MLLM that aligns eight modalities to language using one unified framework. As shown in Fig. 1, OneLLM consists of lightweight modality tokenizers, a universal encoder, a universal projection module (UPM), and an LLM. In contrast to prior works, the encoder and projection module in OneLLM are shared across all modalities. The modality-specific tokenizers, each comprised of only one convolution layer, convert input signals into a sequence of tokens. Additionally, we add learnable modality tokens to enable modality switching and transform input tokens of diverse lengths into tokens of a fixed length.
Training a model of this complexity from scratch poses significant challenges. We start from a vision LLM and align other modalities to the LLM in a progressive way. Specifically, (i) we build a vision LLM with pretrained CLIP-ViT as the image encoder, accompanied by several transformer layers as the image projection module, and LLaMA2 as the LLM. After pretraining on massive paired image-text data, the projection module learns to map visual representations into the embedding space of LLM. (ii) To align with more modalities, we need a universal encoder and projection module. As discussed before, the pretrained CLIP-ViT is possible to serve as a universal encoder. For UPM, we propose to mix multiple image projection experts as a universal X-to-language interface. To increase the model capability, we also design a dynamic router to control the weight of each expert for the given inputs, which turns UPM into soft mixtures-of-experts . Finally, we progressively align more modalities with the LLM based on their data magnitude.
We also curate a large-scale multimodal instruction dataset, including captioning, question answering, and reasoning tasks across eight modalities: image, audio, video, point clouds, depth/normal map, Inertial Measurement Unit (IMU), and functional Magnetic Resonance Imaging (fMRI). By finetuning on this dataset, OneLLM has strong multimodal understanding, reasoning, and instruction-following capabilities. We evaluate OneLLM on multimodal captioning, question answering and reasoning benchmarks where it achieves superior performance than previous specialized models and MLLMs. In conclusion, we summary our contributions as:
We propose a unified framework to align multimodal inputs with language. Different from existing works with modality-specific encoders, we show that a unified multimodal encoder, which leverages a pretrained vision-language model and a mixture of projection experts, can serve as a general and scalable component for MLLMs.
To the best of our knowledge, OneLLM is the first MLLM that integrates eight distinct modalities within a single model. With the unified framework and progressive multimodal alignment pipeline, OneLLM can be easily extended to incorporate more data modalities.
We curate a large-scale multimodal instruction dataset. OneLLM finetuned on this dataset achieves superior performance on multimodal tasks, outperforming both specialist models and existing MLLMs.
Related Work
Large Vision-Language Models. Large Language Models (LLMs) have gained a lot of attention recently. Therefore, extending LLMs to the vision domain is an emergent and rapidly growing research area. Flamingo is a pioneer to inject frozen visual features into LLM with cross-attention layers, achieving superior performance on a wide range of vision-language tasks. BLIP2 uses a Q-Former to aggregate visual features into a few tokens aligned with LLM. Recently, with the popularity of instruction-following LLMs, vision LLMs have experienced a new explosion. LLaMA-Adapter connects pretrained CLIP and LLaMA with parameter-efficient fine-tuning methods, which can tackle close-set visual question answering and image captioning tasks. Subsequent works propose to train such model on large-scale image-text data, enabling it to complete various instructions about images. Among them, LLaVA adopt a linear layer to directly project visual tokens into LLMs, while MiniGPT-4 and some other works resample visual tokens into fixed-length tokens, reducing the computation cost of LLMs. Our work also belongs to the later branch. We preset learnable tokens for each modality (i.e., modality tokens), which are then used to aggregate input information and generate fixed-length tokens for all modalities.
Multimodal Large Language Models. In addition to vision LLMs, recent works proposed to extend LLMs to other modalities, such as audio , video and point cloud . These works make it possible to unify multiple modalities into one LLM. X-LLM adopts modality-specific Q-Former and adapters to connect pretrained image, audio and video encoders with LLMs. ChatBridge and AnyMAL follow a similar architecture with X-LLM but adopts Perceiver and linear layers respectively to align modality encoders with LLMs. Meanwhile, PandaGPT and ImageBind-LLM utilize ImageBind as the modality encoder and therefore naturally support multimodal inputs. However, current MLLMs are limited to supporting common modalities such as image, audio and video. It remains unclear how to expand MLLMs to more modalities with a unified framework. In this work, we propose a unified multimodal encoder to align all modalities with language. We show that one universal encoder and projection module can effectively map multimodal inputs to LLM. To our knowledge, OneLLM is first MLLM capable of supporting eight distinct modalities.
Multimodal-Text Alignment. Aligning multiple modalities into one joint embedding space is important for cross-modal tasks, which can be divided into two lines of works: discriminative alignment and generative alignment. The most representative work of discriminative alignment is CLIP , which utilize contrastive learning to align image and text. Follow-up works extend CLIP to audio-text , video-text , point-text etc. Besides, ImageBind proposes to bind various modalities to images with contrastive learning. On the other hand, generative alignment has attracted much attention in the era of LLM. GIT aligns image and text using a generative image-to-text transformer. BLIP2 proposes generative pretraining to connect frozen vision encoder and LLM. VALOR and VAST extends the training paradigm of BLIP2 to more modalities such as audio and video. Our work also belongs to generative alignment. In contrast to prior works, we directly align mutlimodal inputs to LLMs, thus getting rid of the stage of training modality encoders.
Method
In this section, we will first introduce the architecture of OneLLM (Sec. 3.1) and then present our two training phases: progressive multimodal alignment (Sec. 3.2) and unified multimodal instruction tuning (Sec. 3.3).
Fig. 2 depicts the four main components of OneLLM: modality-specific tokenizers, a universal encoder, a universal projection module (UPM) and an LLM. Detailed descriptions are presented in the following sections.
Universal Encoder. As discussed in Sec. 1, frozen pretrained transformers demonstrate strong modality transfer capability . Therefore, we leverage pretrained vision-language models as the universal encoder for all modalities. Vision-language models, when trained on extensive image-text data, typically learn robust alignment between vision and language, so they can be easily transferred to other modalities. In OneLLM, we use CLIP-ViT as a universal computation engine. Following previous works , we keep the parameters of CLIP-ViT frozen during training. Note that for video signals, we will feed all video frames into the encoder in parallel and perform token-wise averaging between frames to speed up training. Other strategies, such as token concatenation, may further enhance the model’s video understanding capability.
Universal Projection Module. In contrast to existing works with modality-specific projection, we propose a Universal Projection Module (UPM) to project any modality into LLM’s embedding space. As shown in Fig. 2, UPM consists of projection experts , where each expert is a stack of transformer layers pretrained on image-text data (will discuss in Sec. 3.2). Although one expert can also realize any modality-to-LLM projection, our empirical findings suggest that multiple experts are more effective and scalable. When scaling to more modalities, we only need to add a few parallel experts.
LLM. We employ the open-source LLaMA2 as the LLM in our framework. The input to LLM includes projected modality tokens and the text prompt after word embedding. Note we always put modality tokens at the beginning of the input sequence for simplicity. Then LLM is asked to generate appropriate response conditioned on modality tokens and text prompt.
2 Progressive Multimodal Alignment
Image-text alignment has been well investigated in previous works . Therefore, a naive approach for multimodal alignment is to jointly train the model on multimodal-text data. However, training models directly on multimodal data can lead to biased representations between modalities due to the imbalance of data scale. Here we propose to train an image-to-text model as initialization and progressively ground other modalities into LLM.
Multimodal-Text Alignment. We formulate multimodal-text alignment as a continual learning process . At timestamp , we have trained the model on a set of modalities , and the current training data is from . To prevent catastrophic forgetting, we will sample evenly from both previous trained data and current data. In our case, we divide multimodal-text alignment into multiple training stages based on their data magnitude: stage I (image), stage II (video, audio and point cloud) and stage III (depth/normal map, IMU and fMRI). If we want to support new modalities, we can repeat the training episode, i.e., sampling a similar amount of data from previous modalities and jointly training the model with the current modalities.
Multimodal-Text Dataset. We collect X-text pairs for each modality. The image-text pairs include LAION-400M and LAION-COCO . The training data for video, audio and point clouds are WebVid-2.5M , WavCaps and Cap3D , respectively. Since there is no large-scale depth/normal map-text data, we use pretrained DPT model to generate depth/normal map. The source images and text and from CC3M . For IMU-text pairs, we use the IMU sensor data of Ego4D . For fMRI-text pairs, we use fMRI signals from the NSD dataset and take the captions associated with the visual stimuli as text annotations. Note that the input to LLM is the concatenation of modality tokens and caption tokens. We do not add system prompts at this stage to reduce the number of tokens and speed up training.
3 Unified Multimodal Instruction Tuning
After multimodal-text alignment, OneLLM becomes a multimodal captioning model which can generate a short description for any input. To fully unleash OneLLM’s multimodal understanding and reasoning capabilities, we curate a large-scale multimodal instruction tuning dataset to further finetune OneLLM.
Multimodal Instruction Tuning Dataset. We collect instruction tuning (IT) dataset for each modality. Following previous works , the image IT datasets are sampled from the following datasets: LLaVA-150K , COCO Caption , VQAv2 , GQA , OKVQA , A-OKVQA , OCRVQA , RefCOCO and Visual Genome . The video IT datasets include MSRVTT-Cap , MSRVTT-QA and video instruction data from . The audio IT datasets include AudioCaps and audio conversation data from . The point cloud IT dataset is a 70K point cloud description, conversation and reasoning dataset from . The depth/normal map IT datasets are generated from image IT datasets: we random sample 50K visual instruction data from LLaVA-150K and generate depth/normal map using DPT model . For IMU and fMRI IT datasets, we also random sample a subset from Ego4D and NSD , respectively. Finally, our mutlimodal IT datasets have about 2M items, covering multiple tasks such as detailed description/reasoning, conversation, short question answering and captioning.
Prompt Design. Given the diverse modalities and tasks within our multimodal IT datasets, we carefully design the prompts to avoid conflicts between them. (a) When utilizing IT datasets generated by GPT4 (e.g., LLaVA-150K), we adopt the original prompts provided by these datasets. (b) For captioning tasks, we empoly the prompt: Provide a one-sentence caption for the provided {modal}. (c) For open-ended question answering tasks, we enhance the question with Answer the question using a single word or phrase. (d) For question answering tasks with options, the prompt is: {Question} {Options} Answer with the option’s letter from the given choices directly. (e) For IMU and fMRI datasets, we apply prompt such as Describe the motion and Describe this scene based on fMRI data. Despite using these fixed prompts, our experiments indicate that OneLLM is capable of generalizing to open-ended prompts during inference. For detailed prompts on each task and modality, please check out Sec. C.4 of the appendix.
In the instruction tuning stage, we organize the input sequence as: where is the modality tokens, is the system prompt, corresponds to the -th instruction-answer pair in a conversation. Note that for multimodal inputs involving multiple modalities, such as audio-visual tasks , we position all modality tokens at the start of the input sequence.
We fully finetune the LLM and keep rest parameters frozen. Although recent works often employ parameter-efficient methods , we empirically show that the full finetuning approach more effectively harnesses the multimodal capabilities of OneLLM, particularly with the utilization of smaller LLMs (e.g., LLaMA2-7B).
Experiment
Training Details. We use AdamW optimizer with =0.9, =0.95 and weight decay of 0.1. We apply a linear learning rate warmup during the first 2K iterations. For stage I, we train OneLLM on 16 A100 GPUs for 200K iterations. The effective batch size (using gradient accumulation) is 5120. The maximum learning rate is 5e-5. For stage II (resp. III), we train OneLLM on 8 GPUs for 200K (resp. 100K) with an effective batch size of 1080 and maximum learning rate of 1e-5. In the instruction tuning stage, we train OneLLM on 8 GPUs for 1 epoch (96K) with an effective batch size of 512 and maximum learning rate of 2e-5.
2 Quantitative Evaluation
We evaluate OneLLM on multimodal tasks and put evaluation details to Sec. D of the appendix.
Image-Text Evaluation. In Tab. 1, we evaluate OneLLM on visual question answering (VQA), image captioning and recent multimodal benchmarks. For VQA tasks, OneLLM-7B outperforms other MMLLMs such as ChatBridge-13B and AnyMAL-13B by a large margin. Our 7B model is even better than AnyMAL with 70B parameters. For image captioning tasks, OneLLM-7B is on-par with ChatBridge-13B. Although OneLLM is not specifically designed for vision tasks, our results demonstrate that OneLLM can also reach the leading level in vision specialized LLMs, and the gap between MMLLMs and vision LLMs has further narrowed.
Video-Text Evaluation. As shown in Tab. 2, we evaluate OneLLM on video QA and captioning tasks. Our model outperforms both MLLMs (ChatBridge and AnyMAL) and video-specific models (FrozenBiLM and InternVideo ) in video QA tasks. Notably, our training datasets do not include video QA data like NextQA and How2QA , which are video QA tasks that provide answer options. However, our model’s training on similar VQA datasets (e.g., A-OKVQA ) has evidently enhanced its emergent cross-modal capabilities, contributing to the improved performance in video QA tasks.
Audio-Text Evaluation. We evaluate OnLLM on audio captioning and QA tasks. In Tab. 3, we outperforms both ChatBridge and LTU on Clotho Caption . Notably, our zero-shot result on Clotho AQA is on par with fully finetuned Pengi . Similar to our conclusion on video QA, we believe that the captioning task requires more dataset-specific training, while the QA task may be a more accurate measure of the model’s inherent zero-shot understanding capabilities.
Audio-Video-Text Evaluation. We evaluate OneLLM on audio-video-text tasks, such as QA (MUSIC AVQA ), captioning (VALOR-32K ) and dialog completion (AVSD ) based on the video and background audio. As shown in Tab. 4, OneLLM-7B surpasses ChatBridge-13B on all three datasets. Note that ChatBridge was trained on an audio-visual dataset , while OneLLM has not been trained on any audio-visual datasets. Since all modalities in OneLLM are well aligned with language, we can directly input video and audio signals to OneLLM during inference.
Point Cloud-Text Evaluation. In Tab. 5, We evaluate OneLLM on point cloud captioning and classification tasks. OneLLM can achieve excellent captioning results due to our carefully designed instruction prompts for switching between tasks (Sec. 3.3), while InstructBLIP and PointLLM struggle to generate short and accurate captions. On the classification task, OneLLM can also achieve comparable results to PointLLM.
Depth/Normal Map-Text Evaluation. Since there are currently no QA and captioning tasks using depth/normal maps, we evaluate OneLLM on two scene classification datasets . The performance, as displayed in Tab. 6, reveals that OneLLM achieves superior zero-shot classification accuracy compared to CLIP. These results affirm that OneLLM trained on synthetic depth/normal map data can adapt to real world scenarios.
IMU-Text and fMRI-Text Evaluation. Since IMU/fMRI to text generation are seldom explored in previous literature, we solely report our results on IMU/fMRI captioning. For IMU captioning on Ego4D , we evaluate OneLLM on a held-out subset with 2000 items. The CIDEr and ROUGE-L score are 24.9 and 19.5, respectively. For fMRI captioning on NSD , we evaluate OneLLM on its testing set, where OneLLM achieves 31.7 CIDEr and 25.1 ROUGE-L.
3 Ablation Experiments
In this section, we will explore some key designs of OneLLM. Our ablation experiments are conducted on a subset of the training data, which only includes multimodal alignment and instruction tuning datasets of image, audio and video, except for studies on the number of experts. Other settings remain unchanged if not specified.
Separate Training vs. Joint Training. An important question for MLLMs is whether a jointly trained MLLM is better than modality-specific MLLM? To address this, we compare the performance of separately trained MLLMs against a jointly trained MLLM in Tab. 7 (a). In separate training, the model can only access its own data; in joint training, the model is jointly trained on all data. On two image-text tasks NoCaps and VQAv2, we can see that separately and jointly trained models achieve comparable results; While separately trained audio and video models are much worse than the jointly trained model on ClothoQA and MSVDQA, respectively. This suggest that joint training substantially benefits data-scarce modalities (e.g., audio and video), by allowing for the transfer of learned knowledge (e.g., question answering) across modalities.
Image Alignment Benefits Multimodal Alignment. Tab. 7 (b) demonstrate that OneLLM with image-text alignment can help multimodal-text alignment. If we directly align all modalities with text using a random initialized model (i.e. universal projection module), the performance on image and video will drop significantly. Instead, OneLLM with image-text pretraining can better balance different modalities.
Number of Projection Experts. The number of projection experts in UPM is closely related to the number of modalities that OneLLM can accommodate. As shown in Tab. 7, OneLLM with three projection experts is enough to hold all modalities. Increasing the number of experts does not bring about the desired improvement, while the results with one expert is also not satisfactory.
Router Type. The modality router is to link multiple projection experts into a single module. Here we discuss three types of router: constant router, sparse router and the default soft router. (a) Constant router links experts with a constant number . The output of constant router is . (b) Sparse router only selects one expert with the maximum routing weight. The output is where . As shown in Tab. 7 (d), soft router outperforms other two routers, indicating its effectiveness for dynamic routing of multimodal signals.
4 Qualitative Analysis
Fig. 3 gives some qualitative results of OneLLM on eight modalities. We show OneLLM can (a) understand both visual and textual content in images, (b) leverage temporal information in videos, (c) do creative writing based on audio content, (d) understand the details of 3D shapes, (e) analyze visual scenes recorded in fMRI data, (f) guess the person’s action based on motion data, and (g)-(h) scene understanding using depth/normal map. Due to space limit, we put more qualitative results to Sec. F of the appendix.
Conclusion
In this work, we introduce OneLLM, an MLLM that aligns eight modalities with language using a unified framework. Initially, we first train a basic vision LLM. Building on this, we design a multimodal framework with a universal encoder, a UPM and an LLM. By a progressive alignment pipeline, OneLLM can handle multimodal inputs with a single model. Furthermore, we curate a large-scale multimodal instruction dataset to fully unleash OneLLM’s instruction-following capability. Finally, we evaluate OneLLM on 25 diverse benchmarks, showing its excellent performance.
Limitation and Future Work. Our work faces two primary challenges: (i) The absence of large-scale, high-quality datasets for modalities beyond image, which leads to a certain gap between OneLLM and specialized models on these modalities. (ii) Fine-grained multimodal understanding in high-resolution images, long sequences video and audio etc. In the future, we will collect high-quality datasets and design new encoders to realize fine-grained multimodal understanding, e.g., supporting varying length inputs .
Appendix A Appendix Overview
Sec. C: Additional Implementation Details.
Appendix B Additional Ablation Experiments
In the main paper, we follow previous works and set a frozen CLIP-ViT as the universal encoder. Here we explore other design choices such as trainable CLIP-ViT and DINOv2 as the encoder.
We first turn on all the parameters in the multimodal-text alignment stage. As shown in Tab. 8, the performance for visual modalities (image and video) dropped significantly, while the result for audio QA (ClothoQA) improved by 4.7%. We think trainable CLIP will break the pretrained vision-language representations but can leave more space for learning other modalities. However, considering the memory usage (46Gb vs. 74Gb), frozen CLIP will be a better choice for our framework.
Beyond Vision-Language Encoder.
In addition to the vision-language encoder CLIP-ViT, we also explore other models, such as the self-supervised vision model DINOv2 , as the universal encoder. In Tab. 8, we noticed that the performance of OneLLM using DINOv2 is lower than the model using CLIP-ViT because DINOv2 is not aligned with language and we need to learn the vision-language alignment from scratch.
Appendix C Additional Implementation Details
The modality tokenizer is to transform input signal into a sequence of tokens. Here we will introduce the tokenizer of each modality in detail.
We use the same tokenizer setting for visual modalities, i.e., image, video, depth/normal map. The visual tokenizer is a single 2D convolution layer:
Audio Tokenizer.
Point Tokenizer.
IMU Tokenizer.
fMRI Tokenizer.
C.2 Multimodal-Text Alignment Dataset
We summary the multimodal-text alignment dataset in Tab. 9. For depth/normal-text pairs, we adopt DPT model pretrained on ominidata to generate depth/normal map. The source dataset is a subset of CC3M , around 0.5M image-text pairs. For IMU-text pairs, we use the IMU sensor data of Ego4D and the corresponding video narrations (i.e., text annotations). For fMRI-text pairs, we use the subj01 imaging session of NSD and follow the same data split with . Note that the visual stimulus, i.e., images shown to participants, are from MS COCO . Therefore, we use the image captions in COCO Captions as text annotations of fMRI-text pairs.
C.3 Multimodal Instruction Tuning Dataset
We summary the multimodal instruction tuning dataset in Tab. 9.
C.4 Prompt Design
The prompt formats for each dataset are shown in Tab. 10.
Appendix D Evaluation Details
In this section, we first list the evaluation prompts for each dataset in Tab. 11. Then we will give more evaluation details.
We evaluate all datasets using their official evaluation protocols. As shown in Tab. 11, for QA tasks with options, we ask OneLLM to directly predict the option letters; For open-ended QA tasks, we ask OneLLM to predict a single word or phase. For captioning tasks, we ask OneLLM to generate a one-sentence caption. Note that for audio-video-text tasks, the input sequence to the LLM is: {Video Tokens} {Audio Tokens} {Text Prompts}.
Point Cloud Tasks.
Our evaluation on point cloud tasks mainly follows PointLLM . For the point cloud classification task, we use the same prompt as PointLLM: What is this, and evaluate the accuracy using GPT4.
Depth/Normal Map Tasks.
For scene classification using depth/normal map, we first prepend the category list to the beginning of prompt, then we ask OneLLM to choose one class for the list.
IMU/fMRI Tasks.
We evaluate on IMU/fMRI captioning tasks. The prompts are the same as their training prompts: Describe the motion for IMU captioning and Describe the scene based on fMRI data for fMRI captioning.
Appendix E Comparison with Prior Works
The main difference between OneLLM and previous MLLMs is that we show a unified encoder is sufficient to align multi-modalities with LLMs. As shown in Tab. 12, OneLLM with one universal encoder, one projection module and less parameters (0.6B) can unify more modalities into one framework. The results in the main paper (Tab.1-6) also demonstrate that OneLLM can achieves better performance to previous works. The ablation experiments in Tab.7 (a) also show that jointly training all modalities with our unified framework can benefit data-scarce modalities. Here we are not trying to prove that OneLLM’s architecture is optimal, but to show the possibility of building MLLMs using a unified and scalable framework.
Appendix F Additional Qualitative Results
In this section, we provide more qualitative results in Fig. 4, Fig. 5 and Fig. 6.