X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Artemis Panagopoulou, Le Xue, Ning Yu, Junnan Li, Dongxu Li, Shafiq Joty, Ran Xu, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles
Introduction
Humans inherently utilize multiple senses to interpret their surroundings and formulate decisions. By equipping artificial agents with cross-modal reasoning, Cross-modal reasoning is the ability to integrate and discriminate information from multiple modalities over text, in contrast to “multimodal reasoning,” traditionally reserved for vision-language tasks. we can foster the development of systems with a more comprehensive understanding of their environment, allowing them to discern patterns and make inferences that are not apparent when analyzing modalities separately. This has motivated the advent of Multimodal Language Models (MLMs) , which transfer the remarkable abilities of Large Language Models (LLMs) to static vision.
Recent advancements seek to extend the reasoning capabilities of MLMs by incorporating audio and video, either through the introduction of pre-trained cross-modal representations mapped through their transitive relation to images , training foundation models across multiple modalities or by training projection modules to align multimodalities to the representation space of LLMs . While effective, the former often necessitates task-specific fine-tuning and the latter refines models on joint-modality data which can be high-resource demanding in terms of collection and computational resources.
We present X-InstructBLIP github.com/artemisp/LAVIS-XInstructBLIP.git, a scalable and extendable framework (Figure 1) that enables learning with uni-modal data without constraints imposed by pre-training a universal cross-modal embedding space or the computational cost and potential over-fitting risks associated with unfreezing the LLM parameters. X-InstructBLIP seamlessly incorporates a multitude of modalities in an independent manner, eliminating the necessity for joint modality datasets while preserving the capacity to execute cross-modality tasks.
Our approach foregoes the need for extensive modality specific-pre-training and joint-modality data by utilizing the Q-Former module, initialized with image-text pre-trained weights from BLIP-2 and finetuned on unimodal datasets to map inputs from the separate modality embedding spaces to a frozen LLM.
Given the scarcity of instruction tuning data for a spectrum of modalities, we introduce a simple yet potent approach: a three-stage query data augmentation technique to leverage open-source LLMs to extract instruction-tuning data from captioning datasets.
Figure 2 shows illustrative results, highlighting the versatility of our framework. Quantitatively, X-InstructBLIP performs comparably to existing single modality models, and exhibits emergent competence in cross-modality tasks. To quantify and challenge this emergent ability we introduce DisCRn, an automatically curated Discriminatory Cross-modal Reasoning challenge dataset, which requires models to distinguish between diverse combinations of modalities, such as audio-video and 3D-image.
Our contributions are summarized as follows: (i) We present a simple and effective, scalable cross-modal framework to empower LLMs to handle a diverse range of tasks across a variety of modalities, without requiring modality-specific pre-training. Our results show that in spite of each modality (images, video, audio, and 3D) undergoing individual alignment to LLMs, our instruction-aware representations generalize to cross-modality tasks, enabling a seamless integration of more modalities. (ii) We introduce an approach for crafting instruction-tuning datasets from a variety of modalities, leveraging only readily available captioning data and open-source language models. Illustratively, we transform the established Cap3D and AudioCaps datasets into QA collections, resulting in 250k and 24k entries, respectively. (iii) We construct DisCRn, the first dataset designed for evaluating instruction-based cross-modal discriminative reasoning. It includes 36k examples across various modalities such as video, audio, 3D, and images; and introduces a novel task: to discern among various input modalities rather than integrating them, thereby providing a fresh and demanding benchmark in this intricate area of study.
Related Work
Cross-modal Language Models: Recent years have seen a surge in models capable of executing a spectrum of vision-language tasks, leading to the creation of Multimodal Language Models (MLMs). These models align multimodal inputs and the LLM’s latent space through various techniques, such as unified pre-training , vision-to-language alignment through textual feature extraction , vision-encoder optimization , and linear , transformer-based , or auto-encoder based projections . More relevant to this work are approaches that learn intermediate vision-informed language token representations either interleaved in LLM layers or as direct inputs .
Such approaches have been expanded beyond images to learning separate audio and video projections to pre-trained LLMs. In 3D, Hong et al. project point-cloud scenes to LLMs through rendered image features. Cross-modal approaches focus on joint audio-video such as unified pre-training and, more pertinently to our work, projection learning . In this work, we show that the multimodal instruction tuning framework from InstructBLIP can be effectively adapted to encompass video, audio, and 3D point cloud modalities.
Multimodal Multi-Input Language Tasks: The advancements in single input vision-language tasks have paved the way for the development of tasks necessitating models to concurrently reason about multiple non-linguistic inputs, such as engaging in spatial reasoning across multiple images , deliberating over a series of slides , responding to queries necessitating cross-modal reasoning across images and tables , or executing a range of instruction-based tasks involving multiple image inputs . Despite their complexity, these tasks predominantly operate within the realms of image-text modalities.
Even though cross-modal tasks exist, predominantly requiring models to reason over joint audio and video , there is a gap in the evaluation of models’ generative capabilities in reasoning about cross-modal inputs contrastively. While models are often optimized on contrastive objectives , even in cross-modal settings , their evaluation is predominantly confined to classification tasks or utilizing the contrastively learned representations for downstream tasks. To address this gap, we introduce DisCRn, a task requiring contrastive reasoning across cross-modal inputs in a practical open generation setting, providing a way to evaluate a model’s ability to translate features of various modalities from its internal representations to its generative output distribution.
Method
Figure 1 depicts an overview of the model’s architecture which extends the instruction-aware projection method introduced in Dai et al. to an arbitrary number of modalities through independent fine-tuning of modality specific Q-Former projections to a frozen LLM. Figure 3(a) illustrates the modality-to-LLM alignment procedure by highlighting all the components associated with each modality.
Algorithm 1 outlines the X-InstructBLIP alignment framework. In essence, it involves the following steps for each example of a text instruction paired with an extra-linguistic input: (1) Tokenization of the text instruction and embedding of the non-textual input using a frozen pre-trained encoder. (2) Feeding both the normalized encoding of the extra-linguistic input and the tokenized instruction into a Q-Former module, accompanied by a set of learnable query embeddings. (3) Transformation of these query embeddings by the Q-Former, through cross-attention layers in alternate layers of the transformer blocks to conditionally adapt to the inputs. (4) Projection of the modified query embeddings into the embedding space of the frozen LLM through a trainable linear layer.
Formally, for any given modality , a frozen pre-trained encoder embeds the raw input from the modality’s input space into an embedding space . This encoding process is represented as , where is an instance from the raw input of .
Each modality Q-Former transforms the query embeddings conditioned on both into instruction-aware language representations of . The Q-Former module consists of two transformer submodules that share the same self-attention layers: one submodule interacts with the output of the modality encoder and the other is a text transformer that serves as both an encoder and decoder, hence the use of the term tokens to refer to the learnable representation. Each Q-Former is initialized with the pre-trained weights from BLIP-2 , without the cross-attention layers due to a dimension mismatch between the image encoder in BLIP-2 and the other modality encoders. The modality embedding interacts with the instruction text and input query tokens via cross-attention layers inserted every other transformer block. The transformed result is termed output query tokens.
For sequential data, such as video and audio, we extract query tokens; each frame is encoded and then processed separately by the Q-Former module. For the sake of readability we omit the iteration over from the equations.
The output query tokens are linearly projected to the frozen LLM’s space through a learnable projection layer specific to each modality. This is necessary since the Q-Former’s tokenization space differs from that of the frozen pretrained LLM. Let h be the LLM’s tokenizer, E the LLM’s embedding layer, the modality cue, the example text input, and the corresponding target phrase. The resulting LLM input tokens is formed by
where denotes concatenation. The model is optimized under a causal language modeling objective :
is the cross entropy loss, the Q-Former parameters, is the target sequence, is the LLM’s prediction.
Note that the parameters of the LLM remain frozen during the alignment learning process and are therefore not a part of the optimization objective. We recognize the potential significance of investigating scenarios where the LLM parameters are subjected to either full or partial optimization. This could encompass both independent and joint optimization strategies with the Q-Former modules across diverse modalities. However, such an investigation falls outside the purview of our current study. Our focus is primarily on assessing the performance and adaptability of the Q-Former module in various modalities, specifically within the framework of using a frozen LLM. This approach strategically avoids issues related to catastrophic forgetting , both in the language and cross-modal contexts through the integration of new modalities.
Datasets
X-InstructBLIP is optimized and evaluated on a collection of pre-existing and automaticaly generated datasets succintly presented in figure 4, discussed in section 4.1, with more details available in the supplementary material. Section 5.3.3 introduces the Discriminatory Cross-modal Reasoning challenge dataset DisCRn The term discriminative reasoning, adapted from Xu et al. , refers to the ability to distinguish between sets of inputs, as opposed to joint reasoning, the synthesis of information from aligned sources. used to evaluate the emergent capabilities of X-InstructBLIP (section 4.2).
Instruction Data Augmentation: Extracting instruction-aware representations necessitates diverse instruction-related tasks across all modalities. Notably, datasets for 3D and audio modalities are marjorly caption-centric. To address this, we leverage the open-source large language model google/flan-t5-xxl to automatically generate question-answer pairs for the 3D and audio modalities from their corresponding captions. The process begins by prompting the model with captions to generate potential answers. These answers are then used to prompt the model to generate candidate questions. If the model’s response to a question, using the caption as context aligns closely with the initial answer (achieving a Levenshtein similarity score above 0.9), the example is added to our dataset. This procedure yields 250k examples using 3D data from Cap3D A subset consisting of 5k point clouds is deliberately held-out from Cap3D for the construction of DisCRn, as introduced in section 4.2. This exclusion is maintained both in the captioning and QA configurations. and 24k examples for audio data from AudioCaps . See the supplementary material for a breakdown of the data generation process and data distribution.
2 Discriminative Cross-modal Reasoning
X-InstructBLIP offers a distinct emergent capability: reasoning across different modalities, despite individual modality training. This highlights the model’s versatility and potential scalability across numerous modalities. To study this cross-modal reasoning capability, we present a discriminatory cross-modal reasoning (DisCRn) challenge dataset. As Figure 5 shows, this task requires the model to discern between the properties of two entities across modalities by selecting which one satisfies a queried property. This task mandates the model to not only discriminate the inherent characteristics of the involved modalities but also to consider their relative positioning in the input. This strategic imposition serves to diminish reliance on simplistic text-matching heuristics, order bias, or potential deceptive correlations between modalities.
To generate the dataset, we employed the same google/flan-t5-xxl model, utilized for instruction data augmentation. The process is initiated by prompting the language model in a Chain-of-Thought manner to generate a set of properties for each dataset instance. Each instance is then paired with a random entity from the dataset to form a (question, answer, explanation) triplet by prompting the language model with three in-context examples to leverage in-context-learning . A pivotal step in this creation process is a round-trip-consistency check: an example is integrated into the final dataset only when the model’s predictions on the generated question, given the captions, match the example answer, exhibiting a Levenshtein distance above 0.9. This refined dataset encompasses 8,802 audio-video samples sourced from the AudioCaps validation set, and 29,072 image-point cloud instances from a reserved subset of 5k point clouds from Cap3D . Each instance in the dataset is coupled with two representations corresponding to the captions: (audio, video) from AudioCaps and (point cloud, images) from Cap3D. Given that the arrangement of the data can be altered, this allows for maintaining a balanced set of answers. This balance pertains not only to the position of the answers but also to the answer modality. Human performance on the task stands at 90% indicating its high quality. More dataset creation and distribution details are in the supplementary material.
Experiments
We study the effectiveness of X-InstructBLIP as a comprehensive solution for incorporating cross-modality into pre-trained frozen LLMs. Following a debrief on the implementation details in section 5.1, we establish a set of comparisons to evaluate X-InstructBLIP’s effectiveness as a general framework (section 5.2). We also verify the model’s competitiveness in individual modality-to-text tasks (section 5.3.1). More notably, we explore X-InstructBLIP’s emergent cross-modality reasoning ability (sections 5.3.2 and 5.3.3) even in the absence of joint optimization.
X-InstructBLIP is built on the LAVIS library’s framework atop of the Vicuna v1.1 7b and 13b models . Each Q-Former optimizes 188M trainable parameters and learns query tokens with a a hidden dimension of size 768. Table 1 lists the frozen pre-trained encoders used for each modality. Raw inputs undergo standardized pre-processing prior to encoding: images are resized to resolution with random cropping and normalization; audio files undergo mono conversion and filter bank pre-processing followed by normalization as in Chen et al. over two 5-second frames; videos are uniformly sampled to 5 frames subject to the same pre-processing as images, and 3D point clouds are uniformly sampled and normalized to 8k points as in Salesforce , Xue et al. . All modality Q-Formers are pre-initialized with BLIP-2 stage-1 weights except for the video Q-Former which is initialized from the last iteration of the corresponding image Q-Former and optimized for 15k/5k steps for the Vicuna 7b and 13b models respectively.
We optimize the model on 8 A100 40GB GPUs using AdamW with parameters = 0.9 and = 0.999, and a weight decay of 0.05. The learning rate warms up linearly over the initial 1,000 steps from to , followed by a cosine decay to a minimum of 0. Consistent hyper-parameters and templates are maintained for the evaluated tasks, with small modifications tailored to each modality. A single best model is selected for each modality. See details in the supplementary material.
2 Baselines
Our primary aim is to demonstrate the adaptability of our framework across various modalities without relying on large-scale pre-training stages or joint modality data. To ensure our approach’s effectiveness and comparability, we juxtapose its performance to other methods that employ projections to pre-trained frozen or partially frozen LLMs, wherever possible. This serves as a mere sanity check, verifying that our method is both effective and competitive.
Linear Projection: Although Dai et al. have shown the benefits of instruction-aware Q-Formers over their base setup, there is limited prior work directly comparing Q-Formers with simpler architectures. To address this, we implemented a linear projection baseline to effectively replace the 3D and Audio Q-Formers. This projection maps the outputs of the 3D and Audio encoders to 32 outputs, otherwise maintaining the same training setup.
X-InstructBLIP : We also explore the effectiveness of the BLIP-2 initialization process by training the audio Q-Former in X-InstructBLIP (7b) using a random initialization approach. This experiment demonstrates improved performance, indicating that it’s possible to integrate new modalities into our framework without the necessity for extensive pre-training, since from the modalities considered, audio is the least likely to benefit from image-text pre-training. Future research should delve into the effects of modality-specific pre-training, as they are outside of our scope.
X-InstructBLIP †: Trained similarly to X-InstructBLIP (7b), with the distinction that the modality type is not prepended to the modality’s LLM input tokens before feeding into the LLM for training and inference.
3 Results
Across the paper we reserve the following notation in presenting the results: Underlined numbers indicate in-domain evaluations. Bold numbers indicate the top zero-shot performance. Blue numbers indicate the second best zero-shot performance. Models denoted with 7b and 13b indicate the underlying Vicuna model size. Gray shaded rows correspond to current work. Numbers in (parentheses) show improvement over the linear projection baseline. In terms of evaluation, we report CIDEr for captioning, and Top-1 accuracy for question answering and classification tasks.
We evaluate X-InstructBLIP’s performance across a range of single modality to text tasks, illustrating its versatility and efficacy across all four explored modalities. Tables 2, 3, 4, and 6 summarize X-InstructBLIP’s out-domain performance across 3D, audio, image, and silent video modalities.
3D: Table 2 shows the results on zero-shot classification on ModelNet40 under two setups: classification in closed vocabulary using loss ranking and open generation where the model is prompted to describe the 3d model and correctness is validated if a single class from the 40 candidates is present in the generation. X-InstructBLIP significantly outperforms the InstructBLIP baseline, which processes a single view rendering of the point cloud, achieving twice its accuracy.
This is further bolstered by Figure 6 where we employ TSne decomposition to visualize the ULIP-2, X-InstructBLIP and Linear Projection representations respectively. The Linear Projection seems to break the separation of the classes, a result supported by its lower performance of 16.0 and 19.5 points in classification and open generation accuracy compared to X-InstructBLIP.
Audio: In Table 3, X-InstructBLIP’s performance in audio classification, question answering, and captioning tasks is detailed, utilizing datasets from ESC50 , ClothoAQA , and Clotho . Classification is evaluated both in close (cls) and open generation settings, similar to 3D. X-InstructBLIP effectively distinguishes audio with high accuracy, surpassing ImageBind a multimodal contrastive method that aligns audio to multiple modalities though transitive correspondence to images by 10 points.
Notably, X-InstructBLIP trails by 6 points, suggesting initial challenges in audio class representation. However, this gap narrows in open generation and captioning tasks, indicating nuanced differences in model performance. The most significant improvement is observed in question answering, indicating that BLIP-2 weight initialization appears to enhance instruction awareness more than direct audio-language alignment, corroborated by the gap in closed vocabulary classification performance. Prefix Effect: Tables 2 and 3 show that cuing the model with the modality benefits even single modality performance. This is likely due to the fact that the Q-Former is relieved from the extra burden to encode the type of modality and instead reserves bandwidth for semantic information.
Image: While we do not expect large variations in performance compared to InstructBLIP, we present results on image captioning and visual question answering in Table 4 as a sanity check. X-InstructBLIP outperforms InstructBLIP on VizWiz , but we note a small drop in performance compared to InstructBLIP despite the similar fine-tuning setup. We hypothesize that this minor decrement in performance is attributed to the expanded template space which introduces a trade-off of generalization and performance. Table 5 compares performance between InstructBLIP (7b) and X-InstructBLIP (7b) on NoCaps , using prompts not encountered in the optimization of either model. While X-InstructBLIP exhibits some performance variability, it maintains more than half the standard deviation of InstructBLIP. This variance can be attributed to the expanded vocabulary in our templates, allowing the Q-Former to better associate an instruction with a specific task. For example, in the case of prompt P2: Provide a recap of what is happening in the picture", InstructBLIP maintains high performance as it closely resembles an in-domain prompt “Use a few words to illustrate what is happening in the picture". Note that the performance drop in InstructBLIP is mostly attributed to the language model resorting to generating longer descriptions when the Q-Former outputs have not captured the task, resulting in hallucinations in later stages of generation.
Silent Video: Table 6 evaluates X-InstructBLIP on out-of-domain video tasks. We compare performance with prominent baselines that rely on frozen or partially frozen LLMs and show comparable or improved performance on Video Question Answering (VQA).
However, due to the nature of the instruction aware setup, X-InstructBLIP is tuned on other QA tasks, thus having an advantage over VideoLLaMA and FrozenBiLM . As we show in the supplementary, the Video Q-Former component of X-InstructBLIP, initialized with the Image Q-Former’s weights, reaches convergence in performance remarkably fast, within about 1,000 iterations. This rapid convergence underscores the importance of fine-tuning the Video Q-Former with high-quality data. We observe that prioritizing MSRVTT data over the noisier WebVideo2M dataset is crucial, as the latter quickly leads to a decline in performance due to its less reliable captions. For context, VideoLLaMA is pre-trained on the entire WebVideo dataset of 2 million videos, in addition to BLIP-2 image pre-training. FrozenBiLM, on the other hand, undergoes two epochs of training on the larger WebVideo10M.
3.2 Cross-Modal Joint Reasoning
Despite each modality projection being trained individually, X-InstructBLIP shows strong joint modality reasoning. Table 7 demonstrates X-InstructBLIP’s capability to reason jointly over video (V) and audio (A). Notably, X-InstructBLIP (7b) is capable of synergizing inputs, displaying an improvement in performance compared to utilizing a single modality when the model is cued with the different modalities both in MusicAVQA and VATEX Captioning tasks. However, this behavior is not consistent for the model without prefix cues. Initially, it was theorized that the model’s inability to differentiate tokens corresponding to each modality, treating them instead as a continuous stream, might be the cause. However, the results from the image-3D crossmodal reasoning task where the prefix-less model outperforms the prefixed one by 10 points challenge this view. It appears that the inclusion of cues may be prompting the model to encode modality-specific information, which is beneficial in joint reasoning scenarios. This specialized encoding does not, however, prime the model to recognize and process characteristics usually associated with other modalities, required for enhanced performance in contrastive tasks. The underlying rationale is that the language model, already tuned to generate modality-relevant outputs, leads the Q-Former to primarily receive feedback on modality-specific generation during training. This mechanism could also account for the unexpectedly improved performance in individual modality tasks.
3.3 Cross-Modal Discriminative Reasoning
We assess the ability of X-InstructBLIP in executing discriminatory reasoning across different modalities using our newly introduced DisCRn benchmark, detailed in section 4.2. We frame the problem as a realistic open generation problem. The LLM is prefixed with the instruction:
You are given two inputs. Select exactly one of the two by reference to its relative position (first or second, left or right) that best answers the question.
In prompting X-InstructBLIP (7b) we found that using a Q-Former captioning prompt different from the comparative prompt provided to the LLM model induced a more general representation that was more applicable for the comparative task, as such we employ this approach for the results in Table 8. This is likely due to the lack of comparative data in fine-tuning since each modality Q-Former is trained separately. Future work can explore the effect of different prompts conditioned on different parameters in the instruction-aware training setup (e.g. data, templates, joint training, and LLM partial or full optimization). For the video-audio comparison, we select two frames for each modality to allow for a more balanced generation influence.
To benchmark our model’s capabilities, we incorporate a robust captioning baseline by substituting the query outputs with captions corresponding to the modalities using the Vicuna 7b model. The 13b model baseline yields below 5% accuracy, due to inability to produce short responses required for the task, hence is omitted. For images, 3D, and video modalities, we elicit captions by prompting InstructBLIP to Describe the image/video. For 3D inputs, a randomly chosen rendering view of the point cloud served to InstructBLIP. For video we follow Dai et al. and sample four frames and encode them as separate images whose output token representations are concatenated as input to the model. Audio is captioned using WavCaps .
X-InstructBLIP outperforms the audio-video and image-3D baselines by 3.2 and 7.7 points in accuracy respectively. Replacing one of the Q-Formers for with an equivalent linear projection module drops the performance to over half in image-3D and more than 10 points in audio-video.
It is worth noting that using a small sub-sample of the data we observed that the task is prompt sensitive, mainly in the language only setting. We leave it to future work to systematically evaluate the model’s ability on the task based on different prompts and in context examples.
Conclusion
This study introduces X-InstuctBLIP, a scalable framework for independently aligning multiple modalities to a frozen LLM demonstrating competitive results compared to leading methods across all addressed modalities. The framework exhibits emergent crossmodal reasoning, despite separate optimization of each modality. To test the model’s ability a new crossmodal discriminatory reasoning task DisCRn is introduced and show that the model can outperform strong captioning baselines across all four examined modalities. Despite the effectiveness of the method, the task remains an open challenge. We also find complexities and unanswered questions within each modality, paving the way for future explorations across and within modalities.
References
Data Generation
For the audio and 3D modalities, the available range of tasks for instruction tuning is relatively limited. To address this challenge we follow a common paradigm in the literature and extract question-answer pairs are derived from captioning datasets, specifically from captions consisting of 10 words or more. Figure 7 delineates the procedure to automatically generate question answering data from captioning datasets. The google/flan-t5-xxl model from huggingface-transformers is employed, and is prompted to produce candidate single-word answers based on the caption. Subsequently, the model is tasked with generating a relevant question using the answer and context as inputs. The method of round-trip-consistency is utilized to sift through and retain only those question-answer pairs that align with the context. This alignment is verified by ensuring that the Levenshtein partial similarity between the predicted and initial answers is greater than 0.90, calculated using the Fuzzy Wuzzy Python package. Subsequently, we apply a string matching post-processing to filter out instances that do not conform to the prescribed format. As a result, 250,070/1,157 suitable training/validation examples are derived from an initial 661,576/5,000 3D-caption samples from the Cap3D dataset, and 24,156/1,653 training/validation examples are derived from 38,695/1,900 original audio-caption samples from the AudioCaps dataset. Moreover, for 3D data, it is imperative to ensure that the question-answer pairs do not allude to color. This is due to the fact that the 3D encoder does not capture color characteristics. To achieve this, the language model is directed to reformulate the captions by omitting any references to color, prompted as: Rewrite the sentence {caption} by eliminating any color mentions, prior to implementing the round-trip-consistency check. A short human evaluation on 50 samples for each modality shows that 90% of the generated audio and 82% of the 3D data is correct. Table 9 presents a random sample of the generated data and table 10 provides an overview of the datasets’s distribution statistics. It is worth noting that the error cases are typically due to non-sensical questions, rather than wrong answers. For example the following pairs were marked as non-sensical: What is the sewing machine running at? speed, What does the steam whistle do? hisses, What is the 3D model of a brick wall with holes and stacked cubes, resembling? elements, and What is the hat with? pattern.
2 Cross-modal Discriminative Reasoning Data Generation
To assess the cross-modal reasoning capabilities of X-InstructBLIP, we devised a unique task that repurposes existing captioning datasets, specifically focusing on data representable in multiple modalities. We chose the AudioCaps validation dataset and reserved a subset of 5k examples from Cap3D as our validation dataset, ensuring that the 3D Q-Former is not exposed to this subset during the training phase in either captioning or 3DQA settings.
The audio data from AudioCaps originates from Youtube videos, allowing us to download the corresponding video files using their YouTube IDs. For Cap3D, we employed the associated point clouds and randomly selected one rendered image from the available eight view angles.
A depiction of the data generation procedure, also outlined in the main text, is provided in Figure 8. During the evaluation, we maintain a balance, ensuring each option (A or B) serves as the ground truth 50% of the time. Given that this problem is structured as an open vocabulary generation task, we expanded the ground truth answer space to include synonyms and equivalent expressions, such as [{answer modality}, left, 1st, 1, first, input 1, entity 1, object 1, input A, entity A, object A] and [{answer modality}, right, 2nd, second, input 2, entity 2, object 2, input B, entity B, object B], corresponding to whether the first or the second input is the ground truth. The human performance on a subsample of 100 examples of the dataset is 90%. Figure 5 in the main paper presents a sample of the generated data and table 11 provides an overview of the datasets’s distribution statistics.
Video Q-Former Fine-Tuning Versus Image Initialization
To explore the impact of further training the Image Q-Formers on video data, Table 12 presents the results of evaluating video tasks using the weights from the Image Q-Formers. It is evident that training on video data enhances performance. However, it’s worth noting that the Video Q-Formers reach convergence at an earlier stage (15k and 5k iterations for Vicuna7b and Vicuna13b, respectively). This is likely because the Q-Formers have already achieved semantic understanding during the image alignment phase, requiring minimal additional training to capture the nuances of sequential video projections. The higher drop in performance in MSVD captioning compared to VATEX is likely due to the closer similarity between MSVD and the held-in MSRVTT dataset distribution. There is a notably lower drop in performance for Video QA tasks, owing to the more constraint nature of the task - training on videos would not significantly increase the performance since the answer is typically constrained in one frame , and as such processing that frame would be almost equivalent to processing it in the image. The improvement probably stems from identifying the answer across a longer sequence of query tokens.
In-Domain Evaluations
Table 13 presents in-domain performance for a sample of datasets seen in training across all four modalities. It’s important to clarify that when we refer to ‘in-domain,’ we are specifically referring to datasets that were sampled during the training process. However, it’s crucial to note that this does not constitute explicit fine-tuning, as there is no guarantee that the Q-Former has encountered the entirety of the dataset during its training.
Training Details
Table 14 compiles the training hyperparameters employed for each modality and model. The X-InstructBLIP†variant is trained similarly to X-InstructBLIP, with the notable distinction that the modality type is not prepended to the modality’s query outputs, both during training and inference. Following Dai et al. that noted that sampling ratios play an important role in training we perform some minor modifications in the sampling ratios that we show in tables 16 and 15 are effective in improving performance. The decisions are discussed further below. It is worth noting that due to the large amount of experiments consisting of all modalities, we did not exhaust all possibilities, and there may be better training configurations. We leave this to future work to be explored.
As each modality exhibits unique characteristics, we have customized the training approach for each one. For instance, the 3D and Audio Q-Formers are trained for the maximum number of iterations specified in Table 14.
The Vicuna7b Image Q-Former undergoes training for 735k iterations, utilizing normalized data sampling. Additionally, an extra 40k iterations are performed with the sampling ratio of COCO Captions set to 3.0 while keeping the other ratios consistent with the original sampling. This adjustment leverages the clean annotations of COCO Captions, mitigating noise introduced by larger image datasets. However, this upsampling technique is not applied to the Vicuna13b Image Q-Former, since it appears to lower out of distribution performance in non-captioning tasks as shown in table 16. It could be that due to the smaller batch size, Vicuna13b is less sensitive to noisy data, since it effectively sees less of them. In both cases, the last checkpoint from the iterations specified in Table 14 is chosen, with guidance from the COCO Captions validation dataset. Note that we optimize the Image Q-Former for 10 times more iterrations that InstructBLIP. The reason is that we maintain conformity with the other Q-Formers and do not intitialize the cross-attention layers from BLIP-2 pretraining nor we allow for stage-2 training. Nevertheless, we show that with enough iterations, the cross attention layers can be learned equivalently without the need of the contrastive auxiliary losses of BLIP-2 nor stage-2 training.
The Vicuna7b video Q-Former is initialized from the best Vicuna7b image Q-Former and undergoes validation every 5k iterations on the MSRVTT captioning dataset. The selection process involves choosing the checkpoint that precedes any drop in performance during the subsequent validation rounds even if there is a better performing checkpoint later on in training, to avoid overfitting to the MSRVTT skeletal captions. Table 15 quantitatively shows our observations. Due to the initialization of the video Q-Former with the well trained image Q-Former, the noisy captions of WebVid2M reduce the performance instead of improving it. However, this is corrected when the training consists of cleaner data.
Similarly, the Vicuna13b video Q-Former is initialized from the best checkpoint of the Vicuna13b Image Q-Former and validated every 1k iterations. While we let the Vicuna7b and 13b video Q-Formers train for 15k and 25k respectively, we observe early convergence at 15k and 5k iterations likely due to the pre-initialization with the Image Q-Former. During training, 5 frames are sampled for the Vicuna7b Video Q-Former, while 4 frames are sampled for the Vicuna13b to reduce computational demands. Figure 9 shows that the video performance converges in as little as 1k iterations on an out of domain video captioning dataset.
The best training approach for each model was empirically identified, and it is beyond the scope of the paper to rigorously analyze the reasons of the differences in training across modalities. We leave this to future work.
Evaluation Hyperparameters
During the evaluation of X-InstructBLIP, we adhere to a consistent set of hyperparameters, with minor variations to accommodate the distinct needs of each task. A comprehensive list of these configurations is presented in table 17. In every experiment, we utilize Beam Search for generation, setting the beam size to 5, repetition penalty and temperature equal to 1.5 and 1 respectively. For tasks involving contrastive reasoning across video-audio modalities, a balanced representation and computational efficiency are achieved by querying two frames from both video and audio. The length penalty is typically configured to 1 for long caption tasks, -1 for Visual Question Answering (VQA) tasks requiring short answers, and 0 for short caption tasks. The minimum and maximum length constraints are adapted based on the task: for captions, we maintain a range of 10 to 80; for short-answer VQA tasks, the range is set from 1 to 10; for variable-length captions, the range is between 1 and 80. In the case of the InstructBLIP baseline for video datasets, we borrow the recommended inference setup of sampling 4 frames for the captioning baselines of MSVD and VATEX with the prompt A video that shows and the same generation hyperparameters as X-InstructBLIP.
All our evaluations are reported with batch size 1 due to inconsistencies observed in Huggingface implementations of models with rotary position embeddings when evaluated in batch.
Instruction Tuning Suite
Table 18 presents a comprehensive list of datasets employed in the instruction tuning process for X-InstructBLIP, accompanied by their corresponding dataset sizes. Datasets labeled with ∗∗ have been generated automatically through the round-trip-consistency procedure detailed in section 4.1 of the main paper, with further information provided in Appendix 7. Datasets marked with indicate instances of data loss resulting from file corruption or expired links. Video with id f9_bP219ehQ_63_70 is corrupt and could not be retrieved.
Prompt Templates
X-InstructBLIP has undergone fine-tuning using a diverse array of instruction templates, tailored to cover a wide spectrum of tasks and modalities. For reference, the specific templates corresponding to each modality can be found in the following tables: Table 19 for images, Table 20 for audio, Table 21 for 3D, and Table 22 for videos. Compared to InstructBLIP caption templates have increased from 13 to 32, while question-answering templates have grown from 10 to 21. These enhancements have been strategically incorporated to foster greater adaptability of the model to a wide range of user instructions.
Ethics Statement
In this research, we present a framework for aligning multiple modalities with a frozen large language model (LLM). Our methodology strictly involves the use of publicly available and free datasets, ensuring we do not engage in the collection of private data. However, it is crucial to acknowledge that publicly sourced datasets carry implicit biases . These biases reflect historical and societal inequalities, potentially influencing the model’s outputs. Our framework builds upon a pre-existing frozen LLM. While this approach benefits from the extensive knowledge encoded within the LLM, it is important to recognize that such models can inherently propagate biases present in their training data. Additionally, there is a non-negligible risk of generating false or misleading information. While there exist tools to measure language model toxicity such as Helm , their evaluation datasets are constrained in the language modality, and hence are not applicable to measure toxicity across modalities which is the focus of this work. We leave the generation of cross-modal datasets for toxicity and bias measurement as a future research direction.
Users of our framework should be aware of these limitations and exercise caution, particularly in applications where the accuracy and impartiality of outputs are critical. We advocate for responsible use of our framework, especially in sensitive contexts. Users should critically assess and verify the model’s outputs and consider the potential for reinforcing biases or spreading misinformation. Furthermore, we commit to transparency regarding our model’s capabilities and limitations. All code, data, and model weights will be released to ensure reproducibility and encourage external evaluation and subsequent research.