PandaGPT: One Model To Instruction-Follow Them All

Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, Deng Cai

Introduction

Humans possess remarkable abilities to perceive and understand information from diverse sensory modalities, such as seeing a painting and hearing an audio guide. Analogously, to learn simultaneously, holistically, and directly from many different forms of information holds great promise for enabling machines to have a more comprehensive and better understanding of the world. To this end, there has been an emergent interest in developing artificial intelligence (AI) systems capable of perceiving and understanding information from multiple modalities simultaneously in a manner similar to humans.

However, much of the prior research has focused on tackling individual modalities in isolation. For instance, while significant progress has been made in text-to-image retrieval and generation , visually-grounded instruction following , and speech understanding and generation , these advances have largely been confined to separate combinations of text and other modalities or, at best, a few visual modalities (e.g., image and video). These models are limited in their ability to connect information from different modalities and lack the capacity to perceive and understand multimodal inputs holistically, thereby neglecting the inherent richness and complementary nature of multimodal data.

In this paper, we present PandaGPT, the first general-purpose model capable of instruction-following data from six modalities. PandaGPT leverages the power of multimodal encoders from ImageBind and the expressive language models from Vicuna , demonstrating impressive and emergent cross-modal capabilities across six modalities: image/video, text, audio, depth, thermal, and inertial measurement units (IMU). Crucially, PandaGPT achieves these capabilities despite being only trained on aligned image-text pairs, thanks to the shared embedding space provided by ImageBind.

This integration of multimodal information enables PandaGPT to perform a wide range of tasks, including generating detailed descriptions of images, composing engaging stories inspired by videos, and providing accurate answers to questions about audio inputs. Most interestingly, the core innovation of PandaGPT lies in its ability to naturally compose the semantics of multimodal inputs, which enables a rich set of compositional multimodal tasks across different modalities. For example, it can seamlessly connect the visual appearance of objects in a photo with their corresponding sounds in an audio clip, producing a cohesive and comprehensive understanding of the scene. These cross-modal capabilities empower the model to go beyond traditional unimodal analysis. We hope PandaGPT serves as an initial step toward building AGI that can perceive and understand inputs in different modalities holistically, as humans do.

Related Work

Large language models (LLMs) pre-trained over massive unlabeled text have dominated the field of natural language processing (NLP) today . With alignment techniques such as supervised instruction tuning and reinforcement learning from human feedback , LLMs exhibit surprisingly effective zero- and few-shot generalization abilities to perform almost any NLP tasks. The most successful examples could be OpenAI’s ChatGPT and GPT4 , which have made a profound impact on the entire AI research community and beyond. There also have been enormous open-source efforts to replicate the success, such as BLOOM , LLaMA , Alpaca , Vicuna , OpenAlpaca among many others.

Multi-modal Alignment.

Feature alignment among multiple modalities has attracted great interest for its applications such as cross-modal retrieval . Recently, CLIP learns a joint embedding space for image and text. Flamingo , BLIP-2 , and MAGIC bridge powerful pre-trained vision-only and language-only models and show strong zero-shot abilities. AudioCLIP adds audio into the CLIP framework for audio classification. ImageBind learn a joint embedding across six different modalities (image/video, text, audio, depth, thermal, and IMU data) using image-paired data only. More recently, there has been a surge of interest to combine multi-modal alignment and large language models for multi-modal instruction following. LLaVa , Mini-GPT4 , and Video-LLaMA enable visually-grounded instruction following. DetGPT proposes reasoning-based object detection. SpeechGPT adds speech understanding and generation abilities to LLMs. However, these advances have largely been confined to separate combinations of text and other modalities (e.g., image/video or audio).

Method

PandaGPT combines the multi-modal encoders from ImageBind and the large language models from Vicuna, achieving impressive capabilities in vision- and audio-grounded instruction following tasks. To align the feature space of multimodal encoders from ImageBind and large language models from VicunaWe use the version-0 of Vicuna-13B as our base language model., we train PandaGPT using 160k image-language instruction-following data released by and . Each training instance consists of an image I\mathcal{I} and a multi-turn conversation data (x1,y1,...,xn,yn)(\boldsymbol{x}_{1},\boldsymbol{y}_{1},...,\boldsymbol{x}_{n},\boldsymbol{y}_{n}), where xi\boldsymbol{x}_{i} and yi\boldsymbol{y}_{i} are the human’s instruction and the system’s response at the ii-th turn. To reduce the number of trainable parameters, we only train (i) a linear projection matrix ff to connect the representation produced by ImageBind to Vicuna; and (ii) additional LoRA weights on the Vicuna’s attention modules.The total number of trainable parameters is around 0.4% of the parameters of Vicuna. Figure 1 illustrates the architecture of PandaGPT.

The training objective of PandaGPT is defined as

where θf\theta_{f} and θl\theta_{l} correspond to the learnable parameters of the linear projection matrix and LoRA weights. The hIh_{\mathcal{I}} is the image representation produced by ImageBind and θ={θf,θl,θ1,θ2}\theta=\{\theta_{f},\theta_{l},\theta_{1},\theta_{2}\}, where θ1\theta_{1} and θ2\theta_{2} are frozen parameters of ImageBind and Vicuna. Note that the loss is only computed from the part of system responses during training. We train PandaGPT on the image-language instruction-following dataset for two epochs using a learning rate of 5e-4 with linear decay. The maximum sequence length for Vicuna-13B is set to 400 based on our computation resources (8×\timesA100 40G GPUs). The training takes around 7 hours to complete.

It is worth noting that the current version of PandaGPT is only trained with aligned image-text data. However, by leveraging the binding property across six modalities (image/video, text, audio, depth, thermal, and IMU) inherited from the frozen ImageBind encoders, PandaGPT demonstrates emergent, i.e. zero-shot, cross-modal capabilities across all of the modalities.

Capabilities of PandaGPT

Compared to existing multimodal instruction-following models trained individually for one particular modality, PandaGPT can understand and combine the information in different forms together, including image/video, text, audio, depth (3D), thermal (infrared radiation), and inertial measurement units (IMU) readings. We find that the capabilities of PandaGPT (see concrete examples in Section 6) include but are not limited to:

image/video-grounded question answering: see examples of Figure 2 , 3, and 4.

image/video-inspired creative writing: see examples of Figure 5.

visual and auditory reasoning: see examples of Figure 6, 7, and 8.

multimodal arithmetic: PandaGPT is also capable of working with input composed across modalities. By arithmetically adding information from different modalities as input, PandaGPT can produce results that reflect concepts from different parts. See Figure 9 and 10 for examples of image and audio arithmetic, and see Figure 11 and 12 for examples of video and audio arithmetic.

Limitations

Despite the amazing ability in handling multiple modalities and their combinations. There are multiple ways to further improve PandaGPT.

The training of PandaGPT can be enriched by using other alignment data, for instance, other modalities paired with text (e.g., audio-text pairs).

We only use one embedding vector for the content in other modalities than text, more research into fine-grained feature extraction such as cross-modal attention mechanisms could be beneficial to the performance.

PandaGPT currently only allows multimodal information to be used as input, future possibilities include generating richer multimedia content (e.g., creating images and response in audio).

New benchmarks to evaluate the composition ability of multimodal inputs is demanded.

PandaGPT can also exhibit several common deficiencies of existing language models, including hallucination, toxicity, and stereotypes.

Lastly, we would like to note that PandaGPT is a research prototype and cannot be readily used for real-world applications.

Examples

References