Emu: Generative Pretraining in Multimodality

Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, Xinlong Wang

Introduction

With text corpus at massive scale, Large Language Models (LLMs) with straightforward training objectives such as next-word-prediction learn to understand, reason, and generate text with unprecedented accuracy and fluency, paving the way for diverse real-life applications unthinkable a decade ago. Recent studies have investigated Large Multimodal Models (LMMs) beyond LLMs. Flamingo , which connects a powerful language model with a pretrained vision encoder and inserts learnable layers to capture cross-modality dependencies, demonstrates strong abilities in multimodal zero-shot and in-context learning. Recent works also adopt this framework and build LMM by docking a vision encoder with an LLM.

Effective as they are, these LMMs are mostly trained on image-text pairs or documents, while overlooking video data as another scalable source of interleaved multimodal data. Besides, the commonly used training objective in such LMMs is predicting the next text token , typically with a frozen vision encoder and no supervision for the vision part, which highly restricts the model’s capacity. In this work, we introduce Emu, a large multimodal model that learns from both video and image data interleaved with text, under a unified objective of predicting the next visual or text token in an autoregressive fashion.

Documents interleaved with images (e.g., textbooks, webpages) provide an intuitive representation of complex concepts, and have proved to be effective in empowering models with multimodal in-context learning ability . Videos, which usually contain interleaved image frames and subtitles (Figure 3), are an abundant source of multimodal data that has been largely overlooked. They naturally contain dense visual signals and encode stronger cross-modal correlations with text than regular multimedia documents. Furthermore, public videos (especially user-generated clips) possess richer content diversity than Common Crawlhttps://commoncrawl.org/, from which current training datasets mainly originate.

To take advantage of rich web-scale data with omnivore capacity, we formulate diverse sources of interleaved multimodal data (e.g., videos with subtitles, webpages with images and text) into a unified format of interleaved image embeddings and text tokens (videos are converted into randomly-selected frames and subtitles interleaved into a sequence). Specifically, visual signals are first encoded into embeddings via a visual representation model EVA-CLIP , instead of being converted into discrete tokens. These visual embeddings together with text tokens constitute an interleaved multimodal input sequence.

Pretrained with the unified objective and diverse forms of data stated above, Emu can serve as a generalist interface for both image-to-text and text-to-image tasks by performing various types of completion in a multimodal sequence, i.e., accepting multimodal prompts (e.g., text, images, video, or their interleaved sequence) and outputting multimodal response (for image generation, visual embeddings are decoded by a fine-tuned diffusion model), as illustrated in Figure 1. Further, Emu demonstrates impressive abilities such as in-context text and image generation (the 2nd block of Figure 1), image blending (the 5th row of Figure 1 that combines a cat and a tiger into a cute tiger-cat), video understanding (the last block of Figure 1), and real-world knowledge grounding (Section 5.3).

We evaluate Emu on a broad range of zero-shot and few-shot tasks including image captioning, visual question answering, video question answering, and text-to-image generation. For qualitative demonstration, we also build an effective multimodal assistant via instruction tuning on multimodal conversation data. The instruction-tuned Emu assistant can effectively follow human instructions and interact with users via multimodal response.

Emu: Predict the Next in Multimodality

Emu is a large-scale multimodal model that performs completion in multimodality, i.e., perceiving interleaved multimodal input and generating outputs varying in modalities. As illustrated in Figure 2, Emu consists of four parts: Visual Encoder, Causal Transformer, Multimodal Modeling, and Visual Decoder. We leverage pretrained EVA-CLIP , LLaMA and Stable Diffusion to initialize the Visual Encoder, the Multimodal Modeling LLM and the Visual Decoder, respectively.

Given any sequence with interleaved image, text and video, we first encode the image into dense visual features via EVA-CLIP, then transform the encodings into a fixed number of NN visual causal embeddings via Casual Transformer. Similarly, we encode a video of TT frames into T×NT\times N visual causal embeddings. Two special image tokens [IMG] and [/IMG] are prepended and appended for each image or frame, respectively, to represent the beginning and end of the encoded image/frame embeddings. The visual causal embeddings are combined with text tokens to form multimodal sequences that are fed into the Multimodal Modeling LLM for unified autoregressive modeling. We append and tokens to the start and the end of each sequence. In inference, we fine-tune the Visual Decoder to decode the visual embeddings into a realistic image.

Causal Image-text Transformer. Auto-regressively modeling images in raster order is counter-intuitive and has not demonstrated satisfactory performance, which may be attributed to the fact that images naturally possess 2D structures and are not perceived as sequential signals like text. To better capture the characteristics of images and achieve unified modeling of different modalities, we propose a Causal Transformer module to transform 2D spatial visual signals to 1D causal sequences in a latent space ZZ. Specifically, given an image II with its encodings g(I)g(I) from EVA-CLIP, Causal Transformer accepts randomly initialized embeddings {e1,e2,…,eN}\{e_{1},e_{2},\dots,e_{N}\} as input, and outputs NN embeddings {z1,z2,…,zN}\{z_{1},z_{2},\dots,z_{N}\} that capture the causal dependency of the given image:

The architecture of Causal Transformer is similar to the decoder of Transformer , with each block consisting of a causal self-attention layer, a cross-attention layer, and a feed-forward layer. Different from Q-Former that captures bi-directional relations of input tokens, we use a causal self-attention layer to capture the causal dependency among the input latent embeddings for further unified causal modeling of vision and language modalities. The cross-attention layer aggregates visual information from the image embeddings extracted from EVA-CLIP, where the visual embeddings are treated as keys and values, and the outputs from the previous causal attention layer serve as queries.

Visual Decoder. We use a latent diffusion model to decode visual embeddings into images, and adopt the weights of Stable Diffusion as initialization. Specifically, we feed NN visual embeddings generated by Emu into the diffusion model as conditions for image decoding. We replace the linear projections of the cross-attention modules in Stable Diffusion with new linear layers that accommodate the dimension of Emu and Stable Diffusion.

2 Training Objective

Given an unlabeled web-scale corpora D\mathcal{D} consisting of interleaved multimodal sequences x=(x1,x2,…,xn)x=(x_{1},x_{2},\dots,x_{n}), where xx can be vision-language sequences of various forms, such as image-text pairs, image-text interleaved documents, or videos with subtitles. xix_{i} can be a signal unit (text or image token) from any arbitrary modality. We first convert all continuous 2D signals (images and video frames) into 1D causal latent embedding sequences using Causal Transformer, then insert them back into the corresponding places in the sequence xx. The resulting sequence is represented as u=(u1,u2,…,um)u=(u_{1},u_{2},\dots,u_{m}), where uiu_{i} can be either a discrete text token, or a visual embedding that captures causal dependency with neighboring visual embeddings.

We approximate the likelihood of the web-scale corpora p(x)p(x) with p(u)p(u), and maximize the likelihood in a unified auto-regressive manner as follows:

3 Generalist Interface

The unified auto-regressive modeling of different modalities endows Emu with a powerful ability to serve as a multimodal generalist that can perform many types of completion in a multimodal sequence, i.e., accepting multimodal sequence as input, and outputting signals across vision and language modalities. For example, when using two image-text pairs of the same task as the prompt, Emu automatically infers and completes the corresponding task given a new input, as shown in the second block of Figure 1.

Specifically, given a multimodal context, if the expected output format is text, Emu will use the language modeling head to generate discrete text tokens. If the desired output is image, we will append a [IMG] token at the end of the input sequence, then Emu will autoregressively generate NN visual embeddings that will then be sent to the visual decoder for decoding into a real-world image.

Emu Training

We pretrain Emu with web-scale data across modalities in various forms, including image-text pairs (LAION-2B , LAION-COCO ), interleaved images-text data (MMC4 ), video-text pairs (WebVid-10M ), and our collected interleaved video-text data (YT-Storyboard-1B). All these data are formulated as multimodal sequences, from which Emu learns under the objective of predict-the-next-element in a unified auto-regressive manner. After pretraining, we finetune an Image Decoder to transform visual embeddings into realistic images.

Image-text Pairs. We use the image-text pairs from LAION-2B and LAION-COCO for pretraining. LAION-2B provides images paired with noisy alt-texts from the web, and LAION-COCO is its 600M subset that is captioned by BLIP .

Video-text Pairs. WebVid-10M is an extensive dataset consisting of a large collection of short videos with textual descriptions. These videos are sourced from materials websites with diverse contents and a strong correlation between text and video. We use heuristic rules to remove irrelevant metadata (e.g.resolution of the original video, camera parameters).

Interleaved Image and Text. Large-scale image-text interleaved data plays a crucial role in unlocking the in-context learning ability of multimodal models. We leverage the Multimodal-C4 (MMC4) dataset , an expanded version of the text-only C4 . Multimodal-C4 comprises a collection of approximately 75 million image-text-interleaved documents, with 400 million images and 38 billion tokens in total. From each document, we sample a random subsequence of L = 1024 take up to the first N = 5 images included in the sampled sequence. Additionally, we randomly sample N = 5 images along with their corresponding sentences to construct a subsequence of L = 512.

Interleaved Video and Text. Videos with subtitles also present a promising and scalable source of interleaved multimodal data. We introduce the YT-Storyboard-1B dataset which collects 18 million videos and their corresponding subtitles from YouTubehttps://www.youtube.com using the video-ids provided by the YT-Temporal-1B dataset . Instead of raw videos, we collect storyboard images (about 1.8 billion images in total), a set of thumbnails provided by the YouTube website for quick video viewing. The combination of storyboard thumbnails and subtitles creates a natural interleaved sequence of video and text ordered by timestamps. An example is provided in Figure 3.

More details about the pretraining datasets are deferred to Appendix A.1.1.

2 Pretraining

We initialize Emu’s Visual Encoder with the 1B version of EVA-02-CLIP , and Multimodal Modeling LLM with the 13B version of LLaMA . LLaMA is a decoder-only Transformer and EVA-02-CLIP is a 40-layer ViT . The Causal Transformer comprises 12 blocks, each of which consists of a causal self-attention layer, a cross-attention layer, and a feed-forward layer. Random initialization is used for Causal Transformer. The total number of parameters of Emu is 14B and is trained end-to-end.

We use a batch size of 128 for image-text pair data, 64 for interleaved image-text data, 16 for video-text pair and interleaved video-text data. We adopt the AdamW optimizer with β1\beta_{1} = 0.9, β2\beta_{2} = 0.98, and a weight decay of 0.05. We use a cosine learning rate decay with a peak learning rate of 1e-4 for the Causal Transformer, 3e-5 for LLaMA and 5e-5 for EVA-02-CLIP , and a linear warmup of 2k steps. For each video, we randomly sample 8 frames for pretraining, and all images/frames are resized into 224×\times224 resolution. For image-text pair and interleaved data, we randomly put each image before or after its corresponding sentence. We train the model on 128 NVIDIA 80G-A100 GPUs for 10k steps with around 82M samples (150B tokens in total), and the pretraining takes approximately 2 days.

3 Visual Decoding

After pretraining, we tune the visual decoder with both LAION-COCO and LAION-Aesthetics (a high-aesthetics quality subset of LAION-5B ) image-text pair datasets under text-to-image task. Specifically, We initialize the diffusion model with Stable Diffusion v1.5. We freeze the Visual Encoder, Multimodal Modeling LLM in Emu, and the VAE in diffusion model during training, with only the parameters of U-Net updated. For each training sample, we append the [IMG] token to the end of the input text and feed it into the Multimodal Modeling LLM, which will then generate NN visual embeddings in an auto-regressive manner. These visual causal embeddings are fed into Image Decoder as the condition for image generation training.

We follow the model setups of Stable Diffusion v1.5. We employ AdamW optimizer with β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999 and the weight decay of 1e-2. We train the diffusion model with 32 A100-40G GPUs for 15k iterations. The batch size is set to 50 per GPU, and the learning rate warms up to 1e-4 for the first 5k steps, then decreases to 5e-5 and 1e-5 at 10k and 14k steps respectively. To further improve sample quality, we randomly drop image embeddings condition by 10%10\% of the time during training to enable classifier-free guidance . Please refer to Appendix A.2 for more training details.

Instruction Tuning

Language instruction tuning has helped pretrained language models to align with user intentions and generalize to unseen tasks . We apply multimodal instruction tuning on Emu to align it with human instructions through supervised finetuning on publicly available datasets, including language instructions from ShareGPT and Alpaca , image-text instructions from LLaVA , and video instructions from VideoChat and Video-ChatGPT . Dataset details can be found in Appendix B.1.

In instruction tuning, we freeze all parameters of pretrained Emu, and fine-tune a low-rank adaption (LoRA) module . The main focus of instruction tuning is to align the model with natural language instructions, which are less relevant to vision features. Thus, we attach LoRA modules only to the self-attention layers of the Multimodal Modeling LLM, and add no adaptation to the Vision Encoder. We use a batch size of 128 and train for 10k steps. The learning rate linearly warms up to 1e-5 in the first 500 steps, then decays to zero with a cosine schedule. The overall instruction tuning phase takes around 16 hours with 16 A100-80G GPUs.

All instruction-tuning data are packed with this template:

where [USER] and [ASSISTANT] are special tokens initialized from the embeddings of words ‘user’ and ‘assistant’, respectively. varies depending on the specific task, and detailed system messages used for different types of tasks can be found in Appendix B.2. and are actual slots for human instructions and assistant answers, and only is accounted for loss computation.

Evaluation

We evaluate Emu on a broad range of vision-language tasks including image captioning (MS-COCO ), image question answering (VQAv2 , OKVQA , VizWiz ), visual dialog (VisDial ), video question answering (MSRVTTQA , MSVDQA , NextQA ) and text2image generation(MS-COCO). Details of these benchmarks are described in Appendix C.1. We evaluate our pretrained and instruction-tuned models in zero-shot and few-shot settings.

In the zero-shot setting, the model is tested on tasks and datasets it has never encountered during training. Task-specific prompts are used to indicate different tasks to perform, without any additional tuning for model parameters.

Multimodal Understanding. Table 1 presents the zero-shot multimodal understanding performance of Emu and Emu-I (the instruction-tuned model). We adopted the multimodal Chain-of-Thought prompting technique on the pretrained model following . This approach involves two steps: first asking the model to generate a descriptive caption for visual content, then providing the model with both the generated caption and a task-specific prompt to output the final result. Additionally, to ensure a fair comparison with Flamingo , we also evaluate using the same prompting strategy of Flamingo. These results are obtained by using two text-only examples from the task as prompts. Results evaluated under this strategy are indicated by an asterisk (*). Note that these prompts do not include any images, simulating a few-shot text prompt approach. For more detailed information regarding the evaluation, please refer to Appendix C.2.

On COCO captioning task, Emu achieves impressive zero-shot CIDEr score of 112.4, which outperforms other LMMs by a large margin. In a wide range of image and video question answering tasks, Emu consistently surpasses LMMs like Kosmos-1 and Flamingo-9B. Notably, Emu achieves an accuracy of 34.4% on the complex VizWiz VQA dataset, versus Kosmos-1’s 29.2% and Flamingo-9B’s 28.8%. Emu-I is the instruction-tuned Emu model that achieves notable improvements. Remarkably, even with only 14B parameters, Emu-I can outperform much larger-scale Flamingo-80B model in several tasks such as VQAv2 (57.5% vs. 56.3%), VizWiz (38.1% vs. 31.6%), and MSVDQA (36.4% vs. 35.6%).

We evaluate the zero-shot image generation ability on the validation set of MS-COCO . Following , we randomly sample 30k prompts from the validation set and calculate the zero-shot FID . The results are shown in Table 2. For the generation of both Emu and SDv1.5, we use PNDM scheduler with 50 steps. We also adopt classifier-free guidance for better generation quality. The scaling factor is set to 5.0 and 3.0 for Emu and SDv1.5 respectively, as these settings yield the best performance for both models. Emu achieves better performance compared to a concurrent work GILL , which also generates images with LLMs. However, our model is inferior to SDv1.5 in terms of FID. This is probably because the condition space (image embeddings) of our visual decoder deviates a lot from the condition space (text embeddings) of the diffusion model used as initialization, and our model is trained for a relatively short 15k steps. We believe there might be room to improve via fine-tuning with more steps, or using another visual decoder instead of adopting pretrained diffusion models that condition on text embeddings.

2 Few-shot Evaluation

In few-shot evaluation, the model is prompted with task-specific prompts and a small number of examples collected from the training data to evaluate its in-context learning ability. Evaluation details can be found in Appendix C.3. Table 3 presents the performance of the pretraining model Emu in image and video question answering tasks under the few-shot (k=2,4,8k=2,4,8) evaluation setting. We use the Retrieval In-Context Example Selection (RICES) approach employed in Flamingo . With interleaved data incorporated in the pretraining phase, Emu demonstrates superior performance to Flamingo-9B and Kosmos-1 under almost all scenarios. For example, Emu achieves a VQAv2 accuracy of 58.4% and VizWiz 41.3% under the 4-shot setting, surpassing Flamingo-9B by +2.1% and +6.4%, respectively. For video-text tasks, Emu demonstrates strong performance as well, such as 4-shot 21.8% v.s. Flamingo’s 18.2% on the MSRVTTQA benchmark. Additionally, we can observe a positive correlation between the number of shots kk (k=0,2,4,8k=0,2,4,8) and the performance of Emu. These results demonstrate Emu’s remarkable in-context learning ability.

3 Qualitative Evaluation

Beyond quantitative benchmarks, we conduct adequate qualitative evaluation of Emu. Emu demonstrates impressive capabilities that cannot be evaluated on standard benchmarks, including real-world knowledge grounding (upper right of Figure 4), interleaved multi-image understanding (left side of Figure 4), detailed video understanding (lower right of Figure 4), multimodal assistant (Figure 5), multi-turn dialogue (Figure 6), image blending (Figure 7), and (in-context) text-to-image generation. For in-context text-to-image generation, Emu can generate context-related images (in the first two rows of Figure 8, the generated images share the oil painting style in context, compared with the corresponding images generated without context in the first two rows of Figure 9), and follow context-related instructions, as shown in the 4th row of Figure 1. The in-context ability of the multimodal modeling of Emu (LLM as initialization) is responsible for this brand-new ability of image generation.

We also compare Emu with other state-of-the-art multimodal assistants in terms of the ability to perform typical image captioning tasks (Figure 10) and follow human instructions (Figure 11). In Figure 11, we test a slightly difficult instruction, and only Emu response properly to list 8 books written by Agatha Christie and then recommend one.

Related Work

Multimodal pretraining learns cross-modal interactions from large-scale multimodal data. BEiT series convert visual signals into discrete tokens that can be pretrained same as language, and BEiT-3 achieves exceptional fine-tuning performance with a unified BERT-style masked signal modeling objective. Flamingo bridges powerful yet private pretrained vision and large language models and first demonstrates remarkable multimodal zero-shot and few-shot behaviors. With the increasing impact and accessability of LLMs, recent work has also considered building multimodal models based on LLMs , such as BLIP-series that connect frozen vision and language pretrained models with a Q-Former to bridge the modality gap. These LMMs commonly use predicting the next text token as the training objective and exert no supervision for vision data . Instead, Emu unifies the modeling of vision and language with the objective of predicting the next visual or text token in an autoregressive manner, and further explores videos as a new source of interleaved image-text data. This unified modeling leads to a generalist interface for diverse multimodal tasks that output either image or text. Emerging recent studies attempt to build powerful visual multimodal assistants based on LMMs through constructed conversation data. We also instruction-tune Emu using publicly available datasets and build a multimodal assistant that aligns well with human instructions on both images and videos.

Conclusion

In this work, we present Emu, a Large Multimodal Model (LMM) trained with a unified autoregressive objective of predicting the next element, including both visual and textual tokens. Apart from commonly used image-text pairs and interleaved documents, we explore another scalable data source of image-text interleaved data, i.e., video. Emu trained under such unified objective and diverse data can serve as a generalist interface that is capable of performing diverse multimodal tasks, such as image captioning, image/video question answering, and text-to-image generation, together with new abilities like in-context text and image generation, and image blending. We also build a multimodal assistant instruction-tuned on Emu, which exhibits excellent human-aligned abilities such as multi-turn dialogue. We hope that our work will inspire the community to continue exploring the potential of diverse multimodal data at the web-scale and also the generative pretraining beyond vision and language.

Acknowledgement

We would like to thank Hanxiao Qu, Quanyue Ma, Teng Dai, Yemin Shi, Wenhao Huang, Yue Cao, as well as other colleagues at BAAI for their support to this project.

References

Appendix A Emu training

Image-text Pairs. The LAION-2B dataset is the english subset of Laion-5B and contains large-scale image-text pairs data. LAION-COCO is captioned 600M images from LAION-2B with an ensemble of BLIP and CLIP models. Whereas the text in LAION-COCO exhibits enhanced fluency and relevance to the associated images, it has insufficient text diversity and a potential loss of high-level semantic information, including world knowledge contents presented in the original LAION-2B dataset. Thus, we employ both the LAION-2B and LAION-COCO datasets during Emu pretraining.

Video-text Pairs. Webvid-10M dataset contains a diversity of content with strong correlation between text and video. However, we found that a certain amount of the data contained irrelevant metadata information (e.g.resolution of the original video, camera parameters). To prevent the model from being influenced by these irrelevant details, we use heuristic rules to remove these content. Firstly, we build a word list consisting of irrelevant information. This word list is then utilized as a filtering mechanism to process the raw video text descriptions obtained from the original dataset. As a result, approximately 1 million datasets requiring cleaning are identified. Subsequently, specific rules are designed based on this list to identify and eliminate any words of irrelevant information within the text. Finally, the cleaned texts are subjected to rewriting using the Vicuna-13B , thereby ensuring its fluency and enhancing the overall quality.

Interleaved Image and Text. Multimodal-C4 is used as interleaved image-text data in pretraining. Following OpenFlamingo, we filter images based on CLIP similarity score to ensure the relevance of the images and text in each document. Specifically, any image with a CLIP similarity score below the threshold of 0.32 for all text in the same document is discarded. From each document, we sample a random subsequence of L = 1024 and take up to the first N = 5 images included in the sampled sequence. This process results in long text with the inclusion of multiple images. Additionally, we randomly sample N = 5 images along with their corresponding sentences to construct a subsequence of L = 512. This approach yields N = 5 image-text pairs.

Interleaved Video and Text. Videos with interleaved subtitles text represent a valuable and scalable source of multimodal data that has received limited attention thus far. In our study, we introduced YT-Storyboard-1B dataset, which collected storyboard images from YouTube, utilizing the video-ids provided by the YT-Temporal-1B dataset, which encompasses a vast collection of 18 million videos, equating to a total of 1.8 billion storyboard images. Specifically, for each video, we crawl the storyboard images and subtitles files directly. Where the sampling time between storyboard images is fixed, so the start time of each image can be determined by the order. Subtitle files record the content of each subtitle, as well as the start and end times. Therefore, storyboard images and subtitles can be sorted according to their timestamps and adjacent subtitles can be merged to form an interleaved video-text sequence. By opting to collect storyboard images instead of raw video data, we eliminate the necessity of video decoding. Moreover, this approach leads to a substantial 20-fold reduction in data storage costs, resulting in increased download efficiency.

A.1.2 Training Details

We report the detailed training hyperparameters settings of Emu during the pretraining in Table 4.

A.2 Visual Decoding

LAION-Aesthetics is the subset of LAION-5B which have relatively high aesthetics quality while LAION-COCO has relatively high image-text correlation. To empower the visual decoder to possess the ability of decoding visual embeddings with both high quality and high relevance to text prompts, we use both LAION-COCO and LAION-Aesthetics for visual decoding training. More specifically, we filter all text prompts with length greater than 150 to preserve a large enough batch size and prevent the GPU memory overflow. This rule discards about 8% of LAION-Aesthetics and 0.01% of LAION-COCO data, which has little effect on data diversity.

A.2.2 Training Details

The detailed training setups are listed in Table 5.

Appendix B Instruction Tuning

We collect publicly available language, image and video instruction datasets for instruction tuning.

Language instructions: ShareGPT contains about 70K user dialogues with ChatGPT or GPT-4, and Alpaca dataset contains 52K instruction-following data generated using self-instruct from OpenAI’s text-davinci-003.

Image instructions: we use LLaVA dataset consisting of three types of visual instructions, conversation, detailed description, and complex reasoning, with a total number of 158K image-text instruction-following samples. In our preliminary experiments, we found the instruction-tuned model often generates instruction-irrelevant detailed descriptions of the image. Thus, we remove the detailed description subset of LLaVA. We also find a bad pattern ’on top of the back of’ in the model’s response, and we filter all data that contains this pattern. The resulting 130K LLaVA subset is used for instruction tuning.

Video instructions: we use VideoChat-11K and a subset of Video-ChatGPT-100k as our video-instruction dataset. VideoChat-11K dataset is built from WebVid-10M consisting of 7K detailed video descriptions and 4K video conversations. Video-ChatGPT-100k is built from ActivityNet, and we sample an around 30K subset that includes only videos under one minute.

We use a batch size of 128 and train for 10K steps, with 3 epoches for ShareGPT, Alpaca and LLaVA datasets, and 60K samples for video-instruction data. We attach LoRAs on all linear projections of the self-attention layer, with the LoRA rank and alpha being 16.

B.2 System Messages

We use different system messages for language-instruction, image-instruction and video-instruction datasets, as shown in Table 6.

Appendix C Evaluation

Emu excels at performing diverse types of completion in multimodal sequences by accepting multimodal prompts, including text, images, videos, or their combinations, and generating comprehensive multimodal responses. To evaluate the capabilities of Emu, we conduct extensive benchmark tests covering various tasks, which are summarized in Table 7. Specifically, we meticulously select 9 benchmarks that encompass multimodal image/video and language tasks, including text-to-image generation, visual question answering for both images and videos, and image-based visual dialogue. When benchmarking OKVQA, we use VQAv2 evaluation codehttps://github.com/GT-Vision-Lab/VQA and stem the answers using Porter stemming to consolidate answers following . For other tasks, we either submit our results for evaluation on the official website or use standard evaluation code.

C.2 Zero-shot Evaluation

Prompt Template. To ensure that the model outputs answers in the required style for the benchmark tests, we prompt Emu and Emu-I with task-specific templates, as shown in Table 8. For each type of task, we have developed dedicated templates to structure the model’s output. In these templates, “{question}” will be replaced with the question from the question-answering task, “{history question}” will be replaced with the historical question from the multi-turn visual dialogues, and similarly “history answer” will be replaced with the historical annotated answer from the multi-turn visual dialogues. Then, the image/video will be added before the text as input to the model. Additionally, we implement post-processing techniques to filter out commonly occurring redundant phrases such as “it is”, “it’s”, “a”, “an”, and “the”. Furthermore, the model is required to output “unanswerable” for questions that cannot be answered in the VizWiz dataset. To achieve this, we augment the template by adding the phrase “is the answer known?” and prompt the model to respond with either “yes” or “no” by constraining the model generation. If the model responds with “no”, we immediately return the answer as “unanswerable”. On the other hand, if the model responds with “yes”, we proceed to prompt the model to provide a valid answer.

Multimodal Chain-of-Thought Prompting. To enhance the capabilities of the pretrained model, we utilize the Multimodal Chain-of-Thought prompting technique. Initially, when presented with an image or video, we employ a prompt to guide the model in generating a descriptive caption. Subsequently, the model is given both the caption and a task-specific prompt to generate the final result. The complete prompt template is shown in Table 8, where the “{caption}” tag in template will be replaced with the descriptive text generated by Emu. The experimental results demonstrate that this test-time technique effectively improves the model’s performance without any additional data, leveraging the inherent capability of the model itself.

Text-only Examples Prompting. To ensure a fair comparison with Flamingo, we include results obtained through text-only examples prompting, denoted by an asterisk (*) in Table 1. We adopt the same approach as Flamingo in selecting examples (i.e., RICES). This involves utilizing two text-only examples from the task as prompts, without any accompanying images (similar to the few-shot text prompts). During the evaluation process, we observed that this approach effectively formats the model’s output, regardless of the label format of the datasets and the evaluation metrics employed, enabling a more accurate reflection of its true performance.

C.3 Few-shot Evaluation

In the few-shot evaluation settings, we incorporate a few example samples as prefixes in the template and connected the few-shot examples using “. ”. Additionally, like Flamingo, we employ the Retrieval In-Context Example Selection (RICES) approach to select the few-shot examples.

To implement RICES, we begin by randomly selecting 5000 training set samples for each dataset. Then, using the pretrained EVA-CLIP model, we extract features from both the training set images/videos and the test set images/videos. For each test set sample, we select examples from the training set based on the highest cosine similarity using the extracted features, including them in the prompt. For the video-text task, we retrieve similar videos from the training set by comparing the mean of frame-level visual features extracted from our pretrained EVA-CLIP model.

Furthermore, we discover that the support video examples didn’t require too many frames, which could exceed the LLM’s context length limit. Therefore, we sample 8 frames for the given video and only 2 frames for the corresponding support video examples.

Appendix D Qualitative Cases