Generative Multimodal Models are In-Context Learners

Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, Xinlong Wang

Introduction

Multimodal tasks encompass anything involving understanding and generation in single or multiple modalities , which can be highly diverse and long-tail. Previous multimodal systems largely rely on designing task-specific architecture and collecting a sizable supervised training set, both of which are difficult to scale, particularly when this process needs to be repeated for each new task encountered. By contrast, humans can solve a new task in context, i.e., with only a few demonstrations or simple instructions – a capability that current multimodal models have yet to learn.

Recently, generative pretrained language models have demonstrated strong in-context learning abilities . By training a 37-billion-parameter model Emu2 and thoroughly evaluating it on diverse multimodal tasks, we demonstrate that a scaled-up multimodal generative pretrained model can harness similar in-context learning abilities and effectively generalize to unseen multimodal tasks. Emu2 is trained with a unified autoregressive objective: predict-the-next-multimodal-element (either visual embeddings or textual tokens). In this unified generative pretraining process, large-scale multimodal sequences (e.g., text, image-text pairs, and interleaved image-text-video) are used for model training.

We measure Emu2’s capabilities of learning from a few examples or instructions on standard multimodal datasets, as well as new tasks unseen in the training set. Specifically, Emu2 is evaluated under two scenarios: (a)(a) few-shot setting, where we allow as many examples as possible to fit the context window of the model; and (b)(b) instruction tuning, where the model is tuned to follow specific instructions.

Emu2 achieves promising results in the few-shot setting on a wide range of vision-language tasks. For example, it demonstrates state-of-the-art few-shot performance on multiple visual question-answering datasets. We observe a performance improvement when the number of examples in context increases. Figure 1 illustrates Emu2’s strong multimodal reasoning capabilities for tasks in the wild, e.g., recognition and counting in a specific format. Emu2 also learns to follow visual prompting in context (e.g., the circles laid on the images in Figure 1), even although it struggles at a smaller scale or at zero shot.

As Emu2 is inherently equipped to handle interleaved text-image-video at both input and output, it serves as a powerful and versatile base model for diverse multimodal tasks, by following specific task instructions. For example, after instruct tuning with conversational data, Emu2 achieves state-of-the-art results on visual question-answering tasks, and surpasses previous models of more complex designs. In addition, Emu2 can be fine-tuned to function as a controllable visual generation model of high quality. It is capable of accepting a mixture of text, locations and images as conditions, and generating images that are grounded as specified.

Given the broad spectrum of capabilities displayed by Emu2, we conduct a thorough analysis of its potential societal implications and discuss in detail potential concerns over misuse. By identifying further tasks where Emu2’s in-context learning can further improve, we highlight the necessity for continuous enhancement of the model and the importance of deploying Emu2 responsibly.

Approach

Emu2 is a generative multimodal model that learns with a predict-the-next-element objective in multimodal context. As illustrated in 2, the architecture of Emu2 consists of three components: Visual Encoder, Multimodal Modeling, and Visual Decoder. Each image in the input multimodal sequence is tokenized into continuous embeddings via the Visual Encoder and then interleaved with text tokens for autoregressive Multimodal Modeling. The regressed visual embeddings are then decoded into an image or a video by the Visual Decoder. Specifically, we leverage pretrained EVA-02-CLIP-E-plus , LLaMA-33B and SDXL to initialize the Visual Encoder, Multimodal Modeling, and Visual Decoder, respectively. Compared to Emu , Emu2 embraces a simpler framework which connects the Visual Encoder and Multimodal Modeling through mean pooling each image to 8×88\times 8 image patches, followed by a linear projection, instead of using an additional C-Former .

2 Pretraining

The pretraining data for Emu2 comprises several publicly accessible datasets, including image-text pairs from LAION-2B and CapsFusion-120M , video-text pairs from WebVid-10M , interleaved image-text data from Multimodal-C4 (MMC4) , interleaved video-text data from YT-Storyboard-1B , grounded image-text pairs from GRIT-20M introduced by Kosmos-2 and CapsFusion-grounded-100M curated by CapsFusion-120M. Additionally, language-only data from Pile is included to retain textual reasoning capability.

2.2 Training

Similar to Emu , Emu2 learns with the predict-the-next-element objective within a multimodal sequence. Each image is encoded into N=64N=64 dimension-fixed visual embeddings and then interleaved with text tokens to construct a multimodal sequence. The interleaved sequence is then fed into a Transformer decoder for autoregressive modeling.

Emu2 is first pretrained on image-text and video-text pair data with only captioning loss on the text tokens. The input images are resized to 224×224224\times 224. We adopt the AdamW optimizer with β1=0.9\beta_{1}=0.9, β2=0.95\beta_{2}=0.95, ϵ=1×10−6\epsilon=1\times 10^{-6}. The maximum learning rate is 1×10−41\times 10^{-4} for the linear projection layer, 3×10−53\times 10^{-5} for Multimodel Modeling, and 5×10−55\times 10^{-5} for Visual Encoder. We pretrain Emu2 on 162 million image-text samples and 7 million video-text samples for 35,200 iterations. The global batch size is 6,144 for the image-text pairs and 768 for video-text pairs. The training process is then restarted at a higher 448-pixel resolution for an additional 4,000 iterations.

Then, we freeze the Visual Encoder and only optimize the linear projection layer and Multimodel Modeling with both text classification loss and image regression loss. Additional datasets including image-text interleaved data, video-text interleaved data, grounded image-text pair data, and language-only data are used in the training. All images are resized to 448×448448\times 448, and the maximum learning rate is 1×10−51\times 10^{-5}. We use a global batch size of 12,800 for image-text pair data, 6,400 for video-text pair data, 3,200 for image-text and video-text interleaved data, and 800 for language-only data. The training process spans 20,350 iterations and consumes about 160 million samples of image-text data and 3.8B tokens of language-only data.

2.3 Visual Decoding

We train the Visual Decoder to directly decode visual embeddings generated by the Visual Encoder into image. We use SDXL-base as the initialization of our Visual Decoder, which is fully trained to solve the new task of autoencoding. Specifically, we use NN visual embeddings as the condition input to the Visual Decoder and adjust the dimension of the projection layers in cross-attention modules to match the dimension of visual embeddings.

Unlike Emu where each optimization step of its Visual Decoder requires an autoregressive inference of the language model, Emu2’s visual decoding can be considered as training a detokenizer, which can be trained off-the-shelf without the language model. Once trained, the Visual Decoder together with the Visual Encoder works as an image autoencoder that can tokenize an image into embeddings and detokenize back. During Emu2 inference, it generates NN image embeddings and decodes to an image on the fly.

For the decoding of video data, we train a diffusion-based decoder . Similar to , we adapt a 2D denoising U-Net to 3D style by inserting a 1D temporal convolution following each 2D spatial convolutional layer and extending the spatial attention to spatial-temporal attention. This video decoder is initialized via Stable Diffusion 2.1 and fully trained to generate video clips conditioned on visual embeddings from Emu2.

Training Setup. We use the images in LAION-COCO and LAION-Aesthetics to train the Visual Decoder under the task of image autoencoding. The Visual Encoder and VAE in SDXL are frozen, and only the U-Net is updated during training. We adopt AdamW optimizer with β1=0.9,β2=0.999\beta_{1}=0.9,\beta_{2}=0.999 and the weight decay of 0.01. We use loglog learning rate warm-up and linear learning rate decay with a peak learning rate of 1×10−41\times 10^{-4} for 2,000 and 6,000 steps, respectively. We filter out images whose resolution is lower than 512×512512\times 512. The input to the Visual Encoder is set to 448×448448\times 448, while the output of the Visual Decoder is set to 1024×10241024\times 1024. We also employ the classifier-free guidance , which randomly discards image embeddings with the probability of 10%10\%. The batch size is set to 2,048 in total.

3 Instruction Tuning

Emu2 can be efficiently aligned to follow specific task instructions. We fine-tune the base model with conversational data to yield Emu2-Chat, which is capable of following multimodal questions and making responses in dialogue. Similarly, we derive a controllable visual generation model Emu2-Gen, which is capable of accepting a mix of text, locations, and images as conditions, and generating images that are grounded in the specified text or subject.

Training Data. We adopt a uniform approach to train on both academic-task-oriented datasets and multimodal chat data to empower Emu2-Chat with the instruction-following ability while retaining rich visual knowledge. As academic-task-oriented datasets have brief annotations that limit the model’s capacity to provide more comprehensive and helpful responses, we distinguish between these two data categories by employing different system messages and including instructions with output-format control information as used in . A summary of data used is as follows: (a)(a) Academic-task-oriented data: image captioning , visual question answering , knowledgeable question answering , multimodal classification , and referring expression comprehension . (b)(b) Multimodal chat data: GPT-assisted visual instruction , language instruction , clock reading , and video chat .

Training Objective. In instruction tuning of Emu2-Chat, two special tokens, [USER] and [ASSISTANT], are incorporated into the model to denote roles. These tokens help organize different data types in the following format: “ [USER]: [ASSISTANT]: ”. Here represents system message and varies between the two major task categories (academic-task-oriented and multimodal chat). The section comprises multimodal tokens, including images, videos, and text. Only tokens in the section will be supervised by cross-entropy loss during training.

Training Setup. We use a global batch size of 768 and train for 8k steps. The learning rate linearly warms up to 1×10−51\times 10^{-5} in the first 100 steps, then decays to zero with a cosine schedule. The model is trained using the AdamW optimizer with β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98, ϵ=1×10−6\epsilon=1\times 10^{-6}, and a gradient clipping of 5.0. The sequence length during training is limited to 2048, and any excess beyond that is truncated directly. We consistently employed an input image/video resolution of 448 ×\times 448. For video data, we uniformly sample frames in time as input to the model. The number of sampled frames for each video is randomly chosen from 8, 12, and 16. To capture more intricate spatial details, following the visual encoder stage, we apply mean-pooling to each static image, dividing it into 16 ×\times 16 tokens during instruction fine-tuning. This differs from the pre-training phase, where 8 ×\times 8 tokens were utilized.

3.2 Controllable Visual Generation

Training Data. We leverage a mix of high-quality datasets to unleash the potential of controllable generation in context. We use a grounded image-text pair dataset CapsFusion-grounded-100M and GRIT for grounded text-to-image generation. To mitigate the impact of image backgrounds on the effectiveness of multi-entity subject-driven generation, we employ SAM to preprocess the grounding data, yielding a subset of approximately 5 million samples with segmentation results. Additionally, we leverage InstructPix2Pix constructed by for image editing tasks. For the text-to-image task, we use a filtered subset of CapsFusion , LAION-Aesthetics , SA-1B , and LAION-High-Resolution .

We also collect data from premium sources (e.g., Unsplash ) and outputs from advanced text-to-image systems (e.g., Midjourney-V5 and DALL-E-3 ) for quality fine-tuning. This diverse dataset includes around 500k high-quality image-text pairs. For all the data above, during the training, only samples with image resolutions higher than 448×448448\times 448 were retained to ensure generation quality. More details can be found in the supplementary.

Training Objective. We use the same unified generative pretraining objective to adapt to diverse generation tasks in context. Specifically, a training sample for generation is formulated as: “A photo of

a man

image embedding of object localization image[IMG]image embedding of man[/IMG]sitting next to

a dog

image embedding of object localization image[IMG]image embedding of dog[/IMG][IMG]image embedding of the whole image[/IMG]
”. We represent the coordinates of each object directly in image form by drawing the bounding box of each object at its specified location on a black image. Our Emu2-Gen conducts unified multimodal modeling of the text, object image, and corresponding object localization image. The regression loss only applies to the visual embeddings of the last image. We freeze the Visual Encoder during fine-tuning. We randomly drop tokens of entities and object localization image to enhance model adaptability and robustness. Additionally, we apply data augmentation to each object image, incorporating random background variations and random crop, aiming to reduce the reliance on image backgrounds.

Training Setup. We use a global batch size of 4,096 and train for 3k steps. The learning rate linearly warms up to 5×10−55\times 10^{-5} in the first 100 steps, then decays to zero with a cosine schedule. We further fine-tune for 900 steps using the 500k high-quality pairs with a batch size of 2048.

Evaluation

We evaluate zero-shot and few-shot abilities of Emu2 on OKVQA , VQAv2 , VizWiz , TextVQA , and HatefulMemes tasks. Details of the datasets and prompts can be found in supplementary materials. The results are presented in Table 1. Emu2 demonstrates remarkable in-context ability, showcasing improved performance with more in-context samples seen. Specifically, on VQAv2, VizWiz and TextVQA datasets, Emu2 outperforms Flamingo-80B and IDEFICS-80B under all few-shot settings with a much smaller model scale (37B).

Figure 1 demonstrates Emu2’s few-shot capabilities in the wild. For example, the model learns to classify and count simultaneously in a specific format via a few examples (row 1). Additionally, Emu2 is capable of following visual prompts in context, e.g., the red circles laid on the images (row 2 and 3).

2 Instruction-Following Chat

Our Emu2-Chat is evaluated on academic-task-oriented benchmarks including image question-answering datasets (VQAv2 , OKVQA , GQA , VizWiz , TextVQA ) and video question-answering datasets (MSVD and MSRVTT ). The evaluation also encompassed recent benchmarks for large multimodal models, including SEED-Bench , MM-Vet , TouchStone and MMMU . When evaluated on SEED-Bench, we followed the setup of LLaVa-1.5 by presenting options to the model for completing multiple-choice tasks.

As shown in Table 2, Emu2-Chat consistently outperforms other models in image question-answering tasks, encompassing well-established benchmarks like VQAv2 and GQA. Notably, it shows a noticeable improvement in the OKVQA task, which requires the utilization of external knowledge, showcasing the advantage of our model for mastering real-world knowledge. When it comes to video question-answering, Emu2-Chat demonstrated advantages even though it did not use video question-answering data for training. It achieved an accuracy of 49.0 and 31.4 on the MSVD-QA and MSRVTT-QA tasks, respectively, surpassing InstructBLIP, Emu, and the larger Flamingo-80B. More importantly, our model has also achieved better results on LMM benchmarks. LMM benchmarks such as MM-Vet provide a more comprehensive evaluation of model abilities, including solving complicated tasks. Emu2-Chat achieves a score of 48.5 in MM-Vet and 703.8 in TouchStone, confirming its superior capability in understanding and solving multimodal problems compared to existing models.

3 Controllable Visual Generation

Qualitative Results. Figure 3 presents a visualization of Emu2’s autoencoding results. With Emu2’s Visual Encoder and Visual Decoder, we can tokenize an image into visual embeddings and detokenize them back. Compared with SEED and Emu , Emu2 shows significantly superior results. We also evaluate our image autoencoding results on MS-COCO and achieve a strong 0.907 CLIP-I score. More results are in the supplementary.

As depicted in Figure 4, Emu2-Gen is capable of accepting a mixture of text, locations and images as input, and generating images in context. The model skillfully engages in various controllable visual generation tasks in a zero-shot setting, capitalizing on the in-context learning capabilities in multimodality. Examples in Figure 4 show generated images of three dogs conditioned on different subjects, locations and scenarios. The presented visual samples demonstrate the model’s proficiency in tasks such as re-contextualization, stylization, modification, region-controllable generation, and multi-entity composition.

Zero-shot Text-to-image Generation. We evaluate the zero-shot text-to-image generation capability on 30k randomly sampled data from the MS-COCO validation set. We employ CLIP-ViT-B , following the approach in DALL-E 3, to calculate the CLIP-T score to assess prompt-following ability. Additionally, we utilize CLIP-ViT-L, as in GILL, to compute the CLIP-I score for measuring image similarity. A higher score means the generated image is more similar to the prompt or the real image. Table 3 shows that Emu2-Gen achieves the state-of-the-art performance in terms of both CLIP-I and CLIP-T scores compared to various unimodal generation models and multimodal models. More text-to-image generation cases can be found in supplementary.

Zero-shot Subject-driven Generation. Following Kosmos-G , we also evaluate our model’s subject-driven image editing ability on DreamBench . We generate four images for each prompt, resulting in a total of 3,000 images for a comprehensive evaluation. We employ DINO and CLIP-I to evaluate subject fidelity, and CLIP-T to evaluate text fidelity, aligning with the methodology established by DreamBooth. Notably, Emu2-Gen excels in subject fidelity, as evidenced by its superior performance on DINO and CLIP-I metrics compared to methods like BLIP-Diffusion and Kosmos-G. Emu2-Gen impressively reconstructs subjects with just one image input in zero-shot setting, demonstrating superior subject fidelity through powerful visual decoding. Further illustrative cases are provided in the supplementary, showcasing Emu2-Gen’s proficiency in multi-entity generation.

Related Work

Large Multimodal Models. Recent years have witnessed the rapid growth of large multimodal models . CLIP has pioneered the learning of large multimodal models with a contrastive learning objective on massive image-text pair data, yielding impressive zero-shot performance on image classification tasks. The BEiT series re-imagines visual signals as discrete tokens, allowing for language-model-like pretraining. Flamingo and Kosmos series exhibit promising zero-shot and few-shot multi-modal understanding performance by pretraining on large-scale image-text interleaved data. With the remarkable progress in LLM and its open sourcing, connecting pre-trained vision backbones with LLMs with image-text pairs or visual instruction tuning data becomes a popular solution to efficiently learning large multimodal models. BLIP series , LLaVA and MiniGPT4 show promising results by connecting vision encoders and LLMs with a small intermediate model. A school of successive efforts further improves visual instruction tuning with better overall training pipelines , grounding annotations , and extra tasks . There are early studies on training more unified large multimodal models that are capable of performing multimodal understanding and generation simultaneously. In this paper, we further explore the distinct solution proposed in Emu : learning large multimodal models with generative objectives on both texts and images. By further scaling up generative multimodal models, we demonstrate promising in-context learning abilities like those of LLMs on both text and image generation tasks.

In-Context Learning. Recent advancements in large language models have underscored their capacity for in-context learning , where models leverage a few contextual examples to adapt to new tasks. This phenomenon, particularly evident as LLMs scale up in size and data, has been exploited for complex challenges such as mathematical reasoning , signaling new emergent ability in model behavior . There are a few early attempts in in-context learning in the realm of vision and multimodality. Flamingo integrates visual inputs to LLMs, enabling the in-context learning of visual-linguistic tasks such as image captioning and OCR through language-based interfacing. Painter and SegGPT conduct an early study of visual in-context learning. Inspired by the emerging abilities of large language models, in this work we study the problem of multimodal in-context learning by scaling up generative multimodal models and demonstrating strong results in broad understanding and generation tasks.

Broader Impact and Limitations

Large multimodal models offer a wide range of benefits to society, from enhancing visual navigation and medical diagnostics to increasing accessibility for individuals with visual impairment. The in-context learning capabilities of Emu2 allow it to quickly adapt to new tasks or environments, even with limited data, ushering in numerous potential applications. The generative capabilities of Emu2 can be highly valuable to the creative industries.

However, there are potential downsides in more powerful multimodal models to be considered. The hallucination issue of multimodal models may cause incorrect and unreasonable predictions in certain cases. Emu2 may also generate harmful or biased content like other generative models since the training data may be biased or contain unsuitable content. We are actively working to enhance the robustness of multimodal models, reduce model hallucinations, improve the fairness of training data, and reduce toxic data. We also call on the wider community to pay attention to the potential social impact of multimodal models as they are growing larger and stronger.

One of the limitations of Emu2 is that its in-context learning capability could fail in some complex scenes or tasks, e.g., counting in a crowd. Additionally, there is still a gap between Emu2’s question-answering capability and that of closed multimodal systems. For example, GPT-4V achieves 67.7 MM-Vet score vs. Emu2’s 48.5, although already being state-of-the-art among public models. We believe there is much room to improve as the quality and quantity of training data improve and as model scale continues to grow.

Conclusion

We present a 37 billion-parameter generative multimodal model Emu2 that shows strong performance and versatility on many multimodal tasks in the in-context settings. Emu2 serves as a base model and a general-purpose interface for a variety of multimodal tasks. We demonstrate state-of-the-art results on a broad range of benchmarks of multimodal understanding and generation. Specifically, our model largely surpasses prior work on the lately proposed LMM benchmarks that require more advanced capability compared to classic academic benchmarks. Emu2 also shows remarkable capability of controllable visual generation in multimodal context, e.g., subject-/text-grounded generation. Additionally, we review the limitations and broader social impact of Emu2. Despite discussed weaknesses, these results suggest that generative multimodal model at scale may be an important step towards the development of adaptable, general multimodal systems.

References

Appendix A More Pretraining Details

In pretraining, we exclusively leverage image-text pairs and video-text pairs for stage 1 training. We additionally leverage interleaved and language-only data altogether for stage 2. The integration of visual embeddings with text tokens generates unified multimodal sequences. These sequences are then structured by appending the tokens and to denote the beginning and end of each sequence.

In the pretraining stage, we utilize image-text pairs from LAION-2B and CapsFusion-120M , along with video-text pairs from WebVid-10M . During pretraining stage 2, each image or video is randomly placed before or after its corresponding text with a probability of 0.5, respectively. For each video, we randomly sample 8 frames. To structure the visual embeddings, we append two special tokens, [IMG] and [/IMG], to signify the start and end of the visual embeddings. In the case of videos, where there are TT frames, each frame is encoded into a set of visual embeddings, and a special token, [VIDEO], is prepended to the start of the frame embedding sequence. This design helps distinguish between multiple images and video frames within the multimodal sequences.

We harness the Multimodal-C4 (MMC4) dataset and the YT-Storyboard-1B dataset as expansive sources of image and video-text interleaved data. This approach aims to unlock the in-context learning capability of multimodal models. For each MMC4 document, we randomly sample N = 8 images, accompanied by their corresponding sentences, to construct a subsequence of L = 1024. During pretraining stage 2, each image or frame is randomly positioned before or after its corresponding text with a probability of 0.5. The special tokens used in this interleaved data are consistent with those employed in the image-text pair data.

We curated a dataset of grounded image-text pairs named CapsFusion-grounded-100M, employing data from CapsFusion processed through the dataset construction pipeline proposed by Kosmos-2 . Additionally, we utilized the 20M GRIT dataset introduced by Kosmos-2 . To enhance the diversity and context of the dataset, we randomly positioned each phrase before or after its corresponding coordinates with a probability of 0.7. The bounding box can be represented using its top-left point (x1,y1)(x1,y1) and bottom-right point (x2,y2)(x2,y2). We transform continuous coordinates into 224 discrete tokens , the coordinates of a sample box can be formulated as <loc000loc_{000}><loc000loc_{000}><loc224loc_{224}><loc224loc_{224}>. We added these tokens to the word vocabulary to facilitate unified modeling with text. To distinguish grounding text from regular text strings, we introduced two special tokens, and , marking the beginning and end of the bounding box coordinates. Moreover, to establish the correct association between bounding boxes and their corresponding descriptive phrases, an additional set of special tokens,

and

, was appended. To guide the model in grounding text output to the provided image, we utilized the special token . This comprehensive set of tokens and instructions enriches the training data for effective multimodal modeling and understanding.

To maintain text reasoning capabilities, we engage in joint training with the language modeling dataset Pile . The entire text corpus from Pile is preprocessed offline, and each training sample is tokenized into 2048 tokens using the LLaMA tokenizer. We randomly sample a total of 3.6 billion tokens for pretraining purposes.

A.2 Training Hyperparameters

We report the detailed training hyperparameter settings of Emu2 during the pretraining in Table 5.

A.3 Visual Decoding

We utilize images in LAION-COCO and LAION-Aesthetics to train the Visual Decoder. Images whose resolution is smaller than 512×512512\times 512 are filtered to prevent generating low-quality results. We employ ratio-preserving random scaling followed by random cropping of a square portion from the scaled image to keep all training images unstretched. The original image size and crop coordinates are used as additional conditions following SDXL .

A.3.2 Training Hyperparameters

The detailed hyperparameters of visual decoding training are summarized in Table 6.

Appendix B Instruction-Following Chat

We used two types of training data, academic-task-oriented data and multi-modal chat data, in instruction fine-tuning of Emu2-Chat The academic-task-oriented datasets we utilized comprise image captioning datasets such as COCO Caption , and TextCaps , as well as visual question-answering datasets like VQAv2 , OKVQA , GQA , TextVQA , and multi-modal classification data constructed in M3IT . RefCOCO , RefCOCO+ and RefCOCOg datasets are also used. The public multi-modal chat data we use includes GPT-assisted visual instruction data LLaVa and LLaVaR , language instruction data from ShareGPT and Alpaca , and video instruction data from VideoChat . Beyond these, we constructed instruction fine-tuning data from an analog clock reading dataset . For academic-task-oriented datasets, we use the system message “You are a helpful assistant, dedicated to provide concise and efficient answers.”, and for the multi-modal chat data, the system message is “You are a helpful assistant, dedicated to delivering comprehensive and meticulous responses.”.

B.2 Training Hyperparameters

The detailed training hyper-parameters of Emu2-Chat are summarized in Table 7.

Appendix C Controllable Visual Generation

We use the grounded image-text pairs dataset, i.e., CapsFusion-grounded-100M and GRIT for grounded text-to-image generation. We use SAM to obtain segmentation results for the corresponding grounding boxes. We leverage InstructPix2Pix constructed by for image editing tasks. The sample will be formulated as “[IMG]embedding of origin image[/IMG]instruct editing prompt[IMG]embedding of edited image[/IMG]”. For the text-to-image task, we use a filtered subset the CapsFusion , LAION-Aesthetics , SA-1B , and LAION-High-Resolution .

For high-quality fine-tuning, our datasets were meticulously sourced from premium sources, e.g., Unsplash , and outputs from advanced text-to-image systems, e.g., Midjourney-V5 and DALL-E-3 . This comprehensive approach ensured a diverse and rich dataset, comprising approximately 500,000 instances of high-quality image-text pairs, instrumental in refining and enhancing the aesthetic quality of our Emu2-Gen model’s generated images.

C.2 Training Hyperparameters

We report the detailed training hyperparameter settings of Emu2-Gen during the instruction-tuning in Table 8.

Appendix D Evaluation Details

For few-shot evaluation of Emu2, we adopt the Retrieval In-Context Example Selection (RICES) approach for choosing few-shot examples, following Flamingo and Emu . The chosen few-shot examples will be separated by “. ” and then placed ahead of the test sample. We use the prompt ”[image] based on the picture, [question] short answer:”. For zero-shot evaluation, as no example is given, we find the above simple prompt cannot effectively control the model behavior and the model tends to output a sentence rather than a word or phrase. Thus, we modify the prompt to ”[image] based on the picture, answer in one word or phrase. [question] short answer:”. This adjustment aligns the model’s output more closely with the distribution of the tested datasets, where responses typically consist of a succinct word or phrase. The splits and metrics for each benchmark are detailed in Table 9.

The evaluation of Emu2-Chat follows the assessment method of Emu-I , utilizing generation hyper-parameters with a beam size of 5. For video input, 16 frames are uniformly sampled as visual conditions. In the question-answering benchmark that requires short answers, we employ the system message “You are a helpful assistant, dedicated to provide concise and efficient answers.” along with the output format control information used in . In the benchmark for scoring with GPT-4, we use the system message “You are a helpful assistant, dedicated to delivering comprehensive and meticulous responses.”. We provide an overview of the evaluation benchmarks in Table 9.

For all evaluation of visual generation tasks, we use EulerDiscreteScheduler with 50 diffusion steps. The classifier-free guidance scale is set to 3.0. To evaluate on DreamBench , we select exactly the same image for each object as chosen in Kosmos-G . Similarly to Kosmos-G, we also slightly modified the original prompt for the the original prompt with the prefix ”a” , for example, ”a red {}” is modified to ”{}Make it red”

Appendix E Qualitative Results

We present qualitative cases for Emu2-Gen in Figure 5-11 and for Emu2-Chat in Figure 12-14, respectively.