Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, Aniruddha Kembhavi
Introduction
As AI researchers, we seek to build intelligent agents that can perceive their environment, communicate with others, act in the world, and reason about their interactions. The world is multimodal, so our agents must partake in rich interactions that are multimodal in nature via vision, language, sound, action etc. Psychologists have argued that the redundancy of our sensory systems serves as supervisory mechanisms to improve each other . This provides a natural motivation to create models with similar learning capabilities, supporting many different modalities that can supervise each other during training.
Building models that can parse and produce many modalities is a complex undertaking. Training Large Language Models (LLMs) with billions of parameters, despite only supporting a single modality, is extremely challenging across many fronts – from sourcing and processing massive datasets, ensuring data quality and managing biases, designing effective model architectures, maintaining stable training processes, and instruction tuning to enhance the model’s ability to follow and understand user instructions. These challenges are hugely amplified with the addition of each new modality.
In light of these difficulties, a line of recent works in building multimodal systems has leveraged pre-trained LLMs, with some augmenting with new modality encoders , some adding modality specific decoders and others leveraging the LLM’s capabilities to build modular frameworks . Another line of works on training multimodal models from scratch has focused on generating text output with a few recent works supporting the understanding and generation of two modalities – text and images . Building generative models with a wider coverage of modalities, particularly when training from scratch, remains an open challenge.
In this work, we present Unified-IO 2, a large multimodal model (LMM) that can encode text, image, audio, video, and interleaved sequences and produce text, action, audio, image, and sparse or dense labels. It can output free-form multimodal responses and handle tasks unseen during training through instruction-following. Unified-IO 2 contains 7 billion parameters and is pre-trained from scratch on an extensive variety of multimodal data – 1 billion image-text pairs, 1 trillion text tokens, 180 million video clips, 130 million interleaved image & text, 3 million 3D assets, and 1 million agent trajectories. We further instruction-tune the model with a massive multimodal corpus by combining more than 120 datasets covering 220 tasks across vision, language, audio, and action.
Our pre-training and instruction tuning data, totaling over 600 terabytes, presents significant challenges for training due to its diversity and volume. To effectively facilitate self-supervised learning signals across multiple modalities, we develop a novel multimodal mixture of denoiser objective that combines denoising and generation across modalities. We also develop dynamic packing – an efficient implementation that provides a 4x increase in training throughput to deal with highly variable sequences. To overcome the stability and scalability issues in training, we propose to apply key architectural changes, including 2D rotary embeddings, QK normalization, and scaled cosine attention mechanisms on the perceiver resampler. For instruction tuning, we ensure every task has a clear prompt, either using existing ones or crafting new ones. We also include open-ended tasks and create synthetic tasks for less common modalities to enhance task and instruction variety.
We evaluate Unified-IO 2 on over 35 datasets across the various modalities it supports. Our single model sets the new state of the art on the GRIT benchmark, which includes diverse tasks such as keypoint estimation and surface normal estimation. On vision & language tasks, it matches or outperforms the performance of many recently proposed VLMs that leverage pre-trained LLMs. On image generation, it outperforms the closest competitor that leverages the pre-trained stable diffusion model , especially in terms of faithfulness as per the metrics defined in . It also shows effectiveness in video, natural language, audio, and embodied AI tasks, showcasing versatility despite its broad capability range. Moreover, Unified-IO 2 can follow free-form instructions, including novel ones. Figure 1 offers a glimpse into how it handles various tasks. Further examples, along with the code and models, are accessible on our project website.
Related Work
Inspired by the success of language models as general-purpose text processing systems , there has been a recent wave of multimodal systems trying to achieve similar general-purpose capabilities with additional modalities. A common approach is to use a vision-encoder to build features for input images and then an adapter to map those features into embeddings that can be used as part of the input to an LLM. The network is then trained on paired image/language data to adapt the LLM to the visual features. These models can already perform some tasks zero-shot or with in-context examples , but generally a second stage of visual instruction tuning follows using instructions, visual inputs, and target text triples to increase zero-shot capabilities .
Building upon this design, many researchers have expanded the breadth of tasks these models can support. This includes creating models that can do OCR , visual grounding , image-text-retrieval , additional languages , embodied AI tasks or leverage other expert systems . Other efforts have added new input modalities. This includes video inputs , audio or both . PandaGPT and ImageBind-LLM use the universal encoder ImageBind to encode many kinds of input modalities, and ChatBridge uses a similar universal encoder based on language. While these efforts are effective for understanding tasks, they do not allow complex multimodal generation and often exclude modalities long considered central to computer vision (e.g., ImageBind cannot support sparse annotation of images).
Fewer works have considered multimodal generation. Unified-IO , LaVIT , OFA , Emu and CM3Leon train models to generate tokens that a VQ-GAN can then decode into an image, while GILL , Kosmos-G and SEED generate features that a diffusion model can use, and JAM fuses pre-trained language and image generation models. Unified-IO 2 also uses a VQ-GAN, but supports text, image, and audio generation.
Overall, this shows a strong trend towards expanding the number of supported tasks and modalities. Unified-IO 2 pushes this trend to its limit, including the capabilities of these prior works with few exceptions and the ability to generate outputs in more modalities. Recently, CoDi also achieved similar any-to-any generation capabilities by using multiple independently trained diffusion models and aligning their embedding spaces. Unified-IO 2 has stronger language abilities and can perform well on many more tasks.
A notable feature of Unified-IO 2 is that the model is trained from scratch instead of being initialized with a pre-trained LLM. Prior works following this approach are typically not designed to produce complex generations like free-form text responses, images or sounds, or follow text instructions. Compared to recent general-purpose multimodals models , Unified-IO 2 has a significantly broader scope of tasks and outputs. Training from scratch means that the method can be reproduced without a costly preliminary stage of language model pre-training and is a more natural fit for how humans learn modalities simultaneously through their co-occurrences, not one at a time.
Approach
In this section, we discuss the unified task representation (3.1), the model architecture and techniques to stabilize training (3.2), the multimodal training objective (3.3) and the efficiency optimizations (3.4) used in Unified-IO 2.
Unified-IO 2 processes all modalities with a single, unified encoder-decoder transformer . This is achieved by encoding various inputs and outputs – images, text, audio, action, boxes etc., into sequences of tokens in a shared representation space. Our encoding procedure follows the design of Unified-IO , with several modifications to improve performance and new encoders and decoders for additional modalities. Figure 2 shows an overview of the model. Details about how modalities are encoded are given below.
Text, Sparse Structures, and Action. Text inputs and outputs are tokenized using the byte-pair encoding from LLaMA , which we chose since it supports Unicode symbols and preserves whitespace. Sparse structures such as bounding boxes, keypoints, and camera poses are discretized and then encoded using 1000 special tokens added to the vocabulary . Points are encoded with a sequence of two such tokens (one for and one for ), boxes are encoded with a sequence of four tokens (upper left and lower right corners), and 3D cuboids are represented with 12 tokens that encode the projected center, virtual depth, log-normalized box dimension, and continuous allocentric rotation . For embodied tasks, discrete robot actions are generated as text commands (e.g., “move ahead” to command the robot to move forward in navigation). Special tokens are used to encode the robot’s state, such as its position and rotation. Details are in Appendix B.1.
Images and Dense Structures. Images are encoded with a pre-trained Vision Transformer (ViT) . We concatenate the patch features from the second and second-to-last layers of the ViT to capture both low and high-level visual information. These features are passed through a linear layer to get embeddings that can be used as part of the input sequence for the transformer. To generate images, we use VQ-GAN to convert images into discrete tokens. These tokens are added to the vocabulary and then used as the target output sequence in order to generate an image. For better image quality, we use a dense pre-trained VQ-GAN model with patch size that encodes a image into 1024 tokens with a codebook size of 16512.
Following , we represent per-pixel labels, which include depth, surface normals, and binary segmentation masks, as RGB images that can be generated or encoded with our image generation and encoding abilities. For segmentation, Unified-IO 2 is trained to predict a binary mask given a class and bounding box. An entire image can be segmented by first doing detection, and then querying the model for a segmentation mask for each detected bounding box and class. See Appendix B.1 for details.
Audio. Unified-IO 2 encodes up to 4.08 seconds of audio into a spectrogram (See Appendix B.1 and Table 8). The spectrogram is then encoded with a pre-trained Audio Spectrogram Transformer (AST) , and the input embeddings are built by concatenating the second and second-to-last layer features from the AST and applying a linear layer just as with the image ViT. To generate audio, we use a ViT-VQGAN to convert the audio into discrete tokens. Since there is no public codebase, we implement and train our own ViT-VQGAN with patch size that encodes a spectrogram into 512 tokens with a codebook size of 8196.
Image and Audio History. We allow up to four additional images and audio segments to be given as input, which we refer to as the image or audio history. These elements are also encoded using the ViT or AST, but we then use a perceiver resampler , see Table 8 for hyperparameters, to further compress the features into a smaller number of tokens (32 for images and 16 for audio). This approach greatly reduces the sequence length and allows the model to inspect an image or audio segment in a high level of detail while using elements in the history for context. This history is used to encode previous video frames, previous audio segments, or reference images for tasks such as multi-view image reconstruction or image-conditioned image editing. Eight special tokens are added to the text vocabulary and used to reference the individual elements in these histories in the text input or output.
2 Architecture
Unified-IO 2 uses a transformer encoder-decoder architecture. However, we observe that using a standard implementation following Unified-IO leads to increasingly unstable training as we integrate additional modalities. As shown in Figure 3 (a) and (b), training only on image generation (green curve) results in stable loss and gradient norm convergence. Introducing a combination of image and text tasks (orange curve) slightly increases the gradient norm compared to a single modality, but remains stable. However, the subsequent inclusion of the video modality (blue curve) leads to an unrestrained escalation in the gradient norm. When an XXL version of this model is trained on all modalities, as shown in Figure 3 (c) and (d), the loss explodes after 350k steps, and the next token prediction accuracy significantly drops at 400k steps. To address this, we include various architectural changes that significantly stabilize multimodal training.
2D Rotary Embedding. Instead of relative positional embedding , we apply rotary positional embeddings (RoPE) at each transformer layer. For non-text modalities, we extend RoPE to two-dimensional positions: For any 2D indexes , we split each of the query and key embeddings of the transformer attention heads in half and apply separate rotary embeddings constructed by each of the two coordinates to the halves, see Appendix B.2.
QK Normalization. We observe extremely large values in the multi-head attention logits when including image and audio modalities, which leads to attention weights becoming either 0 or 1 and contributes to training instability. To solve this, following , we apply LayerNorm to the queries and keys before the dot-product attention computation.
Scaled Cosine Attention. We use perceiver resampler to compress each image frame and audio segment into a fixed number of tokens. We found that even with QK normalization, the attention logits in the perceiver can grow to extreme values. Therefore, we apply more strict normalization in the perceiver by using scaled cosine attention , which significantly stabilizes training.
To avoid numerical instabilities, we also enable float32 attention logits. Jointly updating the pre-trained ViT and AST can also cause instabilities. Thus, we freeze the ViT and AST during pretraining and finetune them at the end of instruction tuning. Figure 4 shows that the pre-training loss for our model is stable despite the heterogeneity of input and output modalities.
3 Training Objective
A strong multimodal model has to be exposed to solving diverse sets of problems during pre-training. UL2 proposed the Mixture of Denoisers (MoD), a unified perspective to train LLMs, which combines the span corruption and causal language modeling objectives. Motivated by this, we propose a generalized and unified perspective for multimodal pre-training.
Multimodal Mixture of Denoisers. MoD uses three paradigms: [R] – standard span corruption, [S] – causal language modeling, and [X] – extreme span corruption. For text targets, we follow the UL2 paradigms. For image and audio targets, we define two analogous paradigms: [R] – masked denoising where we randomly mask % of the input image or audio patch features and task the model to re-construct it and [S] – where we ask the model to generate the target modality conditioned only on other input modalities. During training, we prefix the input text with a modality token ([Text], [Image], or [Audio]) and a paradigm token ([R], [S], or [X]) to indicate the task.
Autoregressive with Dynamic Masking. One problem with image and audio masked denoising in an autoregressive manner is an information leak on the decoder side; see Figure 5 (a). The current decoder’s input token (3) is conditioned on enocoder’s information (2, 5) and all previous tokens (s 2) to predict target (4). As a result, the predicted token will be conditioned on 1 even though it was masked in the encoder since it appears in the decoder, which will simplify the task and harm representation learning. Simply masking the token in the decoder, as shown in Figure 5 (b), avoids this information leakage but causes the generation and de-noising tasks to interfere with one another. For example, we found that joint training with generation (50% MAE and 50% causal modeling) significantly reduced image generation performance. Our solution is to mask the token in the decoder except when predicting that token, as shown in Figure 5 (c), which does not interfere with causal prediction whilst mostly eliminating data leakage. For image and audio generation, we also use row, column, and conv-shaped masked sparse attention in the decoder.
4 Efficient Implementation
Training on heavily multimodal data results in highly variable sequence lengths for the transformer’s inputs and outputs, both because modalities are often missing for individual examples and because the number of tokens used to encode particular modalities can vary from just a few tokens (for a sentence) to 1024 tokens (for an output image). To handle this efficiently, we use packing, a process where the tokens of multiple examples are packed into a single sequence, and the attentions are masked to prevent the transformer from cross-attending between examples.
Typically, packing is done during pre-processing, but it is challenging in our setup since our encoders and decoder do not always support it. Instead, we do packing right before and after the transformer encoder-decoder stage, which allows the modality encoders/decoder to run on the unpacked data. During training, we use a heuristic algorithm to re-arrange data being streamed to the model so that long examples are matched with short examples they can be packed with. Packing optimization was also explored in , but not in the streaming setup. Dynamic packing leads to an almost 4x increase in training throughput (Details in Appendix B.3).
5 Optimizer
We use Adafactor as our optimizer with a linear warm-up for the first 5,000 steps and a learning rate decay of . We train with and , where is the step number. We use global norm gradient clipping with a threshold of 1.0 and find that this is crucial to stabilized training. Table 1 gives the details of our different models. For all models, we train M steps – M for pre-training and 1.5M for instruction tuning, respectively. More details in Appendix B.4.
Multimodal Data
One critical difference between Unified-IO 2 and prior work is that we train the model with a diverse set of multimodal data from scratch. This requires curating high-quality, open-source multimodal data for both pre-training (4.1) and instruction tuning (4.2).
Our pre-training data comes from various sources and covers many modalities. We provide a high-level overview and details in Appendix C.
NLP [33%]. We use the publicly available datasets that were employed to train MPT-7B . This dataset emphasizes English natural language text but also contains code and markdown. It includes text from the RedPajama dataset , C4 , Wikipedia, and stack overflow. We follow the proportion suggested by and remove multi-lingual and scientific data.
Image & Text [40%]. Text and image paired data comes from LAION-400M , CC3M , CC12M , and RedCaps . To help train the image-history modality, we also use the interleaved image/text data from OBELICS . We use the last image as the image input and the remaining images as the image history. Special tokens are used to mark where those images occur in the text.
Video & Audio [25%]. Video provides strong self-supervisory signals with high correlations between audio and visual channels. We sample audio and video data from various public datasets including YT-Temporal-1B , ACAV100M , AudioSet , WebVid-10M , HD-VILA-10M and Ego4D .
3D & Embodiment [1%]. For self-supervised 3D and embodiment pre-training, we use CroCo for cross-view generation and denoising; Objaverse for view synthesis; and random trajectories in ProcTHOR and Habitat for the next action and frame predictions.
Augmentation [1%]. While there is a lot of unsupervised data on the web for images, text, video, and audio, options are much more limited for dense and sparse annotations. We propose to solve this through large-scale data augmentation. We consider two types of data augmentation: 1. Automatically generated segmentation data from SAM to train the model to segment an object given a point or bounding box. 2. Synthetic patch-detection data which tasks the model to list the bounding boxes of synthetically added shapes in an image. We additionally train the model to output the total number of patches in the image to pre-train its counting abilities.
Training Sample Construction. During pre-training, most of our data contains various modalities without a supervised target. In these cases, we randomly pick one of the modalities present to be the target output. Then, we either remove that modality from the example or replace it with a corrupted version. Other modalities that might be present in the example are randomly kept or masked to force the model to make predictions using whatever information is left. Figure 7 shows an example when using a video that contains a sequence of image frames, the corresponding audio, and a text transcript. The pre-training sample is constructed by following the procedure: 1. select the target modality; 2. select which other input modalities to keep; 3. select the objective; 4. generate the random input mask depending on the task of denoising or generation; 5. add a prefix token indicating the task.
2 Instruction Tuning Data
Multimodal instruction tuning is the key process to equip the model with diverse skills and capabilities across various modalities and even adapt to new and unique instructions. We construct the multimodal instruction tuning dataset by combining a wide range of supervised datasets and tasks. We ensure every task has a clear prompt, either using existing ones or writing new ones. We also include open-ended tasks and create synthetic tasks for less common modalities to enhance task and instruction variety. Our mixture includes 220 tasks drawn from over 120 external datasets. We provide a high-level overview and examples here and leave details in Appendix D.
Natural Language [25.0%]. For natural language, we use the mixture from FlanV2 and various other instruction following datasets . In addition, we continue pre-training on our unsupervised NLP mixture to help prevent the model from forgetting information learned from pre-training during the extensive instruction tuning stage.
Image Generation [17.6%]. For text-to-image generation, we use the same image & text pairs we used during pre-training. We also include data from that provide better caption quality. We additionally train the model to generate images through view synthesis , image editing , segmentation-based image generation and inpainting .
Audio Generation [7.5%]. This includes text-to-audio datasets with audio in the wild , music , and human speech . We also add pre-training data with the task of predicting the next audio clip in a video. More specifically, we divide the audio into segments and then generate one of them given both the text and previous segments as input.
Image Understanding [17.8%]. We include various data sources from visual question answering , image tagging , region classification , and datasets with open-ended chat-like responses . We also include the multimodal instruction tuning datasets M3IT and MIMIC-IT .
Video Understanding [10.6%]. We include data sources from video captioning , video tagging , and video question answering . We also use examples from M3IT and MIMIC-IT for video instruction following.
Audio Understanding [10.6%]. We include data sources from audio tagging , and audio captioning . We also include data from video action classification with audio in the dataset.
Image Sparse Labelling [7.25%]. These tasks require outputting sparse coordinates based on an input image. We mainly consider object detection , referring expression , 3D detection , camera pose prediction , text detection and human keypoints .
Image Dense Labelling [4.06%]. We do several image labeling tasks, including surface normal estimation , depth estimation , and optical flow . We also train our models on various segmentation tasks, including semantic segmentation, localization segmentation, and referring expression segmentation.
Video Sparse Labelling [3.42%]. We do video detection , single object tracking and video action localization .
Embodied AI [4.33%]. For VIMA-Bench , we use the image input as the initial observation of the environment and the image history for the images or videos in the prompt. We add large-scale manipulation datasets with continuous control in both simulated and real-world environments. We also train on the PointNav task from Habitat Gibson scenes.
The distribution of the instruction tuning data is in Figure 6. Overall, our instruction tuning mixture is composed of 60% prompting data, meaning supervised datasets combined with prompts. To avoid catastrophic forgetting, 30% of the data is carried over from pre-training. Additionally, 6% is task augmentation data we build by constructing novel tasks using existing data sources, which enhances existing tasks and increases task diversity. The remaining 4% consists of free-form text to enable chat-like responses.
Experiments
In this section, we evaluate our pre-trained and instruction-tuned models on a broad range of tasks that require parsing and producing all modalities: images, video, audio, text, and actions. We do not perform task-specific finetuning in any experiments. Details about experimental setups, additional result details, results on natural language tasks, and additional studies for Unified-IO 2’s instruction capabilities are in Appendix E.
We demonstrate the effectiveness of our pre-training by evaluating Unified-IO 2 on commonsense natural language inference (HellaSwag ), text-to-image generation (TIFA ) and text-to-audio generation (AudioCaps ). We also assess spatial and temporal understanding on SEED-Bench , a benchmark for comprehensively evaluating perception and reasoning on image and video modalities. Table 2 shows that Unified-IO 2 achieves comparable or even better performance on both generation and comprehension tasks compared to the task-specific specialist or the universal multimodal model .
Despite extensive multitasking, the results on HellaSwag suggest that Unified-IO 2 has language modeling capabilities between typical 3B and 7B language models. This may be due to that the model sees far fewer tokens compared to language-based LLMs – approximately 250 billion tokens in total. Qualitative results of pre-training are in Appendix E.1.
2 GRIT Results
We evaluate on the General Robust Image Task (GRIT) Benchmark , which includes seven tasks: categorization, localization, VQA, referring expression, instance segmentation, keypoint, and surface normal estimation. Completing all 7 tasks requires understanding image, text, and sparse inputs and generating text, sparse, and dense outputs. Although this is a subset of the modalities Unified-IO 2 supports, we evaluate on GRIT because it provides a standardized and comprehensive benchmark on this set of capabilities. See Appendix E.3 for additional inference details on GRIT.
Results are shown in Table 3. Overall, Unified-IO 2 is state-of-the-art on GRIT, surpassing the previous best model, Unified-IO, by 2.7 points. On individual tasks, we can observe gains in localization (3 points), categorization (14 points), segmentation (2 points), and keypoint (5 points). On VQA, our GRIT evaluations show Unified-IO 2 is better on same-source (84.6 vs. 81.2) questions, suggesting the gap is due to reduced performance on the new-source questions that were constructed from Visual Genome; see Appendix E.3 for additional discussion. Despite being slightly behind Unified-IO, Unified-IO 2 still obtains strong referring expression scores that compare favorably to prior work on generalist multimodal models, see Table 5. Surpassing Unified-IO while also supporting much higher quality image and text generation, along with many more tasks and modalities, illustrates the impressive multi-tasking capabilities of our model. Unified-IO 2 even maintains better overall performance with the 3-billion parameter model (65.2 vs. 64.5), which is roughly equal in size to Unified-IO. Ablation results show average performance, and all individual tasks improve with model size, showing that Unified-IO 2 benefits from scale.
3 Generation Results
Table 4 shows results on tasks that require generating image, audio, and action outputs. We evaluate using TIFA , which measures faithfulness to the prompt using VQA models and has been shown to correlate well with human judgments, and FID on MS COCO . On TIFA, we find that Unified-IO 2 scores close to minDALL-E , and about 10 points ahead of other generalist models such as CoDi and Emu . We attribute this strong image generation ability to extensive pre-training and the use of a fine-grained VQ-GAN. We include examples of our generation results from the TIFA benchmark in the Appendix E.5. Unified-IO 2’s FID scores are slightly higher than the compared models, although we note that qualitatively the generated images are still very smooth and detailed.
For text-to-audio generation, we evaluate on the AudioCaps test set. AudioCaps consists of 10-second audio clips, while our model can generate 4.08-second audio at a time, so we cannot do a direct evaluation on this benchmark. Instead, we generate an audio segment based on the text description and previous audio segments as additional input; see Appendix E.6 for more details. While this is not a directly comparable setup to related work, it still gives a reasonable quantitative measure of our audio generation abilities. Unified-IO 2 scores higher then specialist models except the recent latent diffusion model , which shows it’s competitive audio generation ability.
For action, we evaluate using VIMA-Bench , a robot manipulation benchmark containing 17 tasks with text-image interleaved prompts. Since VIMA’s action space is action primitives, Unified-IO 2 directly predicts all actions at once given the initial observation and multimodal prompt. We report the average success rate for 4-level evaluation protocol and compare with the original casual VIMA policy with object-centric inputs, as well as VIMA-IMG, a Gato -like policy with image inputs like ours.
4 Vision Language Results
We evaluate vision language performance and compare it against other vision/language generalist models, i.e., models that are also designed to perform many tasks and can follow instructions. Results on a collection of 12 vision/language benchmarks are shown in Table 5. SoTA results from specialist models are shown for reference.
Unified-IO 2 achieves strong results on VQA, only passed by the much larger 13B LLaVa model on VQA v2 , and ahead of all other generalist models on ScienceQA and TallyQA . OK-VQA is the exception. We hypothesize that because it requires external knowledge, extensive language pre-training is important for this task, and therefore our reduced performance is since Unified-IO 2 was not pre-trained as extensively on text as the dedicated language models used by Qwen-VL and mPLUG-Owl2 .
On referring expression, Unified-IO 2 is ahead of Shikra and Ferret and matches the scores achieved by Qwen-VL. On captioning, Unified-IO 2 also achieves a strong CIDEr score of 130.3, ahead of Shikra and InstructBLIP but behind Qwen-VL and mPLUG-Owl2.
Finally, we evaluate using three recently proposed evaluation-only benchmarks. MMB (MMBench ) tests multiple facets of vision language understanding with multiple choice questions, while SEED-Bench additionally tests video understanding. We show a detailed breakdown of our score in the Appendix E.4. Regarding the overall score, Unified-IO 2 has the strongest score of any 7B model on the SEED-Bench leaderboardas of 11/17/23, and scores the highest on MMB by 3.8 points. Notably, it excels LLaVa-1.5 13B model in both benchmarks. Unified-IO 2 also reaches 87.7 on the POPE object hallucination benchmark , showing that it is not very prone to object hallucination.
Overall, Unified-IO 2 can match or surpass other vision & language generalist models on these benchmarks despite encompassing many more modalities and supporting high-quality image and audio generation. This shows that its wide breadth of capabilities does not come at the expense of vision/language performance.
5 Video, Audio and other Results
Unified-IO 2 shows reasonable performance on audio and video classification and captioning, as well as video question answering, as shown in Table 6. Notably, Unified-IO 2 outperforms BLIP-2 and InstructBLIP on Seed-Bench Temporal by 8.5 points. Unified-IO 2 also achieves better performance on Kinetics-Sounds than MBT , which is trained solely on that dataset.
We show the single-object 3D detection results in Table 7. Our model shows decent results, similar to Cube-RCNN , on the Objectron benchmark . However, its performance drops significantly in multi-object 3D detection tasks, like those on nuScenes and Hypersim . This could be because only 1.0% of our training data focuses on 3D detection. A potential solution might be to combine 2D and 3D detection techniques.
In COCO object detection, excluding the ‘stuff’ categories, our model reached an average precision (AP) of 47.2, with AP50 at 57.7 and AP75 at 50.0. However, it has difficulties with images containing many objects. Previous research, like Pix2Seq , suggests that autoregressive models face similar challenges, which can be improved with extensive data augmentation. Our model’s data augmentation on object detection is comparatively more limited.
Our model shows weak performance in depth estimation, with an RMSE of 0.623 on NYUv2 depth dataset . However, fine-tuning specifically for this task improved the RMSE to 0.423. In our experiment, we simply normalize the depth map with the max depth value in each dataset. Due to the incompatibility of dense ground-truth depth across different datasets , our model failed to capture the exact scale in the current prompt, which could potentially be solved by using better normalization and metric evaluation.
Appendix E shows qualitative visualizations of other tasks, such as single object tracking, future state prediction of robotic manipulation, and image-based 3D view synthesis, etc. missing
Limitation
Due to memory constraints, we use the base versions of the ViT and AST models for image and audio features throughout the project. Using a larger version of these image and audio encoders could substantially improve performance.
While our image generation is more faithful compared to SD-based methods, its quality doesn’t match that of the stable diffusion model. Additionally, our audio generation is capped at approximately 4 seconds, which restricts the practical application of the audio outputs.
Limited computational resources constrained our exploration of the model’s hyperparameters. It’s likely that using a significantly larger batch size could enhance the model’s performance.
Our model is much less reliable for modalities like depth, video or when requiring more niche abilities like 3D object detection, etc. This is probably due to the limited variety of tasks we have in these areas.
Improving the quality of our data could enhance the model’s performance. However, despite considerable efforts, our human-written prompts still fall short in diversity. We notice a notable decrease in the model’s performance when dealing with new instruction tasks, as opposed to those it was trained on.
Conclusion
We introduced Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. This model was trained from scratch on a wide range of multimodal data and further refined with instruction tuning on a massive multimodal corpus. We developed various architectural changes to stabilize the multimodal training and proposed a multimodal mixture of denoiser objective to effectively utilize the multimodal signals. Our model achieves promising results across a wide range of tasks. We show that going from LLMs to LMMs enables new capabilities and opportunities. In the future, we would like to extend Unified-IO 2 from the encoder-decoder model to a decoder-only model. Additionally, we plan to expand the model’s size, enhance the data quality, and refine the overall model design.
Acknowledgement We thank Klemen Kotar for helping gather Embodied AI pre-training data, Jonathan Frankle from MosaicML for suggesting the mixture of NLP pre-training data, Jack Hessel for interleaved image & text dataset and Micheal Schmitz for helping support the compute infrastructure. We also thank Tanmay Gupta for helpful discussions, as well as Hamish Ivison, and Ananya Harsh Jha for their insightful discussions about model design. We additionally thank Oscar Michel, Yushi Hu and Yanbei Chen for their help editing the paper, and Matt Deitke for help setting up the webpage. Savya Khosla and Derek Hoiem were supported in part by ONR award N00014-23-1-2383. This research was made possible with cloud TPUs from Google’s TPU Research Cloud (TRC).
References
Appendix A Contributions
Jiasen Lu, Christopher Clark, Sangho Lee, and Zichen Zhang collectively contributed to dataset construction, prompt development, and conducted numerous exploratory experiments for this project.
Jiasen Lu led and designed the main idea and scope of the project. Developed the majority of the model pipeline – image and audio tokenizer, main architecture, model stabilization, and training objective. Led and designed the pre-training and instruction tuning data pipelines. Conducted experiments with various model and data hyperparameters, oversaw the model training process, and wrote the paper. Coordinate with the whole team.
Christopher Clark co-led and designed the infrastructure, instruction tuning, and evaluation. Developed the dynamic packing system, modality processing pipeline, and classifier-free guidance for image and audio inference. Added the NLP and many V&L datasets, added many synthetic tasks, and built prompts for instruction-tuning tasks. Ran the evaluation in § 5.1 (NLP), 5.2, and 6 (detection, depth) and wrote the paper.
Sangho Lee core contribution to the pre-training data pipeline. Added all large-scale multimodal pretraining datasets, and video and audio instruction tuning datasets. Developed sample construction pipeline for pre-training. Helped with the model implementation – position encoding, perceiver resamplers, and model stabilization. Ran the evaluation in § 5.1 (audio), 5.3 (audio and image FID), 6 (video and audio understanding) and wrote parts of the paper.
Zichen Zhang core contribution to the instruction tuning data pipeline. Added many V&L, embodiment, video, audio, data augmentation, and all instruction tuning datasets. Built prompts for instruction tuning. Investigated the model architectures and training pipelines and stabilized the training. Ran the experiments in § 5.1 (image TIFA, SeedBench), 5.3 (image TIFA, action), 5.4, wrote parts of the paper, developed the model demo and project page.
Savya Khosla added 3D object detection, optical flow, and multi-point tracking datasets, ran the evaluation of 3D detection, and initiated the demo.
Ryan Marten added part of video and tracking datasets.
Derek Hoiem advised on the research direction.
Aniruddha Kembhavi advised on the research direction and evaluation, helped manage compute resources and wrote the paper.
Appendix B Model Implementation Details
In this section, we present the implementation details of our model.
First, we provide details about how different modalities are represented in our model.
Text representation. The Byte Pair Encoding (BPE) vocabulary size is 32000. Similar to , we add 200 additional special tokens to indicated masked spans when de-noising. We further add 10 special tokens that can be used to reference the image, audio, and history input in the text. Two special tokens are to indicate the and , and 8 special tokens represent individual elements in the image and audio history inputs, both of which have a maximum of 4 frames. We use a maximum of 512 input and output tokens.
Sparse structures representation. We use an additional 1000 special tokens to represent all continuous values, such as points, boxes, camera transformation, and 3D cuboids. Points are represented with coordinates and boxes with coordinates with values normalized by the image size. Camera transformations are represented as polar angle , azimuth angle , and distance . 1000 special tokens to represent discretized angle from to . Following , 3D cuboids are represented with 12 parameters including projected center , virtual depth , log-normalized box dimension , and continuous allocentric rotation .
represent the projected 3D center on the image plane relative to the 2D RoI
For 3D cuboid detection, we use prompts to indicate the target format, such as “Locate all objects in 3D using projected 3D center, virtual depth, log-normalized box size, and rotation in the image.”
Action representation. For embodied navigation tasks, the discrete action space is directly represented as texts, e.g. “forward”, “left”, “right”, “stop”. For object manipulation tasks, the action is represented differently based on the robots. Overall, the positional change (e.g. ), rotational change (e.g. ), and gripper open or close are discretized using the same 1000 special tokens, and we use the text prompt to indicate the input and target format. For tasks that require multi-step planning (e.g. VIMA ), the actions are represented as human-readable texts with the indication of steps, skills used (e.g. pick, place, or push), and discretized positional and rotational parameters. Figure 11 provides a detailed illustration of the robot tasks.
Images representation. Images are encoded with a pre-trained ViT . We use the ViT-B checkpoint trained on LAION 2B datasethttps://github.com/mlfoundations/open_clip. For image inputs, we use a maximum length of 576 tokens (i.e. patch encoding from a image). We concatenate features from the second and second-last layers of the ViT to capture both low and high-level visual information. To generate the image, we encode these images as discrete tokens . Different from Unified-IO , which uses the VQ-GAN trained on ImageNet to convert resolution image into tokens, we use the VQ-GAN trained on the Open Images dataset with a compression ratio of 8 and a vocabulary size of 16384https://github.com/CompVis/taming-transformers. This converts resolution image into tokens. We also compare the VQ-GAN tokenizer with the ViT-VQGAN and MoVQ . We empirically find VQ-GAN leads to best generation results.
Dense structures representation. To handle this modality, we convert per-pixel labels into RGB images. For depth, we construct a grayscale image by normalizing the depth map. For surface normal estimation, we convert the orientations into values. For segmentation, we train Unified-IO 2 to predict a single black-and-white mask for a particular object specified by a class and a bounding box. Instance segmentation (as done in GRIT ) can then be performed by first performing localization for the target class and then performing segmentation for each detected box. Unified-IO instead trains the model to produce an image with a randomly selected color for each instance. We found this makes post-processing difficult since output images sometimes do not exactly follow the color scheme, and the model could struggle with images with many different instances.
Audio representation. This modality encodes a 4.08-second segment of audio. We take the waveform sampled at 16000 Hz and convert it to a log-mel-scaled spectrogram. We compute the spectrogram for an entire audio segment (4.08 seconds) simultaneously. Each window involves 1024 samples and 256 samples ‘hops’ between windows. The resulting spectrogram has a size of 128 mel bins with 256 windows. We chose these hyperparameters largely around efficiency. We then encode this with a pre-trained AST with the patch size of , hence a total of 128 tokens.
To generate audio, we use ViT-VQGAN to convert the spectrograms into discrete tokens. Since the authors of did not release the source code or any pre-trained models, we implement and train our own version of ViT-VQGAN with patch size that encodes a spectrogram into 512 tokens with a codebook size of 8196. The model is trained with the audio on AudioSet , ACAV100M , and YT-Temporal-1B datasets. After getting the log-mel-scaled spectrograms, we use HiFi-GANhttps://github.com/jik876/hifi-gan vocoder to decode the spectrograms back to waveforms. We train the HiFi-GAN using the same parameters shown in Table 8. We trained the model on a mixture of AudioSet and LJSpeech to cover natural sound and human voice.
History representation. Images and audio inputs in this history are first encoded in the same way as image and audio inputs. We then use a perceiver resampler to further compress the image and audio features and produce a fixed number of visual outputs (32) and audio outputs (16) to reduce the total sequence length of the model. As shown in Table 8, we consider a maximum of 4 images and audio segments. In our experiments, we test with two different variants of perceiver implementations: 1) a small group of latent embeddings query each frame/segment individually , 2) a large group of latent embeddings query all history at once. While the second implementation can finely represent the referenced image and audio, the first can preserve better temporal information. Thus, our final implementation uses the first one.
B.2 2D Rotary Embedding
We use a rotary position encoding to model the relative location of input sequences . We chose this primarily because we did not want to use absolute (additive) position embeddings, which would have to be added to the inputs of each encoder, and also wanted to be consistent with the LLaMA position encoding.
The rotary encoding uses no parameters and instead uses a kernel trick to allow the model to recover relative distances between key and query elements in a transformer’s attention head. For text, we apply rotary encoding at each layer of the network. For other modalities, we extend RoPE to two-dimensional cases by splitting each of the query and key embeddings of transformer attention heads in half and apply separate rotary embeddings constructed by each of the two coordinates to the halves.
We treat each token (image, audio, image history, and audio history) as having a 2-dimensional position corresponding to 1) coordinates in the image or audio spectrogram, 2) where and represent the indices of frame and perceiver latent vector in the image or audio history, respectively. Different from , which uses a 4-dimensional position to represent all the inputs, we use a combination of learnable segment (modality) embeddings and rotary encoding.
B.3 Dynamic Packing
Here, we describe the dynamic packing algorithm in more detail. As is standard practice, when batching together inputs, we pad input tensors to a maximum length and use attention masked to prevent the transformer from attending to padding elements. This, however, is highly inefficient in our multi-modal setting because many modalities are not present in most examples, which results in a huge amount of padding. For example, if one example in the batch has an image output, every other example must be padded with 1024 target image tokens, even if their output is in a different modality.
One solution is to arrange batches so that each batch contains examples with similar numbers of tokens in each modality. This is, however, complicated to do in practice since (1) our data does not fit in RAM, so we cannot easily sort and group data this way, especially if needing to match tokens across five input and three output modalities and (2) our coding framework, JAX , does not support variable length tensors when constructing the execution graph which makes handling variable lengths between batches extremely difficult.
Instead, we use packing, a process where the tokens of multiple examples are packed into a single sequence, and the attentions are masked to prevent the transformer from cross-attending between examples. Packing is often done as a pre-processing step when handling text, but this does not work in our setup since some parts of our network cannot operate on packed data (e.g., the VAE or image ViT). Instead, we start with an unpacked batch of examples, run these components first, and then dynamically pack the resulting tokens in a backdrop-compatible way before running the transformer. To run efficiently on TPUs we pack examples using matrix multiplication with carefully constructed one-hot matrices.
To account for all modalities, the maximum sequence length our transformer needs to take as input is 1152, and the maximum target length is 2048. When packing, we can generally pack two examples into an input sequence of 864 and a target sequence of 1280, which gives a roughly 4x speed up due to reduced sequence length and the ability to process two examples simultaneously. When streaming data, packing cannot be done reliably. For example, if two consecutive examples have an image output, they cannot be packed since they will total over 1280 output tokens. To handle this, we use a heuristic algorithm to re-arrange data as it is being streamed. The algorithm keeps a small pool of examples in memory. Given a new example, it pairs it with the largest example in the pool it can be packed with and outputs both as a pair. If no such example exists, it adds the example to the pool. If the pool reaches a maximum size of 10, the largest example is emitted and processed without being packed with another example. We find this occurs less than 0.1% of the time during training.
B.4 Full Model Details
In Table 9, we present the full hyperparameters of our model. During pre-training, we train the UIO-2, UIO-2, and UIO-2 with a batch size of 512 due to memory limit. We sub-sample 50% of the image, audio, and history inputs patches. The total packing length is 864 for the encoder and 1280 for the decoder. During instruction tuning, we train all of our models with a batch size 256 due to computing constraints. We sub-sample 87.5% of the image, audio, and history input patches. The total packing length is 1024 for pretraining and 1280 for instruction tuning. 8-way in-layer parallelism and 64-way data parallelism were used to scale up to the 7B model training.
We train for 1.5 million steps with an effective batch size of 512. This results in training on approximately 1 trillion tokens. During pre-training, we keep at most 50% of the image patches in the image history or image encoder, as is common practice with MAE pre-training . We use up to four images/segments in image/audio history.
Appendix C Pre-Training Details
In this section, we provide additional details about the data Unified-IO 2 is pre-trained on. The datasets we use for pre-training are listed in Table 10. Unless otherwise specified, we use the pre-training objective described in Section 3.3, where one of the present modalities is randomly selected as the target. We sample data to ensure all the output modalities are well represented and to balance how often our various corpora are used based on their size. The distribution is shown in Figure 9.
Text. Our data follows the mixture used by MPT-7B .
Image & Text. Image & text paired data comes from various unsupervised corpora, shown in Table 10. For LAION data, we only generate images from image/text pairs from LAION aesthetic, which contains higher quality images, while we generate text for image/text pairs from LAION 400M. We also only keep images from LAION if they are marked as being unlikely to be NSFW in the LAION metadata. Web images is a dataset of images we download and focuses on icons and stylized images.
Video. We gather a total of 180M short videos from various sources. During training, we pick a random sequence of up to five frames from the video. The first four will be encoded with an image/audio history encoder, while the fifth frame will be encoded with the image/audio encoder. The text matching these frames is encoded with a text encoder along with marker tokens to show where each frame occurred as stated in B.1, or, if the dataset only includes a single caption that is not aligned with individual frames, the entire caption is encoded instead. The text, audio, or image modality can be selected as the target modality. As usual, other modalities are randomly masked, and the target modality is randomly masked or injected with noise in the input. Note we have sub-sampled data from many of these corpora to keep the dataset size more manageable, and sometimes due to broken video links.
Interleaved Image & Text. We primarily use OBELICS , which contains paragraphs and images interleaved together. For each document, we randomly select an image or a paragraph as the target and use up to the previous four (if the target is an image) or five (if the target is a paragraph) images as context. The last image is encoded with the image encoder, and the remaining images are encoded in the image history. The text matching those images is concatenated and interjected with marker tokens to indicate where the images in the image history or image input occur. We either do de-noising, where a noisy version of the target is included in the input, or generation, where the target is not part of the input, although we always include both the text and image input modalities.
In addition, we construct interleaved data by interleaving multiple images and captions from several image/text pair corpora. The images are encoded as the image input and/or the image history, and matching text is constructed by specifying the caption for one, or all, of these images using special tokens to mark which image each caption refers to. For this task, we only target the text modality, and train the model to either (1) de-noise the caption of a single image, (2) generate a caption for a single image that is specified in an input prompt using a marker token or (3) generate a sequence of marker tokens and captions that describe each input image. This task aims to ensure the model learns the semantics of the images in the history and understands the marker tokens.
Multi-View. We train on the cross-view completion task from CroCo , where the model must complete a heavily noised image using an image of the same scene, but from a slightly different angle, as context. The noised input is encoded as an image and the second image is encoded through the image history encoder. In addition, we generate data using Objaverse objects by capturing multiple views of the object in 3D, and either specify the camera coordinates in the input text and train the model to generate a new image matching new camera coordinates, or train the model to predict how the camera has moved between different images. We further augment the view synthesis task by providing in-context examples. For example, by giving one or more examples of the views and transformations in the image history, the model predicts the new view from the new camera transformation specified by the prompt. Both tasks aim to improve the model’s 3D understanding during pre-training.
Agent Trajectory. We use scripted shortest path trajectories in ProcTHOR and human-collected demonstrations in Habitat . While the original datasets are for object navigation with relatively long episode lengths, we only subsample from the last few frames for image history and image input such that mostly the target object is within the observation. The task is randomly selected from 1) generating the next visual observation frame as the target image, 2) predicting the next positional observation coordinates as the text target, and 3) predicting the next action as the text target. 1) requires inferring from the image and image history input and the last action specified in the text input, 2) further requires the location information, and 3) is based on the target object name and visual observations for the next action prediction.
Synthetic. We add two synthetic tasks. First, we use the automatically annotated data from Segment Anything . We give the model either a set of points or a bounding box as input and train it to generate a segmentation mask as output. Second, we add artificial patches of various shapes and colors to images from other unsupervised datasets and train the model to output their locations in order to train the model to generate sparse coordinates as output. We additionally train the model to output the total number of patches on the image to pre-train its counting abilities.
Appendix D Instruction Tuning Details
In this section, we provide additional details about the instruction tuning data and individual tasks Unified-IO 2 supports. An overview of the instruction tuning data is shown in Table 11. We show a visualization including individual datasets in Figure 10. We sample broad categories of tasks evenly and then generally sample individual datasets in proportion to the square root of their size, although with some minor hand-engineered adjustments to downweight noisy datasets or upweight very rare tasks.
For natural language data we use the mixture from FlanV2 , which in turn includes data from Muffin , T0-SF , NIV2 , and CoT annotations, as well data from Alpaca , Dolly , Open Assistant , and MDPP . In addition, we continue pre-training on our unsupervised NLP mixture from our fine-tuning stage to ensure the model does not forget information learned from unsupervised data during the extensive instruction-tuning stage.
D.2 Image Generation
For text-to-image generation, we use the same image/text pairs we used during pre-training, as well as localized narratives from Open Images and captions from COCO and Visual Genome (VG) . Our prompts for these tasks specify that the image might be noisy or approximate for unsupervised corpora (e.g. “Generate an image that roughly matches this text: {caption}”) and give hints as to the style for supervised corpora (e.g. “What do you see in this image? Plainly describe the individual element you observe.” for localized narratives) to help disambiguate the stylistic differences between the datasets. We use simple prompts (e.g. “Caption this image.”) for the COCO captions.
We additionally train the model to generate images through view synthesis as was done during pre-training. We also integrate data for image editing and image editing based on various dense control signals such as depth maps, edges, segmentation, etc. Following , and the segmentation-based image generation from Unified-IO using data from COCO and LVIS . Finally, we train on inpainting by masking a region of an input image that contains an object and training the model to generate the complete image given the object name and location. We derive data for this task from the object annotation data in COCO, Open Images, and VG.
During inference, we use top-p sampling, also known as nucleus sampling , for generating images with the temperature and . We also enable classifier-free guidance by replacing the prompt with the un-informative prompt “An image of a random picture.” 10% of the time during training. That prompt is then used as the classifier-free prompt with a guidance scale of during inference.
D.3 Audio Generation
Datasets for audio generation from text include AudioCaps , Clotho , MACS , MusicCaps , and LJSpeech . During training, we divided the audio into 4-second-long segments and then generated one segment of the target audio, giving both the text and any previous segments as input. We also train on the next-frame prediction task, which aims to generate the audio for the next frame in a video from YT-Temporal-1B .
Our prompts for these tasks specify the characteristics of target audio; e.g., “Generate the sound/music based on the description: {caption}” for natural sound and music, respectively, and “Speak: {passage}” for speech. We use the same sampling method as the image generation, the top-p sampling with the temperature and . We do not use the classifier-free guidance because it can lead to poor performance. When generating audio longer than 4.08 seconds during inference, we generate an initial segment that is 4.08 seconds long and then extend it by generating additional segments using previous audio segments as the audio history input.
D.4 Image Understanding
These tasks require generating text in response to a query about an image or a pair of images. We use the data from M3IT and MIMIC-IT , as well as a variety of other additional sources. For VQA, we add GQA , TallyQA , OK-VQA , A-OKVQA , OCR-based VQA datasets , Visual Genome, ScienceQA , VCR and VizWiz . For image tagging we add Caltech Birds , iNaturalist , Sun397 , and Places365 . For region classification, we add examples derived from object annotation from Open Images, VG, and COCO. We categorize datasets with open-ended responses such as LLaVa , Visual Storytelling , and Visual Dialog as visual instruction following, and we categorize NLVR and the “spot the differences” tasks from MIMIC-IT as image pair QA. For image pair QA tasks, we encode the second image in the image history modality.
We also add a grounded relationship prediction task using data from Visual Genome and VSR as well as image captioning using the same supervised sources we use for image generation.
We again put stylistic hints in the prompts for these tasks. For example, in VQA and captioning datasets, we specify to return a short answer (e.g. “Answer this question very succinctly: {question}”), which we find is critical to allow the model to produce longer, more natural responses when asked user questions. Likewise, we roughly specify the kind of class to output for image tagging, e.g., “”What is the scientific name of this animal?” for the iNaturalist dataset.
D.5 Image Sparse Labelling
These tasks require outputting sparse coordinates based on an input image. We use Open Images, Visual Genome, and COCO for object detection and localization, which requires detecting all objects belonging to a specific class and three COCO referring expression datasets for referring expressions.
In addition, we train on the OmniLabel 3D detection dataset by generating the projected 3D center, virtual depth, log-normalized box size, and rotation of each 3D box, again by normalizing these values between 0 and 1 and then encoding them using the location tokens. We also added the camera pose prediction tasks using Objaverse objects that were used during pre-training.
We include 3 text detection datasets from COCO-Text , including finding the bounding box of an input text string for multiple text strings or finding and listing all text along with their bounding boxes in an image.
Lastly, we do keypoint detection using COCO pose data. For keypoint detection, we input a bounding box around a person in the image and train the model to return a list of keypoints for that person. During inference, we first localize all people in the image and then use each returned bounding box as a keypoint query to find that person’s keypoints. During training, the model predicts “MISSING” for keypoints that are not visible (e.g. “right elbow: MISSING”). During inference, we use a masking function over the model’s logit to force it to guess a valid point for each keypoint since the keypoint metric does not award points for correctly identifying a keypoint as being not visible.
D.6 Image Dense Labelling
We do several image labeling tasks, including surface normal estimation on FramNet , BlendedMVS and Taskonomy , depth on NYU Depth , and optical flow on Flying Chairs and MPI Sintel .
We additionally train on several segmentation tasks: semantic segmentation (segmenting a particular class), localization segmentation (segmenting an object in an input bounding box), and referring expression segmentation (segmenting an object matching a referring expression). Data comes from Open Images, COCO, LVIS, and referring expressions from the COCO refexp datasets . To do instance segmentation, as needed for GRIT, we first do localization on the target class and then perform localized segmentation on each returned bounding box.
During inference, we do temperature sampling with a top-p of 0.95 as before, but without classifier-free guidance. For segmentation, we find it beneficial to increase the value of p to 0.97.
D.7 Video Understanding
These tasks require generating text in response to a query about a video. For video captioning, we add VATEX and MSR-VTT . For action classification (video tagging), we add UCF101 , Kinetics-710 , Something-Something v2 and EPIC-KITCHENS-100 . We also use examples from EPIC-KITCHENS-100 for action anticipation. For video question answering, we add MSRVTT-QA , MSVD-QA , STAR and M4-ViteVQA . Lastly, we use examples from M3IT and MIMIC-IT for the video instruction following.
To cover the visual content of the entire video with a small number of frames (5), we use the segment-based sampling following ; we first divide the video into five segments of equal duration and then randomly sample one frame from each of the segments during training, and the middle frame at inference. We use the first four frames as the image history input and the final frame as the image input for action classification and video captioning. We empirically found that using the third frame as the image input while using the other frames as the image history input performs better for video question answering.
We use similar prompts to those for image understanding tasks, e.g., “Write a short description of this video.”, “The question {question} can be answered using the video. A short answer is” and “What are they doing in this video? Short answer:” in video captioning, video question answering, and video tagging, respectively, for ensuring a short answer.
D.8 Video Sparse Labelling
We do single object tracking and spatial-temporal action localization on video data. We train on YouTube-BB , LaSOT and GOT-10k by inputting bounding boxes around a target object in each of previous frames and having the model return the next location as a bounding box (“Anticipate the object’s next location from all previous images and the location of the object in those frames: {locations}.”). We also train the model on AVA by inputting a video snippet consisting of five frames and requiring the model to detect all actions of humans appearing in the middle (third) frame of the video snippet (“Given the temporal context from the video, detect all of the humans performing actions in the image.”). Note that we provide the video snippet, not a single video frame, because some of the actions require temporal context to answer (e.g., stand and sit) correctly. We use the final/middle frame of five consecutive frames in the video as the image input and the other frames as the image history input for single object tracking and action localization, respectively.
D.9 Audio Understanding
We train the model on audio tagging and audio captioning tasks. For audio tagging, we add AudioSet , VGG-Sound , and MACS. For audio captioning, we use the same datasets as text-to-audio generation, that is, AudioCaps, Clotho, MACS, MusicCaps, and LJSpeech. For audio-visual action classification, we train on Kinetics-Sounds and VGG-Sound.
We again use stylistic hints in the prompts for these tasks. For example, we specify the characteristics of target audio (e.g., “Describe the music.” and “Transcribe the audio to text.” for MusicCaps and LJSpeech, respectively), enforce a short answer (e.g., “What is this in the audio? Short answer:” and “Give a short description of this audio.”), and specify the kind of class to output for audio tagging, e.g., “This audio depicts a scene of a” for MACS. We use the same prompts as video tagging for audio-visual action classification.
We use the same sampling strategy as the video understanding; we sample five audio segments with uniform intervals from the whole audio and use the middle/final audio segment as the audio input while using the other segments as the audio history input for audio classification and audio captioning, respectively.
D.10 Embodied AI
With the action representation described in B.1, we seamlessly add large-scale manipulation datasets Language Table , BridgeData V2 , and FrankaKitchen with the continuous control in both simulated and real-world environments. The model directly predicts the next action as the text target based on the current observation as image input, previous frames as image history, and language instruction and previous actions as text inputs.
D.11 Task Augmentation
In addition to these sources, we derive several additional tasks that use the same supervised annotations as other tasks but require performing slightly different functions. We call this task augmentation. The new tasks include prompts that specify the desired output. These tasks serve to add diversity to our instruction following data. We review the task augmentation we construct below.
Segmentation. We build several augmentations of the segmentation tasks, including (1) segmenting pixels belonging to one of a set of 2-4 categories, possibly including categories that do not exist in the image, (2) segmenting pixels belonging to a class and are within an input bounding box and (3) build a map of pixels that do not belong to a set 1-4 classes. Prompts are designed for these that state the requirement, e.g., “Show pixels that are part of chair, paper and in
Detection and Referring Expression. For detection, localization, and referring expressions, we also train the model to output various properties of the output bounding boxes instead of the boxes themselves. Properties include the width, height, area, left/right/top/bottom half, center coordinates, distance from the left/right/top/bottom edge of the image, or the coordinates of different corners of the bounding box. We also change the format of the output bounding box (e.g., instead of format), and change whether the model labels the boxes with the object category or not.
For detection, we train the model to detect any object belonging to a set of 1-4 classes. For referring expressions, we train the model to locate multiple referring expressions from a single query. In this case, we sometimes train the model to predict a property of both referenced boxes instead of outputting the directly, for example, which box is the smallest, which is the largest, the area of intersection, a box containing both boxes, etc.
Relationship Prediction. We train the model to list all relationships between a particular object in the image and any other object. A bounding box and category specify the target object. Similarly, we train the model to predict all relationships between any instance of a particular class of objects and any other object in the image.
Captioning. For captioning, we train the model to generate a caption that is longer or shorter than a given character or word length or contains a particular word or set of words. We also randomly require the caption to start with a particular prefix. Again, these requirements are specified in the prompt, for example, “Generate a caption longer than five words for this image. Start your output with the text ‘My caption is:’”.
Surface Normal Estimation. For surface normal estimation, we train the model to generate RGB images that encode the pixel orientation differently. This includes changing which RGB channels correspond to the x, y, and z orientations and only including a subset of those orientations. We also include tasks that require specifying the x, y, and z orientation at a particular point specified in the prompt using location tokens. Finally, we include tasks requiring segmentation masks over pixels with particular orientations, e.g., “Build a binary mask over surfaces with an upward orientation”.
Embodied AI. We further augment the embodiment datasets with the video QA and goal image generation tasks. The QA augmentation aims for the robot’s planning and affordance. For example, given a robot video trajectory, the model is supposed to predict the plan (caption), or whether a given action is reasonable from the language instruction. Applying image editing in embodied space, we further let the model generate the goal or subgoal images based on the initial visual observation in the image input and the language prompt in the text input. While recent works show that embodiment QA with VLM and (sub-)goal generation with diffusion model are effective in the decision-making downstream tasks, our model combines the both augmentation strategies.
Appendix E Experiment Details
In the main paper, we evaluate the effectiveness of our pre-training by evaluating Unified-IO 2 quantitively on a variety of benchmarks. Here, we qualitatively show the visualizations from the pre-trained UIO-2model. Table 12 shows audio generation from text (top) and text + video (bottom). We can see the pre-trained model learns text-to-speech synthesis through video pre-training, and the model can also synthesize music that matches the video input. Figure 12 shows the future frame prediction samples given the initial input image and action sequence. Figure 13 shows the image generation samples given prompts. The model has a good understanding of different objects. However, it struggles to generate the correct text from the given caption.
E.2 NLP Results
We present results on a set of NLP tasks to evaluate the model’s language understanding abilities. We evaluate using the EleutherAI LM-Eval harness , tasks are evaluated zero-shot using the default prompts without any adjustments aside from adding the [Text] [S] prefix used for all text generation tasks. We evaluate on HellaSwag and a selection of other question answering benchmarks: MMLU , ARC , and BoolQ . Results are shown in Table 14. Baselines were evaluated in the same setting, i.e., zero-shot, with the default prompts, and using LM-Eval. Unified-IO 2 is generally ahead of Open LLaMA 3B but behind LLaMA.
E.3 GRIT Details
We present GRIT results in more detail in Table 13. Notably, Unified-IO 2 is the first unified model to pass the Masked R-CNN baseline for localization and goes a long way toward closing the gap between SAM and unified models on segmentation.
For GRIT VQA, looking at the scores from GRIT on different VQA subsets, we find that Unified-IO 2 does better on the same-source subset (84.6 vs 58.5) but worse on the new-source subset (57.7 vs 67.2). Same-source questions come from VQA 2.0, and new-source questions come from VG, so the difference can be attributed to the kinds of questions being asked. Qualitatively, it is hard to understand why the scores differ on these subsets since the GRIT ablation questions lack ground truth annotations. However, we notice the models often produce different answers when faced with ambiguous questions (e.g. “What color is black on the horse”, “hair” for Unified-IO vs. “mane” for Unified-IO 2), so one possibility is that Unified-IO 2 does not match the VG answer style as well as Unified-IO, which would likely be due to differences in the kind of VQA training data the models were trained on.
For GRIT localization, we find the model can struggle with images with many instances of the target class, particularly when using beam search. We hypothesize that this is because the probability mass can get split between many similar location tokens, resulting in EOS becoming the most probable token even if its probability is low. As a solution, during inference, we only output EOS if the EOS token itself has a probability of over 0.5, which we find significantly improves the performance on crowded images. In rare cases, we observe this leads to the model generating bounding boxes for the same instance multiple times. As a solution, we apply Non-maximum suppression with a higher threshold of 0.8 to remove these duplicates. We apply this inference trick for localization and when doing the initial localization step for the keypoint and segmentation tasks.
E.4 Multimodal Benchmark Details
We now provide the breakdown results for the evaluation-only multimodal benchmarks, POPE and SEED-Bench . POPE is the object hallucination benchmark, requiring ‘yes’ or ‘no’ answers. As shown in Table 15, our largest model achieves the highest F1 score in all 3 dimensions. Interestingly, smaller models favored ‘no’ responses, possibly due to a bias from negative examples encountered during the instruction tuning phase. SEED-Bench offers 19k multiple-choice questions with human annotations for evaluating multimodal models across 12 dimensions, including spatial (Image) and temporal (Video) understanding. As shown in Table 16, our XXL model outperforms all other 7B vision/video language models, and is even slightly better than the LLaVA-1.5 13B model. Notably, our XL (3B) model has already outperformed all other counterparts in the temporal understanding split. While recent video language models have shown proficiency in conventional video tasks like video tagging and captioning, their performance in SEED-Bench’s temporal understanding is even worse than that of vision language models, which might be attributed to their limited instruction-following capabilities.
E.5 Image Generation Details
Figure 15 shows generated images for the TIFA benchmark captions using several baselines as well as UIO-2. We use the official implementation code (Emu and CoDi ) or the images shared in the official GitHub repository of TIFAhttps://github.com/Yushi-Hu/tifa/tree/main/human_annotations (Stable Diffusion v1.5 and miniDALL-E ) for baselines. All the baselines except miniDALL-E use the Stable Diffusion decoder trained on large-scale, high-quality image datasets, generating images of high fidelity. However, they often generate images that do not fully follow the input captions while Unified-IO 2 generates faithful images.
For text-to-image generation on MS COCO , we follow the standard convention ; we evaluate on a subset of 30K captions sampled from the validation set.We use the evaluation code at https://github.com/MinfengZhu/DM-GAN Following , we generate 8 images for each caption and select the best one using CLIP text-image similarity . Despite classifier-free guidance resulting in generated images of qualitatively higher quality, the computed FID score is significantly worse compared to what would have been achieved without employing it (33.77 vs 13.39); see Figure 14.
E.6 Audio Generation Details
For text-to-audio generation, we evaluate on the AudioCaps test set. Note that we cannot do an apples-to-apples comparison with other methods because AudioCaps consists of 10-second audio clips while our model can generate 4.08-second audio at a time. Instead, we evaluate the dataset in the following setup: we first sample four 2-second audio segments, convert them to log-mel-spectrograms with zero-padding, and generate the following audio with the prompt “Generate the following sound based on what you heard and the description: {caption}”. We convert the model output, that is, a log-mel-scaled spectrogram, into a waveform using the pretrained HiFi-GAN, and compare the ground-truth audio and generated audio using computational metrics including Fréchet Audio Distance , Inception Score and Kullback–Leibler divergence. We use the same evaluation code as AudioLDMhttps://github.com/haoheliu/audioldm_eval . We show the audio generation examples in Table 17 and audio-visual qualitative examples in Table 18.
E.7 Video and Audio Understanding Details
We consider classification and question-answering tasks as open-ended answer generation and use the Exact Match (EM) to measure the performance. We also tried to formulate the classification task as multiple-choice answering and generate answers by computing the logit for each dataset label and selecting the one with the highest logit, but the performance boost was quite marginal. Note that we do not train our model directly on the Kinetics-400 ; we instead train on Kinetics-710, a mixture of three different datasets belonging to the Kinetics family, that is, Kinetics-400, 600, and 700. Our model achieves top-1 accuracy 79.1 (vs. instruction tuning only: 73.8) when further finetuning on Kinetics-400 for 5 epochs, following . For Kinetics-Sounds, leveraging both audio and visual inputs largely improves performance (audio-visual: 89.3 vs. video-only: 87.4 vs. audio-only: 38.2). For captioning tasks, we use CIDEr as the evaluation metric. Figure 17 shows the qualitative examples for video understanding tasks.
E.8 Embodiment Details
In VIMA-Bench , there are 4 levels of evaluation protocols: L1 object placement, L2 novel combination, L3 novel object, and L4 novel task. Results and comparisons are shown in Table 19. The inputs for the autoregressive transformer model VIMA are object tokens consisting of cropped images and bounding boxes; image patch tokens encoded by ViT for VIMA-Gato ; image patch tokens encoded by ViT, further downsampled by a perceiver module for VIMA-Flamingo ; and single image token encoded by ViT for VIMA-GPT . The output of those baselines is all next-step action prediction. Since our model has to predict all actions at once only with the initial observation, the task setting is then more challenging than the casual policy learning baselines. Nevertheless, our models still outperform counterparts that input image or image patches for all 4 levels and are only behind the object-centric method . In Figure 18, we show the future state prediction examples on robotic manipulation tasks. Given the input state image and natural language prompt, our model can successfully synthesize the target image state.
E.9 Other Tasks
Figure 16 shows single object tracking examples from the LaSOT dataset. Note that Unified-IO 2 does not use specific class labels for tracking and tracks small moving objects such as a table tennis paddle well. Figure 19 presents qualitative examples of 3D object detection from the Objectron dataset . As outlined in our main paper, Unified-IO 2 exhibits suboptimal performance in benchmarks for multi-object 3D detection. Additionally, Figure 20 illustrates examples of image-based 3D view synthesis using the Objaverse dataset . While the model produces coherent results, it faces challenges in accurately representing relative camera transformations.