SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Yuying Ge, Sijie Zhao, Jinguo Zhu, Yixiao Ge, Kun Yi, Lin Song, Chen Li, Xiaohan Ding, Ying Shan

Introduction

In recent years, Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in comprehending multimodal data through leveraging the strong generality of LLMs . Some pioneering work further empower LLMs with the ability to generate images beyond texts. For example, our previous work SEED-LLaMA can handle a variety of tasks and excel in academic benchmarks through unifying multimodal comprehension and generation. However, the accuracy and diversity of its generated content still fall short of real-world needs. In this work, we focus on bridging this gap through upgrading SEED-LLaMA with enhanced capabilities for real-world applications.

Specifically, in order to make a multimodal foundation model applicable in real-world scenarios, we incorporate two enhanced features: (1) understanding images of arbitrary sizes and ratios, and (2) multi-granularity image generation, encompassing both high-level instructional image generation and low-level image manipulation tasks. These attributes can form the basis for a multimodal foundation model’s effective application in an open-world context, since a multimodal foundation model has to accommodate various downstream tasks requiring different levels of visual semantics.

In this paper, we introduce SEED-X, a unified and versatile multimodal foundation model as a follow-up work of SEED-LLaMA, which seamlessly integrates the features mentioned above. After different instruction tuning, SEED-X can function as various multimodal AI assistants in the real world, capable of addressing various user needs through generating proper texts and images as shown in Fig. 1. Specifically, our instruction-tuned models can act as an interactive designer, generating images while illustrating creative intent, offering modification suggestions and showcasing visualizations based on user’s input images. Additionally, they can act as knowledgeable personal assistants, comprehending images of various sizes and providing relevant suggestions. Moreover, they can generate more diverse outputs, such as slide layouts for slide creation, and interleaved image-text content for storytelling. SEED-X signifies a notable advancement towards a versatile agent for users in the real world.

To endow SEED-X with the aforementioned characteristics, our approach incorporates (1) a visual tokenizer to unify image comprehension and generation, where its multi-granularity de-tokenization phase facilitates image generation and high-precision image manipulation, and (2) an MLLM with dynamic resolution image encoding to enable the comprehension of images with arbitrary sizes and aspect ratios. Specifically, we utilize a pre-trained ViT as the visual tokenizer and train a visual de-tokenizer initialized from SDXL to decode realistic images by taking the ViT features as input. Given the high-level ViT features as input, our visual de-tokenizer can reconstruct images semantically aligned with the original images. To realize the retention of fine-grained details of the input image to satisfy image manipulation, we further fine-tune the visual de-tokenizer to take an extra condition image as input in the latent space (See Fig. 2). The ViT features serve as a bridge to decouple the training of the visual (de-)tokenizer and the MLLM in SEED-X. The dynamic resolution image encoding divides an input image into sub-images and adds extrapolatable 2D positional embeddings to the ViT features of each sub-image, allowing the MLLM to scale to any image resolution. As for the image generation training objective, a fixed number of learnable queries are fed into the MLLM, where the output hidden states are trained to reconstruct the ViT features of the target images. During inference, the image de-tokenizer can take both the output features from the MLLM and the condition image provided by users as input, ensuring that the decoded image can possess high-level semantics that meet the multimodal instructions and retain the low-level details.

We pre-train SEED-X on massive multimodal data, including image-caption pairs, grounded image-text data, interleaved image-text data, OCR data, and pure texts. We further apply multimodal instruction tuning to align SEED-X with human instructions across various domains, utilizing both existing datasets and newly collected datasets that cover image editing, text-rich, grounded and referencing QA, and slide generation tasks. The extensive evaluations on MLLM benchmarks demonstrate that our instruction-tuned model not only achieves competitive performance in multimodal comprehension, but also achieves state-of-the-art results in image generation compared to existing MLLMs on SEED-Bench-2 .

All the models, codes, and datasets will be made publicly available. We hope our work can bring insights to the community about the potential of multimodal foundation models in real-world scenarios through unifying multi-granularity comprehension and generation.

Related Work

With the rapid development of Multimodal Large Language Models (MLLM), recent studies have been working on unified MLLMs that are capable of multimodal comprehension and generation as shown in Tab. 1. Some work perform multimodal autoregression with a unified next-word-prediction objective through using a discrete visual tokenizer. For example, SEED-LLaMA pre-trains a discrete image tokenizer, where the visual codes can be decoded to realistic images using the off-the-shelf SD model, and performs multimodal pre-training and instruction tuning through predicting the next discrete token. Some research efforts have delved into multimodal autoregression with continuous representations, where each image is tokenized into embeddings via a visual encoder, and then interleaved with text tokens for autoregressive modeling. During inference, the regressed visual embeddings will be decoded into an image by a visual decoder. Additionally, some studies enable image generation in a non-autoregressive manner through utilizing learnable queries to obtain visual representations from MLLMs, which are further fed into a image decoder to generate images. The latest work, Mini-Gemini, generates text prompts using MLLMs and then leverages the existing SDXL to generate images.

In this work, we present SEED-X, a unified and versatile foundation model, through upgrading SEED-LLaMA with enhanced capabilities for real-world applications. As shown in Tab. 1, SEED-X supports object detection and dynamic resolution image encoding for multi-granularity comprehension, and high-level instructional image generation and low-level image manipulation for multi-granularity generation, which are not realized in previous work.

Method

In SEED-X, we adopt a visual tokenizer to unify image comprehension and generation, and pre-train a multi-granularity de-tokenizer to facilitate image generation and high-precision image manipulation in a two-stage manner. In the first stage, as shown in Fig. 3 (left), we utilize a pre-trained ViT as the visual tokenizer and pre-train a visual de-tokenizer to decode realistic images by taking the features of the ViT as inputs in the first stage. Specifically, NN visual embeddings from the ViT tokenizer (N=64N=64 after average pooling) are fed into a learnable module as the inputs of the U-Net of the pre-trained SD-XL (replacing the original text features). The learnable module consists of four cross-attention layers to connect the visual tokenizer and the U-Net. We optimize the parameters of the learnable module and keys and values within the U-Net on the images from JourneyDB , LAION-Aesthetics , Unsplash , and LAION-COCO . As shown in Fig. 2, compared with SEED , our visual de-tokenizer can decode images that are more semantically aligned with the original images by taking the ViT features as inputs.

In the second stage, as shown in Fig. 3 (right), we further fine-tune the visual de-tokenizer to take an extra condition image as inputs for the retention of low-level details. Specifically, we follow InstructPix2Pix to encode the condition image into the latent space via the VAE encoder, and concatenate them with the noisy latent as the input of U-Net. The channel number of the U-Net convolutional layer is expanded from 4 to 8, and all parameters of U-Net are optimized. We fine-tune the visual de-tokenizer on MagicBrush and in-house image editing data, as well as the pure images in the first stage, where the conditional inputs are set to zeros. As shown in Fig. 2, by incorporating the condition image as an additional input besides the high-level image features, our visual de-tokenizer can recover the fine-grained details of the original image.

2 Dynamic Resolution Image Encoding

Current MLLMs require to resize the input images to a pre-defined resolution (typically a square size), which corresponds to the training resolution of the vision encoder, which can result in the loss of fine-grained information. In this work, we propose dynamic resolution image encoding to enable the processing of images with arbitrary sizes and aspect ratios by dividing the image into a grid comprising of sub-images. Specifically, for the visual encoder with the training resolution Ht×WtH_{t}\times W_{t}, we first up-sample the input image with the size H×WH\times W to the size of {Nh∗Ht}×{Nw∗Wt}\{N_{h}*H_{t}\}\times\{N_{w}*W_{t}\}. The grid size Nh×NwN_{h}\times N_{w}, are determined by

We also resize the original image to the size of Ht×WtH_{t}\times W_{t} to provide global visual context. All sub-images and the resized global image are fed into the visual encoder to obtain the features, which are concatenated as the input of the LLM.

To enable the LLM to be aware of the positional information of each sub-image within the original image, we add extrapolatable 2D positional embeddings to the visual features of each sub-image. Specifically, for a sub-image with a normalized center location (xc,yc)(x_{c},y_{c}) in the grid, where 0.0\textlessxc,yc\textless1.00.0\textless x_{c},y_{c}\textless 1.0, its learnable positional embedding pp is computed:

ll, rr, tt, and bb represent four learnable position embeddings indicating left, right, top and bottom respectively. Consequently, our visual encoder can handle inputs with any arbitrary sizes and aspect ratios, even if the image resolution was not encountered during training.

3 Multimodal Pre-training and Instruction Tuning

As shown in Fig. 4, SEED-X adopts next-word prediction and image feature regression training objectives on interleaved visual and textual data. Specifically, we perform dynamic resolution encoding of each image in the multimodal sequence, and their features along with text tokens are fed into the pretrained LLM. In order to equip the model with detection and referencing abilities, we add 224 bbox tokens, designated for representing bounding box coordinates, represented by ‘ ’ with specical tokens at the beginning and end of the bounding box. The text and added bbox tokens are training through predicting the next token with cross-entropy loss.

We employ NN learnable queries (N=64N=64 to align with the visual de-tokenizer) to obtain the output visual representations from the LLM, which are trained to reconstruct the features of the pre-trained ViT tokenizer with a Mean Squared Error (MSE) loss. We add two special tokens ‘’ and ‘’ to represent the beginning and the end of the query embeddings, and the ‘’ is trained to predict where an image emerges. In doing so, we utilize the pre-trained ViT tokenizer as a bridge to decouple the training of a visual de-tokenizer and the MLLM for image generation. During inference, the regressed visual representations from SEED-X are fed into the visual de-tokenizer to decode realistic images.

We pre-train SEED-X initialized from Llama2-chat-13B using LoRA on massive multimodal data, including image-captions pairs, grounded image-texts, interleaved image-text data, OCR data and pure texts. We perform pre-training with 48 H800-80G GPUs (10 days) on a total of 158M samples. See Appendix. A and Appendix. B for more details.

3.2 Training Stage II: Multimodal Instruction Tuning

We perform multimodal instruction tuning through fine-tuning SEED-X using a LoRA module with both public datasets and in-house data covering image editing, text-rich, grounded and referencing QA, and slide generation tasks. The details of datasets can be found in Appendix. A. We fine-tune SEED-X with conversational and image generation data to yield a general instruction-tuned model SEED-X-I, which can follow multimodal instructions and make responses with images, texts and bounding boxes in multi-turn conversation. We further fine-tune the foundation model SEED-X on specialized datasets, resulting in a series of instruction-tuned models tailored for specific tasks, including SEED-X-Edit, SEED-X-PPT, SEED-X-Story and SEED-X-Try-on. The proficient capabilities of these instruction-tuned model across various domains demonstrate the versatility of our pre-trained foundation model SEED-X. We perform instruction tuning on the foundation model SEED-X across different datasets, resulting in various models with distinct capabilities. Note that we do not have an all-in-one instruction-tuned model that encompasses all abilities, which will be explored for future work.

Experiments

We evaluate SEED-X-I on benchmarks specifically designed for evaluating MLLMs, since recent work point out that traditional VQA benchmarks are not tailored for evaluating MLLMs with open-from output. As shown in Tab. 2, SEED-X-I achieves competitive performance in both the image comprehension and generation tasks. For example, SEED-X-I with dynamic resolution inputs achieves results on MMBench that are second only to GPT-4V, while SEED-X-I with fixed image size surpasses the performance of GPT-4V. Since MLLM Benchmarks lack a large number of data with unusual aspect ratios, the advantages of dynamic resolution encoding cannot be fully demonstrated, resulting in SEED-X-I with fixed image size achieving overall better performance. SEED-X-I also shows promising results for comprehending multi-image and interleaved image-text content in SEED-Bench-2 . Compared with previous work that unify comprehension and generation within an LLM, SEED-X-I achieves the state-of-the-art performance in P3 level of SEED-Bench-2 including the evaluation of text-to-image generation, next image prediction and text-image creation.

2 Qualitative Evaluation

SEED-X can be effectively instruction tuned to function as various multimodal AI assistants in the real world across different domains after integrating two enhanced features, including the comprehension of images of arbitrary sizes and ratios, and multi-granularity image generation, encompassing both high-level instructional image generation and low-level image manipulation tasks. As shown in Fig. 1 and Fig. 5, our instruction tuned models can serve as an interactive designer, which can generate images without descriptive captions while illustrate creative intent, and showcase visualizations of modified images. For example, it can explain the design idea of concept image for AGI and a two-story cabin. It can create an imaginative illustration for the novel without the need of describing the scene with languages. It can further offer modification suggestions of the user’s room and showcase the visualization. Additionally, the instruction tuned models can act as an knowledgeable personal assistant, comprehending images of arbitrary sizes and providing relevant suggestions. For example, it can identify foods suitable for fat reduction in the refrigerator, display appropriate clothing based on the screenshot of weather forecasts.

2.2 Image Generation and Manipulation.

We compare previous MLLMs that are capable of generating images for text-to-image generation in Fig. 7 of Appendix. Our instruction tuned model can generate images that are more aligned with the elements in the caption and possess artistic qualities. Through utilizing a pre-trained ViT Tokenizer as the bridge to decouple the training of visual de-tokenizer and the MLLM, our pre-trained model SEED-X can effectively realize high-quality image generation, which is a fundamental capability to be applied in real-world scenarios.

We compare image manipulation with previous MLLMs including Emu2-Gen , Gemini , MGIE and Mini-Gemini . As shown in Fig. 6, we can observe that SEED-X-Edit can more effectively adhere to editing instructions while maintaining the low-level details of the input image. For instance, SEED-X-Edit can accurately add sunglasses to the dog on the right, while both Emu2-Gen and MGIE fail to follow the instruction, resulting in sunglasses being added to both dogs. Additionally, SEED-X-Edit successfully eliminates the dog in the baby image while preserving the background details and the baby’s features. In contrast, Emu2-Gen fails to retain the fine details of the input image, and MGIE is unsuccessful in removing the dog. Note that Gemini lacks the ability to edit images as it retrieves images on the Internet. Here the presence of black images is due to its failure to display images related to human portraits. Mini-Gemini generates text prompts as the input of a pre-trained SDXL model, which can not preserve the visual details of the input image. The examples show the effectiveness of our instruction model for high-precision image manipulation. Our MLLM accurately predicts visual semantic representations based on an input image and a language instruction, which serve as input for the U-Net. The visual de-tokenizer can further condition on the input image, ensuring the preservation of fine-grained details in the decoded images.

2.3 Multimodal Comprehension.

We provide qualitative examples of multimodal comprehension by SEED-X-I in Fig. 8 and Fig. 9 of Appendix. SEED-X-I can realize fine-grained object detection and perception, text-rich comprehension, fundamental mathematical computation, world-knowledge and commonsense reasoning, diagram understanding, which are crucial capabilities for its application in real-world scenarios.

Conclusion

We present SEED-X, a versatile foundation model, which can serve as various multimodal AI assistants in the real world after instruction tuning. In order to make a multimodal foundation model applicable in open-world context, we integrate two enhanced features into SEED-X including image comprehension of arbitrary sizes and ratios, and multi-granularity image generation, which encompasses both high-level instructional image generation and low-level image manipulation. We hope that SEED-X can inspire future research into the potential of MLLMs in the real-world scenarios through unifying multi-granularity comprehension and generation.

References

Appendix A Pre-training and Instruction Tuning Datasets

As listed in Tab. 3, we pre-train SEED-X and conduct instruction tuning on a large variety of both public datasets and in-house data. For multimodal pre-training, we utilize image-caption pairs, grounded image-caption pairs, interleaved image and text content, OCR data and pure text data. The images of LAION-COCO and SAM are re-captioned for a more detailed descriptive caption to improve both image comprehension and generation.

For instruction tuning, we utilize various public VQA datasets, and further curate text-rich QA, grounded and referencing QA to enhance the model’s capability of comprehending text-rich images and detecting objects that requires reasoning. We use multiple conversational datasets, which are specifically collected for MLLMs with open-form text output. We use the same image-caption pairs as in the pre-training phase to maintain the model’s ability to generate images. For the image manipulation, since the high-precision editing dataset MagicBrush is only at the level of thousands, we employ a series of models to collect a dataset of millions of image editing examples, which are used for both training the visual de-tokenizer and SEED-X-Edit. We further collected data on slides, obtaining images, captions, and layouts for training slide generation.

Appendix B Implementation Details

Visual Tokenization and De-tokenization. We use the visual encoder from Qwen-vl as the ViT Tokenizer and adopt 1D average pooling to obtain N=64N=64 visual embeddings. These visual embeddings are fed into four layers of cross-attention as the input of the U-Net initialized from SDXL . In the first stage, we optimize the parameters of the cross-attention layers and the keys and values within the U-Net on the images from JourneyDB , LAION-Aesthetics , Unsplash , and LAION-COCO . We train the visual de-tokenizer on 32 A100-40G GPUs with 42K training steps, where the learning rate is set to 1e-4 with cosine decay.

In the second stage, we encode the condition image into the latent space via the VAE encoder, and concatenate them with the noisy latent as the input of U-Net. The channel number of the U-Net convolutional layer is expanded from 4 to 8, and all parameters of U-Net are optimized. We pre-train the visual conditioner on MagicBrush and in-house image editing data, as well as the image-caption pairs in the first stage, where the conditional inputs are set to zeros. We fine-tune the visual de-tokenizer on 32 A100-40G GPUs with 30K training steps, where the learning rate is set to 1e-4 with cosine decay.

Multimodal Pre-training and Instruction Tuning. We utilize the visual encoder from Qwen-vl as the ViT Tokenizer and initialize a cross-attention layer to obtain N=64N=64 visual embedding as the input of the LLM initialized from Llama2-chat-13B. We initialize N=64N=64 learnable queries and the output hidden states from them are fed into a cross-attention layer to reconstruct N=64N=64 visual embeddings from the ViT Tokenizer. We optimize the LLM using LoRA and optimize the parameters of the input cross-attention layer, output cross-attention layer, extrapolatable 2D positional embeddings, and LoRA on image-captions pairs, grounded image-texts, interleaved image-text data, OCR data and pure texts. We perform pre-training with 48 H800-80G GPUs (10 days) on a total of 158M samples, where the learning rate is set to 1e-4 with cosine decay.

For the instruction tuning, we fine-tune a LoRA module on the pre-trained model, and optimize the parameters of the input cross-attention layer, output cross-attention layer, extrapolatable 2D positional embeddings, and LoRA. We fine-tune SEED-X with conversational and image generation data to yield a general instruction-tuned model SEED-X-I. We further fine-tune SEED-X on specialized datasets, resulting in a series of instruction-tuned models tailored for specific tasks, including SEED-X-Edit, SEED-X-PPT, SEED-X-Story and SEED-X-Try-on.

Appendix C Qualitative Examples

Text-to-image Generation, Fig. 7 visualizes the comparison between MLLMs for text-to-image generation including Next-GPT , SEED-LLaMA-I, Emu2-Gen and Gemini . Compared with previous MLLMs, our instruction tuned model can generate images that are more aligned with the elements in the descriptive caption and possess artistic qualities. For example, images generated by SEED-X-I vividly and accurately depicts “person standing in a small boat”, “a gleaming sword on its back”, “an oriental landscape painting”, “tiger with vivid colors” in the captions. Through utilizing a pre-trained ViT Tokenizer as the bridge to decouple the training of visual de-tokenizer and the MLLM, our pre-trained model SEED-X can effectively realize high-quality image generation, which is a fundamental capability for applying multimodal models in real-world scenarios.

Image Manipulation. We compare image manipulation with previous MLLMs including Emu2-Gen , Gemini , MGIE and Mini-Gemini . Language-guided image manipulation presents a significant challenge as the model must be capable of comprehending free-form instructions and generating images with the low-level details of the input image preserved. As shown in Fig. 6, we can observe that SEED-X-Edit can more effectively adhere to editing instructions while maintaining the low-level details of the input image. For instance, SEED-X-Edit can accurately add sunglasses to the dog on the right, while both Emu2-Gen and MGIE fail to follow the instruction, resulting in sunglasses being added to both dogs. Additionally, SEED-X-Edit successfully eliminates the dog in the baby image while preserving the low-level background details and the baby’s features. In contrast, Emu2-Gen fails to retain the fine details of the input image, and MGIE is unsuccessful in removing the dog. Note that Gemini lacks the ability to edit images as it retrieves images on the Internet. Here the presence of black images is due to its failure to display images related to human portraits. Mini-Gemini generates text prompts as the input of a pre-trained SDXL model, which can not preserve the visual details of the input image. The examples demonstrate the effectiveness of our instruction model for high-precision image manipulation. Our MLLM accurately predicts visual semantic representations based on an input image and a language instruction, which serve as input for the U-Net. The visual de-tokenizer can further condition on the input image, ensuring the preservation of fine-grained details in the decoded images.

Multimodal Comprehension We show qualitative examples of multimodal comprehension by SEED-X-I in Fig. 8 and Fig. 9. SEED-X-I can realize fine-grained object detection and perception, text-rich comprehension, fundamental mathematical computation, world-knowledge and commonsense reasoning, diagram understanding, which are crucial capabilities for its application in the real world.