Making LLaMA SEE and Draw with SEED Tokenizer

Yuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge, Chen Li, Xintao Wang, Ying Shan

Introduction

In recent years, Large Language Models (LLMs) pre-trained on massive text corpus with straightforward training objectives such as next-word prediction have exhibited remarkable abilities to understand, reason, and generate texts across a variety of open-ended tasks. Recent studies further exploit the strong generality of LLMs to improve visual understanding or generation tasks, collectively referred to as Multimodal LLM (MLLM). While these studies have contributed to technological advancements, MLLMs have yet to achieve the remarkable success of LLMs in terms of emergent capabilities. We have made a bold assumption that the premise for the emergence of multimodal capabilities is that text and images can be represented and processed interchangeably in a unified autoregressive Transformer.

We posit that a proper visual tokenizer is the key as it can facilitate the follow-up multimodal training by (i) easing the semantic alignment between visual and word tokens, and (ii) enabling LLM’s original training recipe (i.e., next-word prediction) for multimodal data without specific adaptation for visual tokens. Representing images as a sequence of discrete IDs is naturally compatible with the autoregressive training objective of LLMs. But unfortunately, works that utilize discretized visual tokens for multimodal tasks have receded from prominence, as such models generally rely on super-scale training to converge, leading to substantial training costs. Moreover, our previous work empirically found that the dominant tokenizer VQ-VAE in existing works captures too low-level information for LLMs to effectively perform multimodal comprehension tasks. Existing image tokenizers fail to meet the requirements of unifying the generation of images and texts and facilitating multimodal training.

To this end, we introduce SEED, a VQ-based image tokenizer that produces discrete visual codes with 1D causal dependency and necessary high-level semantics for both visual comprehension and generation tasks, as shown in Fig. 2 (a). The off-the-shelf LLMs can be readily equipped with SEED by treating discrete visual tokens as new words and updating the vocabulary. We would like to emphasize the design principles of SEED. (1) Why causal-dependent tokens? Existing visual tokens (e.g., from VQ-VAE or CLIP-ViT ) are generated using 2D context, which is incompatible with the unidirectional attention in dominant LLMs and counterintuitive for text-to-image tasks requiring raster order prediction. Thus, we convert 2D raster-ordered embeddings into a sequence of semantic codes with 1D causal dependency. (2) Why high-level semantics? Since visual and textual tokens in LLMs are expected to be interoperable—sharing weights and training objectives—they should encompass the same degree of semantics to prevent misalignment, i.e., the high-level semantics inherently present in words.

Specifically, the SEED tokenizer is composed of a ViT encoder , Causal Q-Former, VQ Codebook , multi-layer perceptron (MLP), and a UNet decoder . The ViT encoder and UNet decoder are directly derived from the pre-trained BLIP-2 and unCLIP-SD model , respectively. (1) Tokenize: Causal Q-Former converts 2D raster-ordered features produced by the ViT encoder into a sequence of causal semantic embeddings, which are further discretized by the VQ Codebook. (2) De-Tokenize: The discrete visual codes are decoded into generation embedding via MLP. The generation embedding is aligned with the latent space of unCLIP-SD so that realistic images with consistent semantics can be generated using the off-the-shelf SD-UNet.

We further present SEED-LLaMA by equipping the pre-trained LLM with SEED tokenizer. SEED-LLaMA is pretrained on multimodal data, including image-text pairs, video-text pairs, and interleaved image-text data, toward the training objective of next-word prediction as shown in Fig. 2 (b). Such an easy-to-implement and unified proxy task facilitates scalable multimodal pretraining. We further apply multimodal instruction tuning to align SEED-LLaMA with human instructions through supervised fine-tuning. Our model demonstrates extensive emergent abilities such as multi-turn in-context image and text generation given multimodal instructions as shown in Fig. 1. We also benchmark on a broad range of tasks including image captioning, image/video question answering, and text-to-image generation, receiving competitive performance.

In summary, our contributions are three-fold. (1) We introduce SEED, an advanced image tokenizer, designed based on the insights that visual tokens compatible with LLMs should capture high-level semantics while being generated with 1D causal dependency. The tailored SEED improves the scalability of subsequent multimodal training. (2) We present SEED-LLaMA, composed of a pretrained LLM and SEED tokenizer, through large-scale multimodal pretraining and instruction tuning under the next-word-prediction training objective. It successfully unified multimodal comprehension and generation tasks in one framework. (3) SEED-LLaMA shows competitive results on existing multimodal tasks (e.g., text-to-image, image-to-text) and further demonstrates emergent abilities in multi-turn in-context multimodal understanding, reasoning, and generation.

Related Work

With the impressive success of Large language models (LLMs), recent studies work on Multimodal LLM (MLLM) to improve visual comprehension through utilizing the strong generality of LLMs. Previous work align visual features of pre-trained image encoder with LLMs on image-text datasets. However, these work commonly use the prediction of the next text token as the objective, thus can only output texts.

To empower LLMs with the image generation ability, CogView pre-trains a visual tokenizer by reconstructing image pixels, and fine-tunes GPT with the objective of next-token prediction. GILL learns a mapping between the embeddings of a LLM and a frozen text-to-image generation model. Both work aim to generate images with LLMs, without being explicitly designed for unifying multimodal comprehension and generation.

Our concurrent works both perform multimodal autoregression including the generation of images and texts. CM3Leon utilizes discrete visual codes from a image tokenizer pre-trained on image pixel reconstruction and performs image-to-text and text-to-image autoregression. However, it yields suboptimal performance in visual comprehension tasks (e.g., CIDEr 61.6 vs. ours 126.9 on COCO image captioning) because the image tokenizer captures too low-level information. Emu employs continuous visual representations and is pre-trained on interleaved multimodal sequences through classifying the next text token or regressing the next visual embedding. For image generation, Emu further fine-tunes a SD model to accommodate the output representations from the LLM. By contrast, we pre-train a discrete image tokenizer, where the visual codes can be decoded to realistic images using the off-the-shelf SD model, and perform multimodal autoregressive with a unified next-word-prediction objective, which facilitates scalable multimodal training.

Visual tokenizer aims to represent images as a sequence of discrete tokens. Previous work trains a Vector Quantized Variational AutoEncoders (VQ-VAE) by reconstructing image pixels, which captures only low-level details such as color, texture and edge. Beit v2 trains a visual tokenizer through reconstructing high-level features from the teacher model, but its visual codes from 2D features of a vision transformer are incompatible with the unidirectional attention in dominant LLMs for image generation. By contrast, we present SEED tokenizer, which produces discrete visual codes with 1D causal dependency and high-level semantics.

Method

As shown in Fig. 3, the SEED tokenizer is composed of a ViT encoder , Causal Q-Former, VQ Codebook , multi-layer perceptron (MLP), and a UNet decoder . The ViT encoder and UNet decoder are directly derived from the pre-trained BLIP-2 and unCLIP Stable Diffusion (unCLIP-SD) , respectively. We first train a Causal Q-Former to convert 2D raster-ordered features (16×\times16 tokens) produced by the ViT encoder into a sequence of causal embeddings (32 tokens). We then train a visual codebook to discretize the causal embeddings to quantized visual codes (32 tokens) with causal dependency. We employ a MLP to decode the visual codes into generation embedding (1 token), which is aligned with the latent space of the pre-trained unCLIP-SD conditioned on image embedding. Our previous work aligns generation embeddings with the text embeddings of SD , and we analyze the difference in Sec. 4.3. We pre-train SEED tokenizer on CC3M , Unsplash , LAION-COCO and MS-COCO .

As shown in Fig. 3, a set number of learnable query embeddings (32 tokens) and features of a pre-trained ViT encoder are fed into the Causal Q-former to encode a fixed number of causal embeddings (32 tokens) of the input image. Specifically, the query embeddings can interact with only previous queries through self-attention layers with causal mask, and interact with frozen image features through cross-attention layers. We adopt contrastive learning to optimize Causal Q-former fine-tuned from BLIP-2 Q-Former on image-text pairs. We use contrastive loss to maximize the similarity between the final causal embedding and text features of the corresponding caption.

1.2 Training Stage II: Visual Tokenize and De-tokenize

As shown in Fig. 3, we train a VQ codebook to discretize the causal embeddings (32 tokens) into quantized visual codes (32 tokens). Specifically, a quantizer looks up the nearest neighbor in the codebook for each causal embedding and obtains the corresponding code. We employ a decoder, which is a multi-layer Transformer , to reconstruct the continuous causal embeddings from discrete codes. During training, we maximize the cosine similarity between the output of the decoder and the causal embeddings. We further employ a MLP to reconstruct the image embedding (1 token) of a frozen unCLIP-SD from discrete codes. During training, we minimize the MSE loss between the generation embedding and the image embedding of unCLIP-SD. During inference, the generation embedding are fed into the off-the-shelf SD-UNet to decode realistic images.

2 SEED-LLaMA

As shown in Fig. 4, SEED-LLaMA adopts a unified next-word-prediction training objective on interleaved visual and textual data. Specifically, visual inputs are first discretized into a sequence of causal codes by SEED tokenizer. Then the interleaved visual codes and text tokens are fed into the pretrained LLM for performing multimodal autoregression, where the visual codes are treated as new words and the vocabulary of the LLM is updated accordingly. We maximize the likelihood in a unified autoregressive manner as follows:

where uiu_{i} represents visual code or text token, and Θ\Theta denotes the the parameters of the transformer. We initialize SEED-LLaMA from a pre-trained LLM, and add 8192 visual codes to the vocabulary. The embedding layer and decoder head layer in the transformer are expanded and the parameters of added visual codes are randomly initialized.

For efficiency, we first train SEED-LLaMA using LoRA tuning and together optimize the parameters of the embedding layer and decoder head layer due to the added visual codes. We then merge the parameters of LoRA onto the LLM backbone and fine-tune all parameters except for the embedding layer. We freeze the embedding layer since we observe that fine-tuning it together with other parameters can lead to unstable training loss, which is also reported in BLOOM and GLM-130B . We preprocess the images and videos into discrete tokens beforehand to conserve computational resources. We perform pretraining using two versions of LLM, Vicuna-7B and Llama2-chat-13B, with 64 A100-40G GPUs, and yield SEED-LLaMA-8B (144 hours) and SEED-LLaMA-14B (216 hours), respectively. See Appendix. B for details.

2.2 Training Stage II: Multimodal Instruction Tuning

We perform multimodal instruction tuning on SEED-LLaMA to align it with human instructions through supervised finetuning on public datasets. The details of datasets can be found in Appendix. C. We fine-tune a LoRA module on the pre-trained SEED-LLaMA with the template as below,

Only the content of is accounted for loss. The overall instruction tuning phase takes 16 hours for SEED-LLaMA-8B and 27 hours for SEED-LLaMA-14B with 32 A100-80G GPUs.

Experiment

We evaluate the performance of Causal Q-Former on the image-text retrieval using COCO and Flickr30K . The performance is measured by Recall@K (R@K). Note that we adopt the dual-stream paradigm for inference and remove the image-text-matching (ITM) re-rank module in BLIP-2 for a fair comparison. As shown in Tab. LABEL:tab:retrieval, our Causal Q-former achieves better results than BLIP-2 in terms of an aggregated metric Recall@mean. It demonstrates that the output query embeddings with causal dependency do not drop performance than the output embeddings with bi-directional attention in BLIP-2.

We evaluate causal codes on the image-text retrieval, where the reconstructed embeddings from causal codes are used for retrieval. As shown in Tab. LABEL:tab:retrieval, discrete codes exhibit competitive performance compared to BLIP-2, which demonstrates that the discrete codes from SEED tokenizer capture high-level semantics, which are suitable for visual comprehension.

We further evaluate image reconstruction on COCO and Flickr30K dataset. SEED first discretizes input images into causal codes (32 tokens) and obtain generation embedding (1 token), which are fed into the unCLIP-SD-UNet for reconstruction. We follow GILL to compute the CLIP similarity score as the metric to evaluate the semantic consistency. As shown in Tab. LABEL:tab:clip_score, compared with the upper bound unCLIP-SD, SEED only slightly drops performance.

We visualize the reconstructed images of SEED tokenizer in Fig. 5. Through obtaining the generation embedding from the causal visual codes, realistic images can be generated using the frozen SD-UNet, which maintain consistent semantics with inputs. The above evaluation and visualization demonstrate the versatility of SEED visual tokens for both comprehension and generation tasks.

2 SEED-LLaMA

We evaluate SEED-LLaMA on a wide range of multimodal comprehension tasks including image captioning and image/video question answering. Details of these benchmarks and evaluation metrics are provided in Appendix. D. As shown in Tab. 3, our SEED-LLaMA achieves competitive performance in both the image and video understanding tasks compared with MLLMs that use continuous visual representations. The results demonstrate that our SEED tokenizer can generate discrete visual codes with high-level semantics, which facilities the visual comprehension. We can observe that pretraining from a LLM with larger model size improves performance on SEED-Bench and instruction tuning further contributes to enhanced results. Note that as pointed out by recent work , previous VQA benchmarks listed in Tab. 3 are not tailored for evaluating MLLMs with open-from output, since they require an exact match between the model prediction and the target word or phrase. The qualitative examples of multimodal comprehension is provided in Appendix. E.

We evaluate the text-to-image generation on MS-COCO and Flickr30K and compute the pair-wise CLIP similarity score as the evaluation metric following GILL . As shown in Tab. LABEL:tab:clip_score, images generated by our SEED-LLaMA from textual descriptions show higher similarity with the ground-truth images. The results demonstrate that SEED-LLaMA generates images that are highly correlated with text prompts via a frozen SD-UNet. We show qualitative examples of text-to-image generation in Appendix. E.

2.2 Emergent Ability

Multi-turn In-context Multimodal Generation.

As shown in Fig. 1 and Fig. 6, given multimodal instructions including images and open-form texts from a user, our SEED-LLaMA can respond with synthesized image (e.g., a dog in front of the Golden Gate Bridge), sequentially generated images (e.g., a cartoon cat in different scenes), instruction-followed image (e.g., a closer look-up of a cherry blossom), various forms of texts via creation and real-world knowledge (e.g., a story, a poem and flower identification). The results illustrate the impressive capability of SEED-LLaMA in reasoning and generating long-context multimodal content.

As shown in Fig. 7, our SEED-LLaMA can realize a variety of zero-shot compositional image generation as below,

Stylized Image Generation. SEED-LLaMA can take a text prompt and a style reference image as inputs and produce an output image that adheres to both the style and text prompt.

Image Blending. SEED-LLaMA can take two images as inputs and generate an image that blends the visual components of the input images.

Multimodal Composition. SEED-LLaMA can take an image prompt and a text prompt as inputs and generate a composite image that combines the multimodal inputs.

In-context Generation. SEED-LLaMA can take images, their textual references, and text prompts as inputs and generate context-related images.

3 Ablation Study

The generation embedding of SEED is aligned with the image embedding of unCLIP-SD, and can be decoded to realistic images with the unCLIP-SD-UNet. In our previous work , we train a visual tokenizer SEEDtext\text{SEED}^{\text{text}} through aligning the generation embeddings with the text embeddings (77 tokens) of SD conditioned on texts. As shown in Tab. LABEL:tab:clip_score, the similarity between the reconstructed images of SEEDtext\text{SEED}^{\text{text}} and original images drop heavily. The semantic representations of texts can not fully preserve the rich visual information of images. The visual comparison of the the reconstructed images between SEEDtext\text{SEED}^{\text{text}} and SEED are provided in Appendix. A.

Causal Visual Codes vs. Bilateral Visual Codes.

We train a Causal Q-Former to convert 2D features produced by the ViT encoder into a sequence of causal semantic embeddings, which are further discretized as causal visual codes. To verify whether the causal visual codes are necessary for compatibility with LLM, we train a visual tokenizer SEEDBi\text{SEED}^{\text{Bi}}, which produces bilateral visual codes from a pre-trained Q-Former with bilateral self-attention. We then pre-train SEEDBi\text{SEED}^{\text{Bi}}-LLM∗\text{LLM}^{\ast} and SEED-LLM∗\text{LLM}^{\ast} on image-text pairs and evaluate the text-to-image generation on COCO test set. Given 5000 captions of COCO, SEEDBi\text{SEED}^{\text{Bi}}-LLM only generates 2134 images successfully while SEED-LLM∗\text{LLM}^{\ast} generates 4997 images (Failure cases occur when the model predicts a number of visual tokens not equal to 32). The results demonstrate that the non-causal codes lead to highly unstable model performance since they contradict with the left-to-right autoregressive mechanism of LLM.

We first train SEED-LLaMA using LoRA tuning, and then merge the parameters of LoRA with the original LLM and fine-tune all parameters except for the embedding layer.

To explore whether fully fine-tuning helps, we evaluate the performance of the model before and after fully fine-tuning on image captioning and text-to-image generation, with evaluation metric CIDEr and clip similarity score. Tab. LABEL:tab:ablation shows that fully fine-tuning the LoRA tuned model enhances model’s capability for both image comprehension and generation.

Conclusion

We present SEED, a discrete image tokenizer, designed based on the premise that visual tokens compatible with LLMs should capture high-level semantics while being generated with 1D causal dependency. SEED enables LLMs to be trained with multimodal data following the original recipe of text (i.e., next-word prediction), which is mature and scalable. We further present SEED-LLaMA via multimodal pretraining and instruction tuning on the interleaved visual and textual data with SEED tokenizer. SEED-LLaMA not only exhibits remarkable performance across multimodal comprehension and image generation tasks, but also demonstrates extensive compositional emergent abilities. We hope that SEED would draw increased attention to visual tokenizers. A more rational visual tokenizer could substantially reduce the complexity of multimodal LLM training.

References

Appendix A SEED Tokenizer

The generation embedding of SEED is aligned with the image embedding of unCLIP SD, and can be decoded to realistic images with the unCLIP-SD-UNet. In our previous work , we train a visual tokenizer SEEDtext\text{SEED}^{\text{text}} through aligning the generation embeddings with the text embeddings (77 tokens) of SD , and the generation embeddings can be decoded to images with the SD-UNet. The visual comparison of the the reconstructed images between SEEDtext\text{SEED}^{\text{text}} and SEED are shown in Fig. 8. We can observe that compared with SEEDtext\text{SEED}^{\text{text}}, the images reconstructed by SEED can better preserve the visual information of the original images.

Appendix B Pretraining

As shown in Tab. 5, we utilize diverse categories of datasets as pretraining data, which can be summarized as follows.

We use the image-text pairs from CC3M , Unsplash , LAION-COCO and MS-COCO . We filtered the samples in these datasets based on image resolution, aspect ratio, and visual-textual similarity. We randomly place images or text at the forefront, in order to achieve the generation of captions based on images and vice versa.

We use a large-scale dataset WebVid-10M containing videos and captions. We implemented heuristic rules to exclude extraneous metadata, such as the resolution of the original video and camera parameters. We sample four frames of each video for training.

We use publicly available MMC4 and OBELISC datasets, which were extracted and thoroughly filtered from Common Crawl. Specifically, we employ the MMC4-core split, consisting of 7.3 million samples, and the complete OBELISC dataset, containing 141 million samples. For documents in MMC4, we create a sequence of length 1024 and randomly shuffle the order of images and their corresponding texts (those with the highest CLIP score). As for OBELISC, we generate a sequence of length 1024 based on the order of data in the dataset.

B.2 Pretraining Hyperparameters

We report the detailed pretraining hyperparameters of SEED-LLaMA in Tab. 6.

Appendix C Instruction Tuning

We summarize the datasets and their prompts for supervised instruction tuning of SEED-LLaMA in Tab. 7 and Tab. 8. Note that MagicBrush contains both the single-turn and multi-turn scenarios, and we only use the single-turn for multimodal prompt image generation.

Appendix D Evaluation

In order to assess the multimodal comprehension and image generation ability of SEED-LLaMA, we evaluate SEED-LLaMA on 10 benchmarks as shown in Tab. 9. For the evaluation of image generation, we adopt the CLIP-ViT-L/14 to calculate the CLIP score between the ground-truth image and the generated image. When evaluating SEED-Bench, we adhere to the official guidelines, selecting the option with the highest log likelihood as the response for each multi-choice question (MCQ). For the evaluation on video tasks, we uniformly sample 4 frames for MSVDQA and MSRVTTQA, and 8 frames for NExTQA. For the other tasks, we follow the evaluation procedures in prior works and either submit the results to the official server (VQAv2, VizWiz) or assess them using the official code ourselves.

D.2 Prompt Templates

We summarize the prompt templates used for evaluating SEED-LLaMA in Tab. 10. As the pre-trained SEED-LLaMA with size of 8B and 14B adopt different LLM (Vicuna-7B and Llama2-chat-13B), their prompts differ accordingly.

Appendix E Qualitative Cases

More examples of multi-turn in-context multimodal generation and compositional image generation are shown in Fig. 9 and Fig. 10. Note that generating images with multimodal prompt is not an emergent ability since SEED-LLaMA is fine-tuned on corresponding paired data such as InstructPix2Pix . We showcase qualitative examples of text-to-image generation by SEED-LLaMA in Fig. 11. Given various textual descriptions, our SEED-LLaMA can generate realistic images that aligns with the text prompts. We further provide qualitative examples of multimodal comprehension by SEED-LLaMA in Fig. 12, Fig. 13 and Fig. 14. SEED-LLaMA can realize in-context multi-image understanding, real-world knowledge grounding, complex reasoning, story creation and video understanding.