DreamLLM: Synergistic Multimodal Comprehension and Creation

Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, Xiangwen Kong, Xiangyu Zhang, Kaisheng Ma, Li Yi

Introduction

“What I cannot create, I do not understand.”

Richard P. Feynman, on his blackboard at the time of his death, 1988

Content comprehension and creation in multimodality are crucial and among the ultimate courses of machine intelligence (Sternberg, 1985; Legg & Hutter, 2007). To this end, Multimodal Large Language Models (MLLMs) (Alayrac et al., 2022; Hao et al., 2022; Huang et al., 2023) have emerged as extensions of the successful GPT-style Large Language Models (LLMs) (Brown et al., 2020; Zhang et al., 2022; OpenAI, 2022; 2023; Chen et al., 2023b; Touvron et al., 2023a; b) into visual realm. Recognized as foundation models (Bommasani et al., 2021), MLLMs have achieved unprecedented progress in multimodal comprehension capabilities. These advanced models typically enhance LLMs by incorporating images as multimodal inputs, such as CLIP features (Radford et al., 2021), to facilitate language-output multimodal comprehension. Their aim is to capture multimodal conditional or marginal distributions via a language posterior. However, multimodal creation, which involves generating images, texts, or both, necessitates a universal generative model that simultaneously learns language and image posteriors—currently underexplored.

Until very recently, some concurrent works have shown success in conditional image generation using MLLMs (Koh et al., 2023; Sun et al., 2023b). As depicted in Fig. 1, these methods compel MLLMs to produce either discrete or continuous conditional embeddings that explicitly align with a pretrained CLIP encoder, which could later be used by a pretrained Stable Diffusion (SD) (Rombach et al., 2022) model for image generation. However, due to an inherent modality gap (Liang et al., 2022), CLIP semantics focus predominantly on modality-shared information, often overlooking modality-specific knowledge that could enhance multimodal comprehension. Consequently, these studies have not fully realized the potential learning synergy between multimodal creation and comprehension, have shown only marginal improvements in creativity, and remain deficient in multimodal comprehension.

In this work, we introduce DreamLLM, universally learning image and text posteriors with expected creation & comprehension synergy, based on the following two de-facto designing principles:

Generate Everything as It Is Different from existing works that generate intermediate image representations like CLIP embeddings during training, DreamLLM not only takes all modalities raw data as inputs but also as outputs in a truly end-to-end fashion (i.e., outputs are identical to inputs, see Fig. 1). The challenge lies in enabling MLLMs to learn the image posterior without compromising their comprehension capabilities. To address this, we introduce dream queries, a set of learnable embeddings that encapsulate the semantics encoded by MLLMs. This approach avoids altering the output space of MLLMs. Raw images are then decoded by the SD image decoder conditioned on these semantics. In this fashion, the pretrained SD acts as the score function (Ho et al., 2020). The image posterior is thus modeled by direct sampling in the pixel space, facilitated by score distillation (van den Oord et al., 2018; Poole et al., 2023).

Interleaved Generative Pre-Training (I\mathcal{I}-GPT) DreamLLM is trained to generate interleaved multimodal corpora from the internet (Zhu et al., 2023b), both encoding and decoding interleaved image-text multimodal inputs. Unlike encoding multimodal inputs as in existing methods, decoding interleaved multimodal outputs is challenging due to the complex interleaving layout structures and the long-context requirement of image. Our approach tackles the interleaved layout learning using a unique token that predicts the placement of images within texts. Harnessing DreamLLM’s causal nature, all contents are generated with history multimodal contexts of any length. This interleaved generative pretraining (I\mathcal{I}-GPT) inherently forms all joint, marginal, and conditional distributions of images and texts in the document, leading to a learning synergy that grounds DreamLLM’s comprehension in creation and vice versa.

Extensive experiments across various vision-language comprehension, content creation, and language-only tasks demonstrate DreamLLM’s superior performance as a zero-shot multimodal generalist. For instance, DreamLLM-7B achieves an 8.46 FID on MS-COCO and sets a new standard with 49.1/35.9 scores on MMBench and MM-Vet evaluations, respectively. Moreover, we delve into the learning synergy between comprehension and creation, revealing decent in-context generation capabilities. With I\mathcal{I}-GPT pretraining, DreamLLM generates interleaved documents following human prompts after supervised fine-tuning on instruction-following data, curated with GPT-4. To our knowledge, this work is the first to enable MLLMs to create free-form interleaved content with a learning synergy on both sides. As a foundational learning framework, DreamLLM is adaptable across all modalities, laying a promising foundation for future multimodal learning research.

Background & Problem Statement

Since ϵξ(zt;t)=−σtsξ(zt;t)\bm{\epsilon}_{\xi}(\mathbf{z}_{t};t)=-\sigma_{t}s_{\xi}(\mathbf{z}_{t};t) as derived from Tweedie’s (Efron, 2011; Luo, 2022), it is equivalent to denoising score matching of \gradientztlog⁡pξ(zt)\gradient_{{\mathbf{z}}_{t}}\log p_{\xi}({\mathbf{z}}_{t}) (Hyvärinen, 2005; Vincent, 2011), thus DMs are also called score-function based generative models (Song & Ermon, 2019; 2020; Song et al., 2021; 2023).

Multimodal signals typically exhibit modality-specific information that has distinct structure but complementary semantics (Dong et al., 2023). This complementary property allows us to utilize deep language comprehension to enhance cross-modal image generation (Saharia et al., 2022). However, the potential of multimodal creation to improve comprehension remains largely unexplored.

Existing strategies (Koh et al., 2023; Sun et al., 2023b; Ge et al., 2023) integrate successful Diffusion Models with MLLMs by aligning the semantic spaces of conditional embeddings between CLIP CCLIP\mathcal{C}^{\text{CLIP}} and MLLMs CMLLM\mathcal{C}^{\text{MLLM}}. The objective is to minimize alignment loss Lalign=D(Mψ∘CMLLM,CCLIP)\mathcal{L}_{\text{align}}=D(\mathcal{M}_{\psi}\circ\mathcal{C}^{\text{MLLM}},\mathcal{C}^{\text{CLIP}}), employing a distance metric D(⋅,⋅)D(\cdot,\cdot) and a condition projector Mψ\mathcal{M}_{\psi}. However, CLIP models primarily learn modality-shared semantics, often overlooking modality-specific information due to a modality gap (Liang et al., 2022; Liu et al., 2023d). This explicit alignment with CLIP’s intermediate output space may induce more conflicts than synergies, as MLLMs are forced to generate semantically reduced information, deviating from their original output space. To circumvent these issues, we propose alternative learning methodologies (See Fig. 2), which we elaborate in the ensuing sections.

Learning Objective Our aim is to leverage MLLMs to model distributions via direct pixel space sampling. Here, the pretrained SD functions as a score metric, distilling the learned data distribution. This approach is similar to Score Distillation Sampling (Poole et al., 2023) (SDS, also known as Score Jacobian Chaining (Wang et al., 2023a)). In this context, image posterior is learned in a DeepDream-like manner (Mordvintsev et al., 2015), using MLLMs’ conditional parameterization.

Conditional Embeddings Rather than converting the output space of MLLMs to align with CLIP, we propose to query MLLMs using learned embeddings. Consequently, MLLMs-enriched semantics serve as diffusion conditioning, and the distribution is implicitly modeled through synthesis sampling.

DreamLLM

We introduce DreamLLM, a universal learning framework that facilitates both MLLM’s comprehension and creation capabilities. Our DreamLLM is built with a causal decoder-only LLM Fθ\mathcal{F}_{\theta} as the model foundation, i.e., Vicuna (Chiang et al., 2023) based on LLaMA (Touvron et al., 2023a) trained on ShareGPT (Zheng et al., 2023). We adopt OpenAI’s CLIP-Large (Radford et al., 2021) as the visual encoder Hϕ\mathcal{H}_{\phi}, followed by a linear layer Mζ\mathcal{M}_{\zeta} for visual embedding projection. To synthesize images, we use Stable Diffusion (SD) (Rombach et al., 2022) as the image decoder, and the condition projector Mψ\mathcal{M}_{\psi} is also a linear layer. An overview of the architecture is depicted in Fig. 2.

All natural documents can be regarded as carriers of text-image interleaved information. Text-only, images-only, and text-image pairs data, on the other hand, can be seen as special cases of interleaved corpora with different modality compositions. Thus, it is critical to empower the model with the capability to learn and generate free-form interleaved documents that form all possible distributions.

Interleaved Structure Learning To model the interleaved structure, the interleaved sequence is operated by extending a new special token before images. During training, DreamLLM is trained to predict this token that indicates where an image emerges, and the conditional image synthesis is performed afterward, as introduced next. During inference, DreamLLM will generate an image on its “free will” when this token is predicted.

Conditional Synthesis through Score Distillation To avoid the possible conflicts of CLIP semantics and MLLMs stated in Sec. 2.1, we carefully design a different learning objective and conditional embeddings. Formally, we introduce a series of learnable dream queries with length QQ: d={dq}q=1Q\mathbf{d}=\{\mathbf{d}_{q}\}_{q=1}^{Q}. Considering the tt-th token is predicted as token, the conditional embeddings CK(t)+1DreamLLM\mathcal{C}^{\text{{DreamLLM}}}_{K(t)+1} for the (K(t)+1)(K(t)+1)-th image synthesis can be obtained by causally querying the previous sequences:

Thus, the denoising score matching with latent z\mathbf{z} is motivated in the similar formulation to Eq. 2:

where ξ\xi is not updated since the SD is frozen. Eq. 4 can also be viewed as a generalized formulation of textual inversion (Gal et al., 2023), but all condition embeddings are learnable by model-seeking. From the perspective of score distillation (van den Oord et al., 2018), the KL divergence defined by conditions and the pre-learned score function is equivalently minimized for distilling (Hinton et al., 2015) learned probability density in conditional image synthesis:

Universal Multimodal Generative Modeling An interleaved document sequence x={xt}t=1T\mathbf{x}=\{{\mathbf{x}}_{t}\}_{t=1}^{T} contains both words w={wi}i=1N\mathbf{w}=\{\mathbf{w}_{i}\}_{i=1}^{N} and images I={Ik}k=1K\bm{I}=\{I_{k}\}_{k=1}^{K}. The autoregressive nature forms all possible conditional distributions, such as image conditional multimodal comprehension p(w∣I)p(\mathbf{w}|\bm{I}) or text-to-image synthesis p(I∣w)p(\bm{I}|\mathbf{w}). The images are processed as visual embeddings V\bm{V} for causal comprehension. Assuming that the pretrained SD is an optimal score function, Eq. 5 thus could be viewed as an MLE optimization for the synthesis posterior. Different from Eq. 1, the targeted sequence xt{\mathbf{x}}_{t} now could be both encoded images or words. The objective is thus unified to the MLE of all causally-conditioned posteriors in arbitrary forms:

2 Model Training

In this work, we consider a three-stage training procedure. It can be summarized as follows, and the implementation details, like training data, can be found in Table 9 in Appendix C.

Alignment Training This stage is used to alleviate the gap in multimodality, facilitating the adaptation of multimodal inputs to LLMs. The linear visual projector, linear condition projector, and learnable dream embeddings are pretrained for cross-modal manifold alignment among frozen LLMs, visual encoder, and SD. We use approximately 30M image-text pairs data, training both image-to-text comprehension and text-to-image synthesis.

I\mathcal{I}-GPT Pretraining Following alignment, the LLM undergoes an unfrozen process for I\mathcal{I}-GPT pretraining (detailed in Sec. 3.1). This critical stage facilitates the learning of joint vision-language distributions via generative modeling. Training incorporates approximately 2M selectively filtered documents from MMC4-Core (Zhu et al., 2023b), adhering to a CLIP score threshold of 0.25. Furthermore, we use 2M paired data samples from LAION400M (Schuhmann et al., 2021), captioned by BLIP (Li et al., 2022) (i.e., BLIP-LAION), to enhance text-to-image training and potentially mitigate the impact of some low-quality noisy images and texts from sMMC4.

Supervised Fine-tuning This stage enables the model to perform general multimodal comprehension and creative tasks following human instructions (Ouyang et al., 2022). We utilize approximately 80K visual instruction tuning data collected by Liu et al.. For instruction-following content creation, GPT-4 (OpenAI, 2023) is prompted with document summaries or image captions, collecting approximately 20K instruction-following document synthesis from MMC4 (InstructMMC4) and 20K image synthesis data from BLIP captioned LAION400M (Instruct-BLIP-LAION).

Experiments

DreamLLM is a versatile multimodal generalist that excels at zero-shot or in-context vision-language comprehension and synthesis tasks. In this section, we conduct systematic evaluations for demonstration. See qualitative results in Appendix B and implementation details in Appendix C.

Multimodal comprehension enables humans to interact with agents conditioned on both words and visual content. We evaluate the multimodal vision and language capabilities of DreamLLM across several benchmarks, including image-to-text captioning on COCO (Karpathy & Fei-Fei, 2017) and Image2Paragraph (Krause et al., 2017), general visual question answering (VQA) on VQAv2 (Goyal et al., 2019), OKVQA (Marino et al., 2019), VizWiz (Gurari et al., 2018), and text-related VQA on TextVQA (Singh et al., 2019). Additionally, we conducted a zero-shot evaluation on the recently developed benchmarks of MMBench and MM-Vet to assess the model’s performance in complex multimodal tasks. The results are presented in Table 1 (See Table 5, and Table 6 in Appendix A). All metrics and data splits are listed in Table 10 in Appendix C. We find that i) DreamLLM outperforms other MLLMs across all benchmarks. Notably, DreamLLM-7B surpasses concurrent MLLMs with image synthesis capabilities by a significant margin, achieving +16.6 higher accuracy on VQAv2 compared to Emu-13B. ii) On comprehensive benchmarks like MMBench and MM-Vet, DreamLLM achieves state-of-the-art performance against all 7B counterparts. Detailed analysis revealed superior spatial/relation reasoning capabilities in DreamLLM compared to other MLLMs, likely a result of its image synthesis learning. See qualitative results and comparisons on multimodal dialogue in Appendix B, Appendix B, Fig. 7, Fig. 8, and Fig. 9, in Appendix B.

2 Text-Conditional Image Synthesis

Text-conditional image synthesis is one of the most commonly used techniques for creative content generation that follows human’s fabulous imaginations through free-form languages.

We assess text-conditional image synthesis on the MS-COCO validation set (Lin et al., 2014) and LN-COCO, the COCO subset of Localized Narratives (Pont-Tuset et al., 2020), following prior works (Xu et al., 2018; Yu et al., 2022b). The MS-COCO dataset primarily contains high-level image abstractions with shorter captions, whereas LN-COCO provides more comprehensive image descriptions (Yu et al., 2022b). DreamLLM samples 8 images per text prompt on MS-COCO by CLIP score ranking, following previous works (Ramesh et al., 2022). On LN-COCO, DreamLLM samples one image per prompt without CLIP ranking since the text is too long and exceeds the CLIP length limit. Note that Parti samples 16 images per prompt with CoCa (Yu et al., 2022a). Our evaluation metric is the zero-shot Fréchet Inception Distance (FID) (Heusel et al., 2017), the results of which are presented in Table 2. We note three key observations: i) Our DreamLLM shows a significant FID improvement over the StableDiffusion baseline after stage-I alignment, reducing the score by 3.67 and 11.83 on MS-COCO and LN-COCO respectively. Further, FID improvements of 3.97 and 13.73 are achieved after pretraining and supervised fine-tuning. The substantial improvement on LN-COCO underscores DreamLLM’s superior capability in processing long-context information. ii) When compared to prior specialist models, DreamLLM delivers competitive results based on the SD image decoder. iii) DreamLLM consistently outperforms concurrent MLLMs-based image synthesis methods. For instance, DreamLLM-7B surpasses Emu-13B by a significant 3.20 FID on MS-COCO. See qualitative results on text-to-image synthesis in Fig. 10 and Fig. 11 in Appendix B.

3 Multimodal Joint Creation & Comprehension

Free-form Interleaved Document Creation Instruction tuning endows DreamLLM to act as a multimodal generalist that performs various kinds of tasks by following instructions. Leveraging the interleaved generative modeling from I\mathcal{I}-GPT, DreamLLM can now generate interleaved documents in a free-form manner. In Fig. 3, we showcase the generated interleaved contents based on human instructions. It demonstrates that: i) DreamLLM can generate meaningful responses in accordance with the given instructions. ii) The system can autonomously create images at any specified location by predicting the proposed tokens, thereby eliminating the need for additional human intervention. This is a more user-friendly approach compared to systems like Emu, which necessitate human input for image generation locations. iii) The images generated by DreamLLM accurately correspond to the associated text, a vital attribute for interleaved documents.

Image Quality Document quality can be influenced by factors such as text content, image quality (including image-text alignment), and illustration positioning. To assess the quality of generated documents, we utilized a held-out instruction-following subset from the constructed InstrcutMMC4 as a demonstrative tool. This subset comprises 15K documents across 30 MMC4-defined topics, with 500 samples per topic. We began by evaluating image quality using FID on this subset, generating each image based on the corresponding ground truth texts. The results revealed that when using only matched text inputs for image synthesis, SD achieved an FID score of 74.77. In contrast, our DreamLLM significantly outperforms SD with an FID score of 36.62.

Human Evaluation We perform a comprehensive human evaluation to assess the quality of the generated samples. We randomly selected 150 samples (5 per topic) for instruction-following document generation, mixing the generated and ground truth MMC4 documents without any identifying information. Five unbiased volunteers were then asked to determine whether the given samples were supported. Given the presence of duplicate and low-quality images in MMC4, the supportive rate for MMC4 was only 77.24%. In contrast, our DreamLLM model achieves a supportive rate of 60.68%, surpassing the 30% Turing test requirement. This result indicates that the generated documents contain high-quality images placed logically, demonstrating the effectiveness of our model.

Discussions

To elucidate the synergy between multimodal creation and comprehension, we make the comparison among three methods with DreamLLM architecture, each utilizing identical training data yet differing in their learning objectives: a) the Creation-only baseline, focused solely on text/document-conditional image synthesis; b) the Comprehension-only baseline, dedicated to word generation exclusively; c) the Joint-learning method, which is the default setting of DreamLLM learning both image and language modeling.

Quantitative Analysis As per Table 3, the following observations are made: i) The powerful language comprehension of LLMs significantly enhances the performance of text-to-image specialists like SD, as evidenced by the impressive 8.50 FID (line 1). ii) The use of interleaved data, such as MMC4, can potentially boost multimodal comprehension performance (line 4). iii) The proposed I\mathcal{I}-GPT, further synergizes comprehension and creation with improved performance (line 5). iv) When incorporating CLIP alignment loss LCLIP\mathcal{L}_{\text{CLIP}} stated in Section 2.1, our DreamLLM fails to converge but rather ends in a collapsing point (line 6). This indicates that the queries are adaptively learning the true data distributions, where CLIP semantics are in conflict with MLLM-encoded semantics.

Qualitative Analysis In Fig. 4, we compare answers to some examplar VQA tasks from comprehension-only and joint learning modules, respectively. It can be seen that: i) The joint-learning method exhibits superior multimodal comprehension, particularly in identifying subject relationships and attributes like object size. ii) In multimodal comprehension scenarios involving multiple image inputs, the joint-learning approach demonstrates enhanced precision. This improved performance is a natural outcome of I\mathcal{I}-GPT pretraining, allowing better modeling of multimodal correlations in various interleaved documents.

Multimodal In-Context Generation Multimodal in-context generation is a critical emerging capability for MLLMs (Bommasani et al., 2021; Alayrac et al., 2022). While significant strides have been made in in-context visual question answering, in-context image synthesis remains relatively lacking in exploration. The multimodal context-conditional image synthesis capabilities of DreamLLM, as demonstrated in Fig. 5, offer promising insights into this domain. Tasks such as in-context image edition, subject-driven image generation, and compositional generation, however, pose significant challenges in a zero-shot setting, particularly without downstream fine-tuning as in DreamBooth (Ruiz et al., 2023) or attention modification techniques as in Prompt2Prompt (Hertz et al., 2023). Despite these hurdles, Fig. 5 illustrates DreamLLM’s ability to generate images conditioned on the provided image context. This capability suggests promising potential for DreamLLM in maintaining subject, identity, and semantic context, thereby paving a new way for resolving these complex tasks.

2 What is learned by DreamLLM?

In DreamLLM, the conditional embedding is derived from MLLMs with some learned dream queries. Fig. 6 demonstrates a visualization of the learned cross-attention mechanism between these queries and the diffusion latent. Similar to (Hertz et al., 2023), we visualize the attention map averaged across all timestamps. It is seen that: i) The query attention is structured, disentangled, and semantically-oriented. This is evidenced by the fact that distinct queries adeptly capture different subject and background semantics. ii) Despite varying prompts, attention patterns exhibit remarkable similarity as shown in Fig. 6 (a) and (b). This contrasts with the token attentions from the original SD, which are typically text-token dependent. We postulate that this arises from the model’s causal nature, leading to a consistent semantic structure order.

Related Works

Rapid developments have been witnessed in extending LLMs like LLaMA (Touvron et al., 2023a) to multimodal comprehension that enables human interaction with both words and visual content. One line of work is built by system integration of LLMs with various functioning agents where language acts as general interface (Wu et al., 2023; Gupta & Kembhavi, 2023; Yang et al., 2023b; Liang et al., 2023; Shen et al., 2023; Yang et al., 2023a; Surís et al., 2023), and remarkable success has been demonstrated in such plugin-style frameworks. Another line of work instead explores training LLMs to consume and understand multimodal inputs (Hao et al., 2022; Huang et al., 2023; Chen et al., 2023b) with parameter-efficient tuning (Hu et al., 2022; Alayrac et al., 2022; Li et al., 2023b; Zhang et al., 2023d; Zhu et al., 2023a; Ye et al., 2023) and instruction tuning (Xu et al., 2023; Liu et al., 2023a; Dai et al., 2023). More recently, some approaches have been developed towards visual-interactive multimodal comprehension by precise referring instruction tuning (Zhao et al., 2023a; Peng et al., 2023; Chen et al., 2023a; Zhang et al., 2023f). For cross-modal creation, early works generally tokenize the visual contents into discrete VQ codebooks (van den Oord et al., 2017; Wang et al., 2022; Lu et al., 2023; Diao et al., 2023; Yu et al., 2023a). Recent works instead explore incorporating MLLMs for image synthesis using text-to-image models such as Stable Diffusion, and the objective is to generate conditional embeddings that align pretrained CLIP text (i.e., CLIP) or CLIP variant embeddings (Koh et al., 2023; Ge et al., 2023; Sun et al., 2023a; b).

Conclusions

How can the learning synergy between multimodal content understanding and creation emerge? In this paper, we present DreamLLM, a comprehensive framework for developing MLLMs that not only understands but also creates multimodal content via diffusion models. Through score distillation of conditional-image synthesis distributions, we avoid the need for intermediate representation targets. The employment of interleaved documents further enriches the multimodal distributions, fostering the learning of multimodal encoding and decoding. Our extensive empirical evaluations across diverse VL benchmarks demonstrate the effectiveness of DreamLLM and the emerging learning synergy between multimodal content understanding and creation. Besides, this work initiates the first step towards interleaved content creation. As a general learning framework, we hope it will spur further research in the multimodal machine learning field.

References

Appendix A Additional Experiments

We evaluate the natural language processing capabilities of DreamLLM post-multimodal adaptation learning via zero-shot experiments on language-only tasks. These included commonsense reasoning (PIQA (Bisk et al., 2020), SIQA (Sap et al., 2019), HellaSwag (Zellers et al., 2019), WinoGrande (Sakaguchi et al., 2021)), reading comprehension (BoolQ (Clark et al., 2019)), and a general multi-task benchmark (MMLU 5-shot (Hendrycks et al., 2021)). As Table 4 illustrates, DreamLLM outperforms the Vicuna baseline on most language benchmarks. This suggests that DreamLLM’s multimodal adaptation does not compromise the language learning model’s (LLM) capabilities. When compared to prior Multimodal Language Learning Models (MLLMs), DreamLLM demonstrates superior performance, although this may be attributed to the higher baseline results. This finding suggests that a more robust LLM base model could yield improved results.

A.2 Additional Multimodal Comprehension Results

The evaluation results on MMBench (Liu et al., 2023c) and MM-Vet (Yu et al., 2023b) are presented in Table 5 and Table 6, respectively. The key observations from these results are as follows: i) Our DreamLLM-7B outperforms all other 7B MLLMs, setting a new benchmark in overall performance. Notably, it even exceeds the performance of some 13B models, including LLaVA and MiniGPT-4. ii) A detailed capability evaluation reveals DreamLLM’s superior performance in fine-grained understanding and relational/spatial comprehension. This advantage is likely due to DreamLLM’s unique learning synergy, where image distributions are comprehended not solely through language-posterior comprehension, but also through creation.

Visual Hallucination

Visual hallucination, a phenomenon where Multimodal Large Language Models (MLLMs) generate non-existent objects or identities in images, significantly compromises their multimodal comprehension capabilities and may pose safety risks (MacLeod et al., 2017; Rohrbach et al., 2018). We assess the robustness of DreamLLM against visual hallucination using the recently developed POPE benchmark (Li et al., 2023d). Refer to Table 7 for a detailed comparison with concurrent comprehension-only MLLMs. Our results indicate that DreamLLM-7B exhibits robustness to visual hallucination, matching or surpassing the performance of 13B counterparts. Remarkably, DreamLLM achieves most best or second-best performance in the most challenging setting. We posit that this robust anti-hallucination property stems from a deep understanding of object concepts and semantics, fostered by multimodal creation learning.

A.3 Additional Ablation Study

In Table 8a, we show the results of DreamLLM using different numbers of the proposed learnable queries. i.e., queries. The results show that 64 queries achieve the best result, while 128 may be too many that may impact the performance. However, the choice of query number is also related to the training data size and diffusion model choice. For example, if given more data and a stronger diffusion model image decoder, queries more than 64 may be better.

A.4 Inference Latency

In Table 8b, we present a comparison of real-time inference latency between DreamLLM and SD. Relative to SD, DreamLLM introduces a marginal latency cost of 0.2s on average. This is because the latency primarily stems from the computational demands of the diffusion U-Net denoising, rather than the text condition embedding. To enhance inference efficiency, potential strategies could include the adoption of Consistency Models (Song et al., 2023), or the implementation of model compression techniques such as quantization (Yao et al., 2022; Dong et al., 2022; Shang et al., 2023).

Appendix B Additional Qualitative Examples

In Fig. 10 and Fig. 11, we show the image examples of DreamLLM using the same prompts from previous works for a cross reference and comparison, including DALL-E (Ramesh et al., 2021), DALL-E 2 (i.e., unCLIP) (Ramesh et al., 2022), GLIDE (Nichol et al., 2022), Imagen (Saharia et al., 2022), and Parti (Yu et al., 2022b). Similar to Parti, we have extended some prompts with new sub-prompts for constructing more examples from different prompts.

Multimodal Dialogue

In Tables B and B, we present a comparative analysis of visual question answering results between our model, DreamLLM, and other state-of-the-art models: GPT-4 (OpenAI, 2023), LLaVA (Liu et al., 2023a), BLIP-2 (Li et al., 2022), and OpenFlamingo (Awadalla et al., 2023b). The key findings are as follows: i) DreamLLM surpasses GPT-4 in providing more detailed and precise responses to given questions. ii) While LLaVA (Liu et al., 2023a) also offers detailed responses, it frequently introduces imaginary elements not present in the image. In contrast, DreamLLM delivers more accurate answers, effectively avoiding this visual hallucination issue. This observation aligns with our earlier findings in Table 7, which underscore the robustness of DreamLLM against visual hallucination. Furthermore, we showcase additional qualitative results of the multimodal dialogue in Fig. 7, Fig. 8, and Fig. 9. These figures illustrate DreamLLM’s proficiency in comprehending and generating long-context multimodal information in various input and output formats.

Appendix C Implementation Details

In Table 9, we list the detailed training dataset usage and hyper-parameters. The training data are constructed based on the following datasets: a) LAION400M (Schuhmann et al., 2021), b) LAION-COCO (Schuhmann et al., 2023), c) MMC4 (Zhu et al., 2023b), d) BLIP-LAION (Li et al., 2022) which is filtered and caption by BLIP (Li et al., 2022), e) LLaVAPretrain (Liu et al., 2023a) which contains 558K image-text pairs from BLIP-captioned CC3M (Sharma et al., 2018), SBU (Ordonez et al., 2011), and LAION400M filtered by LLaVA, f) LLaVAInstruct (Liu et al., 2023a), which contains 80K visual instruction-following data constructed by LLaVA, and g) InstructMMC4, which is our instruction-following interleaved document generation data curated by prompting GPT-4 to generate instruction based on the text contents of MMC4. h) Instruct-BLIP-LAION, which is our instruction-following image synthesis data. Similar to InstructMMC4, it is curated by prompting GPT-4 to generate instructions based on image captions. Unless otherwise specified, we randomly sample the indicated number of instances from each dataset during the training process.

C.2 DreamLLM Model

Language Model We use LLaMA-1 (Touvron et al., 2023a) trained on ShareGPT (Zheng et al., 2023) as as the default LLM (i.e., Vicuna-7BVicuna-7B v1.1: https://huggingface.co/lmsys/vicuna-7b-v1.1. (Chiang et al., 2023)) following Liu et al. (2023a) to endow its instruction-following capacity. During training, we use Flash Attention (Dao et al., 2022) and PyTorch FSDP (Zhao et al., 2023b) to accelerate training efficiency.

Visual Encoder The visual encoder is the publicly available OpenAI CLIP-L/14 (Radford et al., 2021) model, which is frozen during the whole process. The images are resized to 224×\times224 resolution to align with the CLIP pretraining settings, resulting in a sequence of 256 total tokens for each image. Following prior VL practice (Lu et al., 2019; Liu et al., 2023a), we append a special token before the image sequence and a special at the end of the sequence.

Diffusion Image Decoder We adopt SDv2.1 (Rombach et al., 2022) trained on 512×\times512 resolution as the default diffusion image decoder. Same as the visual encoder, the SD model is frozen without any modifications or training throughout the whole process. When constructing the SD target to compute the MSE loss, we resize the images to 512 resolution to fit its pretraining configuration.

Dream Query We use dream queries to gather semantic context from MLLMs as introduced before in Sec. 3. Without specifications, we use 64 learnable query embeddings. It is both efficient and effective in generating high-quality images. In order to predict when to generate images, we also introduce the special token, which is appended before the dream query sequence. A is appended at the end of the sequence, similar to image inputs.

Classifier-Free Guidance Classifier-free guidance (CFG) (Ho & Salimans, 2021) has been demonstrated successful in generating photo-realistic contents at the cost of acceptable generation diversity. This technique modifies the objective by ϵ^:=(1+s)ϵξ(xt,t,C)−sϵξ(xt,t,∅)\hat{\bm{\epsilon}}:=(1+s)\bm{\epsilon}_{\xi}({\mathbf{x}}_{t},t,\mathcal{C})-s\bm{\epsilon}_{\xi}({\mathbf{x}}_{t},t,\emptyset), where ∅\emptyset is a special “empty” condition representation and ss is the condition scale. The larger guidance scale generally improves image authenticity while decreasing diversity. We only adopt CFG during inference, and the scale is set to 7.5 by default and 2.0 for MS-COCO text-conditional image generation.

C.3 Evaluation Benchmarks

Systemic evaluations of DreamLLM regarding VL comprehension, content creation, and NLP capabilities have been conducted. See the used benchmarks and datasets listed in Table 9. During the evaluation, we use the prompt templates listed in Fig. 12.

Appendix D Additional Related Works

A flourishing era of Natural Language Processing (NLP) driven by LLMs is being experienced, with the parameter size growing over 100B according to the scaling law (Kaplan et al., 2020). The GPT series of models, starting with GPT-1 (Radford et al., 2018) and followed by GPT-2 (Radford et al., 2019), made significant advancements in few-shot learning by scaling up the number of parameters to 175 billion in GPT-3 (Brown et al., 2020). This breakthrough garnered a lot of attention and paved the way for further research and development in the field. Since then, researchers have focused on developing LLMs by improving the scaling strategy. Several notable efforts include Gopher (Rae et al., 2021), GaLM (Du et al., 2022), FLAN (Wei et al., 2022a), SwitchTransformer (Fedus et al., 2022), Chinchilla (Hoffmann et al., 2022), and PaLM (Chowdhery et al., 2022). Besides, instruction-based tuning techniques are explored for aligning with human preferences (Christiano et al., 2017; Ouyang et al., 2022). Such success of LLMs has been further solidified by the production release of ChatGPT (OpenAI, 2022) and the highly anticipated GPT-4 (OpenAI, 2023). Meanwhile, in the community, the open-source LLMs are achieving remarkable progress in language capabilities compared to their close-source counterparts. For example, OPT (Zhang et al., 2022), BLOOM (Scao et al., 2022), GLM (Zeng et al., 2023), LLaMA (Touvron et al., 2023a; b), and Falcon (Penedo et al., 2023) all raised great attention and are been widely deployed. Other methods attempt to learn from distillation, such as Alpaca (Taori et al., 2023) and Vicuna (Chiang et al., 2023).

D.2 Text-Conditional Content Creation with Diffusion Models

The recent surge in AI-generated content (AIGC) has been primarily driven by diffusion-based methods, particularly in the realm of text-conditional content creation. Saharia et al. (2022) have achieved astonishing advancements in high-resolution image synthesis through large-scale pretrained language models and cascaded DMs. Another paradigm, such as SD, focuses on latent spaces and demonstrates superior efficiency and performance (Rombach et al., 2022; Ramesh et al., 2022; Peebles & Xie, 2022). Recently, Lian et al. (2023) propose to enhance the reasoning capability by constructing layouts with LLMs. Motivated by the great success in 2D, a series of works have significantly propelled the 3D synthesis development (Lin et al., 2023; Wang et al., 2023c) based on Score Distillation Sampling (SDS) (Poole et al., 2023; Wang et al., 2023a) that utilizes pretrained 2D DMs. For text-to-video synthesis, the expansion of pretrained spatial to a spatial-temporal factorized U-Net with joint image and video data training has yielded significant success (Ho et al., 2022a; b).

Appendix E Limitations, Failure Cases & Future Works

While DreamLLM has made significant strides toward the development of versatile, creative, and foundational MLLMs, it still has several limitations.

Model scale. The primary constraint pertains to the scale of the LLMs utilized. Current evaluations mainly employ 7B LLMs as the base model, and despite the impressive results garnered, the potential benefits of larger model sizes, such as 65B or 130B (Kaplan et al., 2020), are worth future exploration.

Training data. The second challenge relates to the quality and quantity of training data (Jia et al., 2021). As the model size and capabilities scale up, a corresponding increase in data is crucial. However, the procurement and refinement of high-quality training data present substantial logistical and financial hurdles. For instance, the open-source interleaved dataset MMC4 contains a significant amount of noise in the form of text and images, like commercial advertisements. This noise could adversely affect the model’s output language and image style.

Prompt sensitivity. The sensitivity of LLMs to human prompts is a known issue (Wei et al., 2022b; Wang et al., 2023b; Zhou et al., 2023), a challenge that extends to MLLMs. For instance, MLLMs’ propensity for detailed responses necessitates tailored prompting to elicit concise and short answers, which is particularly useful when addressing Visual Question Answering (VQA) tasks.

Failure Cases

The main failure cases of DreamLLM are observed for multiple image-based content creations. For instance, when presented with two images and a composite instruction such as “A and B”, DreamLLM sometimes generates a single subject that amalgamates the characteristics of A and B. This output aligns more closely with the directive “A like B”. This phenomenon is not unique to DreamLLM, but is also observed in specialized compositional generation methodologies, such as StructureDiffusion (Feng et al., 2023; Chefer et al., 2023). This recurring issue may be attributed to the inherent complexity of compositional generation tasks, compounded by the severe scarcity of data specific to this domain.

Future Works

As a simple and general multimodal learning framework, our future work aims to enhance the DreamLLM framework by integrating fine-grained visual comprehension via methods like precise referring instruction tuning (Zhao et al., 2023a). We also plan to expand beyond visual and linguistic content comprehension and generation. Several promising research directions include:

Exploring applications of in-context generation capabilities of DreamLLM to complex tasks such as image-to-image translation (Isola et al., 2017; Zhang et al., 2023c; Zhang & Agrawala, 2023).

Utilizing DreamLLM’s context consistency feature for geometry-preserving tasks, including 3D content creation (Poole et al., 2023; Qi et al., 2023b; Liu et al., 2023b), representation learning (Dong et al., 2023; Qi et al., 2023a; Zhang et al., 2023a; e), scene comprehension (Zhang et al., 2023b; Hong et al., 2023), and embodied artificial inteligence (Ichter et al., 2022).

Striving to achieve a unified multimodal zero-shot generalist by extending the scope to various modalities using techniques such as ImageBind (Girdhar et al., 2023) and exploring content creation models in other modalities like audio (Kong et al., 2021).