Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Ling Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu, Stefano Ermon, Bin Cui

Introduction

Recent advancements in diffusion models (Sohl-Dickstein et al., 2015; Dhariwal & Nichol, 2021; Song et al., 2020; Yang et al., 2023c) have significantly improve the synthesis results of text-to-image models, such as Imagen (Saharia et al., 2022), DALL-E 2/3 (Ramesh et al., 2022; Betker et al., 2023) and SDXL (Podell et al., 2023). However, despite their remarkable capabilities in synthesizing realistic images consistent with text prompts, most diffusion models usually struggle to accurately follow some complex prompts (Feng et al., 2022; Lian et al., 2023; Liu et al., 2022; Bar-Tal et al., 2023), which require the model to compose objects with different attributes and relationships into a single image (Huang et al., 2023a).

Some works begin to solve this problem by introducing additional layouts/boxes (Li et al., 2023b; Xie et al., 2023; Yang et al., 2023e; Qu et al., 2023; Chen et al., 2024; Wu et al., 2023b; Lian et al., 2023) as conditions or leveraging prompt-aware attention guidance (Feng et al., 2022; Chefer et al., 2023; Wang et al., 2023) to improve compositional text-to-image synthesis. For example, StructureDiffusion (Feng et al., 2022) incorporates linguistic structures into the guided generation process by manipulating cross-attention maps in diffusion models. GLIGEN (Li et al., 2023b) designs trainable gated self-attention layers to incorporate spatial inputs, such as bounding boxes, while freezing the weights of original diffusion model.

Another potential solution is to leverage image understanding feedback (Huang et al., 2023a; Xu et al., 2023; Sun et al., 2023a; Fang et al., 2023) for refining diffusion generation. For instance, GORS (Huang et al., 2023a) finetunes a pretrained text-to-image model with generated images that highly align with the compositional prompts, where the fine-tuning loss is weighted by the text-image alignment reward. Inspired by the reinforcement learning from human feedback (RLHF) (Ouyang et al., 2022; Stiennon et al., 2020) in natural language processing, ImageReward (Xu et al., 2023) builds a general-purpose reward model to improve text-to-image models in aligning with human preference.

Despite some improvements achieved by these methods, there are still two main limitations in the context of compositional/complex image generation: (i) existing layout-based or attention-based methods can only provide rough and suboptimal spatial guidance, and struggle to deal with overlapped objects (Cao et al., 2023; Hertz et al., 2022; Lian et al., 2023) ; (ii) feedback-based methods require to collect high-quality feedback and incur additional training costs.

To address these limitations, we introduce a new training-free text-to-image generation framework, namely Recaption, Plan and Generate (RPG), unleashing the impressive reasoning ability of multimodal LLMs to enhance the compositionality and controllability of diffusion models. We propose three core strategies in RPG:

Multimodal Recaptioning. We specialize in transforming text prompts into highly descriptive ones, offering informative augmented prompt comprehension and semantic alignment in diffusion models. We use LLMs to decompose the text prompt into distinct subprompts, and recaption them with more detailed descriptions. We use MLLMs to automatically recaption input image for identifying the semantic discrepancies between generated images and target prompt.

Chain-of-Thought Planning. In a pioneering approach, we partition the image space into complementary subregions and assign different subprompts to each subregion, breaking down compositional generation tasks into multiple simpler subtasks. Thoughtfully crafting task instructions and in-context examples, we harness the powerful chain-of-thought reasoning capabilities of MLLMs (Zhang et al., 2023d) for efficient region division. By analyzing the recaptioned intermediate results, we generate detailed rationales and precise instructions for subsequent image compositions.

Complementary Regional Diffusion. Based on the planned non-overlapping subregions and their respective prompts, we propose complementary regional diffusion to enhance the flexibility and precision of compositional text-to-image generation. Specifically, we independently generate image content guided by subprompts within designated rectangle subregion, and subsequently merge them spatially in a resize-and-concatenate approach. This region-specific diffusion effectively addresses the challenge of conflicting overlapped image contents. Furthermore, we extend this framework to accommodate editing tasks by employing contour-based regional diffusion, enabling precise manipulation of inconsistent regions targeted for modification.

This new RPG framework can unify both text-guided image generation and editing tasks in a closed-loop fashion. We compare our RPG framework with previous work in Figure 1 and summarize our main contributions as follows:

We propose a new training-free text-to-image generation framework, namely Recaption, Plan and Generate (RPG), to improve the composibility and controllability of diffusion models to the fullest extent.

RPG is the first to utilize MLLMs as both multimodal recaptioner and CoT planner to reason out more informative instructions for steering diffusion models.

We propose complementary regional diffusion to enable extreme collaboration with MLLMs for compositional image generation and precise image editing.

Our RPG framework is user-friendly, and can be generalized to different MLLM architectures (e.g., MiniGPT-4) and diffusion backbones (e.g., ControlNet).

Extensive qualitative and quantitative comparisons with previous SOTA methods, such as SDXL, DALL-E 3 and InstructPix2Pix, demonstrate our superior text-guided image generation/editing ability.

Method

In this section, we introduce our novel training-free framework - Recaption, Plan and Generate (RPG). We delineate three fundamental strategies of our RPG in text-to-image generation (Section 2.2), as depicted in Figure 4. Specifically, given a complex text prompt that includes multiple entities and relationships, we leverage (multimodal) LLMs to recaption the prompt by decomposing it into a base prompt and highly descriptive subprompts. Subsequently, we utilize multimodal CoT planning to allocate the split (sub)prompts to complementary regions along the spatial axes. Building upon these assignments, we introduce complementary regional diffusion to independently generate image latents and aggregate them in each sampling step.

Our RPG framework exhibits versatility by extending its application to text-guided image editing with minimal adjustments, as exemplified in Section 2.3. For instance, in the recaptioning phase, we utilize MLLMs to analyze the paired target prompt and source image, which results in informative multimodal feedback that captures their cross-modal semantic discrepancies. In multimodal CoT planning, we generate a step-by-step edit plan and produce precise contours for our regional diffusion. Furthermore, we demonstrate the ability to execute our RPG workflow in a closed-loop manner for progressive self-refinement, as showcased in Section 2.3. This approach combines precise contour-based editing with complementary regional diffusion generation.

2 Text-to-image Generation

Let ycy^{c} be a complex user prompt which includes multiple entities with different attributes and relationships. We use MLLMs to identify the key phrases in ycy^{c} to obtain subpormpts denoted as:

where nn denotes the number of key phrases. Inspired by DALL-E 3 (Betker et al., 2023), which uses pre-trained image-to-text (I2T) caption models to generate descriptive prompts for images, and construct new datasets with high-quality image-text pairs. In contrast, we leverage the impressive language understanding and reasoning abilities of LLMs and use the LLM as the text-to-text (T2T) captioner to further recaption each subprompt with more informative detailed descriptions:

In this way, we can produce denser fine-grained details for each subprompt in order to effectively improve the fidelity of generated image, and reduce the semantic discrepancy between prompt and image.

CoT Planning for Region Division

Based on the recaptioned subprompts, we leverage the powerful multimodal chain-of-thought (CoT) reasoning ability of LLMs (Zhang et al., 2023d) to plan the compositions of final image content for diffusion models. Concretely, we divide image space H×W{H\times W} into several complementary regions, and assign each augmented subprompt y^i\hat{y}^{i} to specific region RiR^{i}:

In order to produce meaningful and accurate subregions, we need to carefully specify two components for planning region divisions: (i) region parameters: we define that rows are separated by ”;” and each column is denoted by a series of numbers separated by commas (e.g., ”1,1,1”). To be specific , we first use ”;” to split an image into different rows, then within each row, we use commas to split a row into different regions, see Figure 5 for better comprehension; (ii) region-wise task specifications to instruct MLLMs: we utilize the CoT reasoning of MLLMs with some designed in-context examples to reason out the plan of region division. We here provide a simplified template of our instructions and in-context examples:

1.Task instruction You are an smart region planner for image. You should use split ratio to specify the split method of the image, and then recaption each subregion prompts with more descriptive prompts while maintaining the original meaning. 2.Multi-modal split tutorial …… 3. In-context examples User Prompt: A girl with white ponytail and black dress are chatting with a blonde curly hair girl in a white dress in a cafe. # Key pharses extraction and Recaption # Split ratio Planning # Composition Logic # Aesthetic Considerations: # Final output 4.Trigger CoT reasoning ability of MLLMs User Prompt: An old man with his dog is looking at a parrot on the tree. Reasoning: Let’s think step by step…… To facilitating inferring the region for each subprompt, we adhere to three key principles in designing in-context example and generating informative rationales: (i) the objects with same class name (e.g., five apples) will be separately assign to different regions to ensure the numeric accuracy; (ii) for complex interactions between two entities, we take these two entities as a whole to avoid contradictory overlapped generation results mentioned in (Lian et al., 2023); (iii) If the prompt focuses more on the appearance of a specific entity, we treat the different parts of this entity as different entities (e.g., A green hair twintail in red blouse , wearing blue skirt. ⟹\Longrightarrow green hair twintail, red blouse, blue skirt).

Complementary Regional Diffusion

Recent works (Liu et al., 2022; Wang et al., 2023; Chefer et al., 2023; Feng et al., 2022) have adjusted cross-attention masks or layouts to facilitate compositional generation. However, these approaches predominantly rely on simply stacking latents, leading to conflicts and ambiguous results in overlapped regions. To address this issue, as depicted in Figure 6, we introduce a novel approach called complementary regional diffusion for region-wise generation and image composition. We extract non-overlapping complementary rectangular regions and apply a resize-and-concatenate post-processing step to achieve high-quality compositional generation. Additionally, we enhance coherence by combining the base prompt with recaptioned subprompts to reinforce the conjunction of each generated region and maintain overall image coherence (detailed ablation study in Section 4). This can be represented as:

where ss is a fixed random seed, CRD is the abbreviation for complementary regional diffusion.

More concretely, we construct a prompt batch with base prompt ybase=ycy^{\text{base}}=y^{c} and the recaptioned subprompts:

In each timestep, we deliver the prompt batch into the denoising network and manipulate the cross-attention layers to generate different latents {zt−1i}i=0n\{{\bm{z}}^{i}_{t-1}\}_{i=0}^{n} and zt−1base{\bm{z}}_{t-1}^{\text{base}} in parallel, as demonstrated in Figure 6. We formulate this process as:

where image latent zt{\bm{z}}_{t} is the query and each subprompt y^i\hat{y}^{i} works as a key and value. WQ,WK,WVW_{Q},W_{K},W_{V} are linear projections and dd is the latent projection dimension of the keys and queries. Then, we shall proceed with resizing and concatenating the generated latents {zt−1i}i=0n\{{\bm{z}}_{t-1}^{i}\}^{n}_{i=0}, according to their assigned region numbers (from to nn) and respective proportions. Here we denote each resized latent as:

where h,wh,w are the height and the width of its assigned region RiR^{i}. We directly concatenate them along the spatial axes:

To ensure a coherent transition in the boundaries of different regions and a harmonious fusion between the background and the entities within each region, we use the weighted sum of the base latents zt−1base{\bm{z}}_{t-1}^{\text{\text{base}}} and the concatenated latent zt−1cat{\bm{z}}_{t-1}^{\text{cat}} to produce the final denoising output:

Here β\beta is used to achieve a suitable balance between human aesthetic perception and alignment with the complex text prompt of the generated image. It is worth noting that complementary regional diffusion can generalize to arbitrary diffusion backbones including SDXL (Podell et al., 2023), ConPreDiff (Yang et al., 2023b) and ControlNet (Zhang et al., 2023a), which will be evaluated in Section 3.1.

3 Text-Guided Image Editing

Our RPG can also generalize to text-guided image editing tasks as illustrated in Figure 7. In recaptioning stage, RPG adopts MLLMs as a captioner to recaption the source image, and leverage its powerful reasoning ability to identify the fine-grained semantic discrepancies between the image and target prompt. We directly analyze how the input image x{\bm{x}} aligns with the target prompt ytary^{\text{tar}}. Specifically, we identify the key entities in x{\bm{x}} and ytary^{\text{tar}}:

Then we utilize MLLMs (e.g., GPT4 (OpenAI, 2023), Gemini Pro (Team et al., 2023)) to check the differences between {yi}i=0n\{y^{i}\}_{i=0}^{n} and {ei}i=0m\{e^{i}\}_{i=0}^{m} regarding numeric accuracy, attribute binding and object relationships. The resulting multimodal understanding feedback would be delivered to MLLMs for reason out editing plans.

CoT Planning for Editing

Based on the captured semantic discrepancies between prompt and image, RPG triggers the CoT reasoning ability of MLLMs with high-quality filtered in-context examples, which involves manually designed step-by-step editing cases such as entity missing/redundancy, attribute mismatch, ambiguous relationships. Here, in our RPG, we introduce three main edit operations for dealing with these issues: addition Add()\text{Add}(), deletion Del()\text{Del}(), modification Mod()\text{Mod}(). Take the multimodal feedback as the grounding context, RPG plans out a series of editing instructions. An example Plan(ytar,x)\text{Plan}(y^{\text{tar}},{\bm{x}}) can be denoted as a composed operation list:

where i,j,k<=n,length(Plan(ytar,x0))=Li,j,k<=n,\text{length}(\text{Plan}(y^{\text{tar}},x^{0}))=L. In this way, we are able to decompose original complex editing task into simpler editing tasks for more accurate results.

Contour-based Regional Diffusion

To collaborate more effectively with CoT-planned editing instructions, we generalize our complementary regional diffusion to text-guided editing. We locate and mask the target contour associated with the editing instruction (Kirillov et al., 2023), and apply diffusion-based inpainting (Rombach et al., 2022) to edit the target contour region according to the planned operation list Plan(ytar,x)\text{Plan}(y^{\text{tar}},{\bm{x}}). Compared to traditional methods that utilize cross-attention map swap or replacement (Hertz et al., 2022; Cao et al., 2023) for editing, our mask-and-inpainting method powered by CoT planning enables more accurate and complex editing operations (i.e., addition, deletion and modification).

Multi-Round Editing for Closed-Loop Refinement

Our text-guided image editing workflow is adaptable for a closed-loop self-refined text-to-image generation, which combines the contour-based editing with complementary regional diffusion generation. We could conduct multi-round closed-loop RPG workflow controlled by MLLMs to progressively refine the generated image for aligning closely with the target text prompt. Considering the time efficiency, we set a maximum number of rounds to avoid being trapped in the closed-loop procedure. Based on this closed-loop paradigm, we can unify text-guided generation and editing in our RPG, providing more practical framework for the community.

Experiments

Our RPG is general and extensible, we can incorporate arbitrary MLLM architectures and diffusion backbones into the framework. In our experiment, we choose GPT-4 (OpenAI, 2023) as the recaptioner and CoT planner, and use SDXL (Podell et al., 2023) as the base diffusion backbone to build our RPG framework. Concretely, in order to trigger the CoT planning ability of MLLMs, we carefully design task-aware template and high-quality in-context examples to conduct few-shot prompting. Base prompt and its weighted hyperparameter base ratio are critical in our regional diffusion, we have provide further analysis in Figure 16. When the user prompt includes the entities with same class (e.g., two women, four boys), we need to set higher base ratio to highlight these distinct identities. On the contrary, when user prompt includes the the entities with different class name (e.g., ceramic vase and glass vase), we need lower base ratio to avoid the confusion between the base prompt and subprompts.

Main Results

We compare with previous SOTA text-to-image models DALL-E 3 (Betker et al., 2023), SDXL and LMD+ (Lian et al., 2023) in three main compositional scenarios: (i) Attribute Binding. Each text prompt in this scenario has multiple attributes that bind to different entities. (ii) Numeric Accuracy. Each text prompt in this scenario has multiple entities sharing the same class name, the number of each entity should be greater than or equal to two. (iii) Complex Relationship. Each text prompt in this scenario has multiple entities with different attributes and relationships (e.g., spatial and non-spational). As demonstrated in Table 1, our RPG is significantly superior to previous models in all three scenarios, and achieves remarkable level of both fidelity and precision in aligning with text prompt. We observe that SDXL and DALL-E 3 have poor generation performance regarding numeric accuracy and complex relationship. In contrast, our RPG can effectively plan out precise number of subregions, and utilize proposed complementary regional diffusion to accomplish compositional generation. Compared to LMD+ (Lian et al., 2023), a LLM-grounded layout-based text-to-image diffusion model, our RPG demonstrates both enhanced semantic expression capabilities and image fidelity. We attribute this to our CoT planning and complementary regional diffusion. For quantitative results, we assess the text-image alignment of our method in a comprehensive benchmark, T2I-Compbench (Huang et al., 2023a), which is utilized to evaluate the compositional text-to-image generation capability. In Table 1, we consistently achieve best performance among all methods proposed for both general text-to-image generation and compositional generation, including SOTA model ConPreDiff (Yang et al., 2023b).

Hierarchical Regional Diffusion

We can extend our regional diffusion to a hierarchical format by splitting certain subregion to smaller subregions. As illustrated in Figure 9, when we increase the hierarchies of our region split, RPG can achieve a significant improvement in text-to-image generation. This promising result reveals that our complementary regional diffusion provides a new perspective for handling complex generation tasks and has the potential to generate arbitrarily compositional images.

Generalizing to Various LLMs and Diffusion Backbones

Our RPG framework is of great generalization ability, and can be easily generalized to various (M)LLM architectures (in Figure 10) and diffusion backbones (in Figure 11). We observe that both LLM and diffusion architectures can influence the generation results. We also generalize RPG to ControlNet (Zhang et al., 2023a) for incorporating more conditional modalities. As demonstrated in Figure 3, our RPG can significantly improve the composibility of original ControlNet in both image fidelity and textual semantic alignment.

2 Text-Guided Image Editing

In the qualitative comparison of text-guided image editing, we choose some strong baseline methods, including Prompt2Prompt (Hertz et al., 2022), InstructPix2Pix (Brooks et al., 2023) and MasaCtrl (Cao et al., 2023). Prompt2Prompt and MasaCtrl conduct editing mainly through text-grounded cross-attention swap or replacement, InstructPix2Pix aims to learn a model that can follow human instructions. As presented in Figure 12, RPG produces more precise editing results than previous methods, and our mask-and-inpainting editing strategy can also perfectly preserve the semantic structure of source image.

Multi-Round Editing

We conduct multi-round editing to evaluate the self-refinement with our RPG framework in Figure 13. We conclude that the self-refinement based on RPG can significantly improve precision, demonstrating the effectiveness of our recaptioning-based multimodal feedback and CoT planning. We also find that RPG is able to achieve satisfying editing results within 3 rounds.

Model Analysis

We conduct ablation study about the recaptioning, and show the result in Figure 14. From the result, we observe that without recaptioning, the model tends to ignore some key words in the generated images. Our recaptioning can describe these key words with high-informative and denser details, thus generating more delicate and precise images.

Effect of CoT Planning

In the ablation study about CoT planning, as demonstrated in Figure 15, we observe that the model without CoT planning fail to parse and convey complex relationships from text prompt. In contrast, our CoT planning can help the model better identify fine-grained attributes and relationships from text prompt, and express them through a more realistic planned composition.

Effect of Base Prompt

In RPG, we leverage the generated latent from base prompt in diffusion models to improve the coherence of image compositions. Here we conduct more analysis on it in Figure 16. From the results, we find that the proper ratio of base prompt can benefit the conjunction of different subregions, enabling more natural composition. Another finding is that excessive base ratio may result in undesirable results because of the confusion between the base prompt and regional prompt.

Related Work

Diffusion models (Sohl-Dickstein et al., 2015; Song & Ermon, 2019; Ho et al., 2020; Song & Ermon, 2020; Song et al., 2020) are a promising class of generative models, and Dhariwal & Nichol (2021) have demonstrated the superior image synthesis quality of diffusion model over generative adversarial networks (GANs) (Reed et al., 2016; Creswell et al., 2018). GLIDE (Nichol et al., 2021) and Imagen (Saharia et al., 2022) focus on the text-guided image synthesis, leveraging pre-trained CLIP model (Radford et al., 2021; Raffel et al., 2020) in the image sampling process to improve the semantic alignment between text prompt and generated image. Latent Diffusion Models (LDMs) (Rombach et al., 2022) move the diffusion process from pixel space to latent space for balancing algorithm efficiency and image quality. Recent advancements in text-to-image diffusion models , such as SDXL (Podell et al., 2023) Dreambooth (Ruiz et al., 2023) and DALL-E 3 (Betker et al., 2023), further improve both quality and alignment from different perspectives. Despite their tremendous success, generating high-fidelity images with complex prompt is still challenging (Ramesh et al., 2022; Betker et al., 2023; Huang et al., 2023a). This problem is exacerbated when dealing with compositional descriptions involving spatial relationships, attribute binding and numeric awareness. In this paper, we aim to address this issue by incorporating the powerful CoT reasoning ability of MLLMs into text-to-image diffusion models.

Compositional Diffusion Generation

Recent researches aim to improve compositional ability of text-to-image diffusion models. Some approaches mainly introduce additional modules into diffusion models in training (Li et al., 2023b; Avrahami et al., 2023; Zhang et al., 2023a; Mou et al., 2023; Yang et al., 2023e; Huang et al., 2023b, a). For example, GLIGEN (Li et al., 2023b) and ReCo (Yang et al., 2023e) design position-aware adapters on top of the diffusion models for spatially-conditioned image generation. T2I-Adapter and ControlNet (Zhang et al., 2023a; Mou et al., 2023) specify some high-level features of images for controlling semantic structures (Zhang et al., 2023b). These methods, however, result in additional training and inference costs. Training-free methods aim to steer diffusion models through manipulating latent or cross-attention maps according to spatial or semantic constraints during inference stages (Feng et al., 2022; Liu et al., 2022; Hertz et al., 2022; Cao et al., 2023; Chen et al., 2024; Chefer et al., 2023). Composable Diffusion (Liu et al., 2022) decomposes a compositional prompt into smaller sub-prompts to generate distinct latents and combines them with a score function. Chen et al. (2024) and Lian et al. (2023) utilize the bounding boxes (layouts) to propagate gradients back to the latent and enable the model to manipulate the cross-attention maps towards specific regions. Other methods apply Gaussian kernels (Chefer et al., 2023) or incorporate linguistic features (Feng et al., 2022; Rassin et al., 2023) to manipulate the cross-attention maps. Nevertheless, such manipulation-based methods can only make rough controls, and often lead to unsatisfied compositional generation results, especially when dealing with overlapped objects (Lian et al., 2023; Cao et al., 2023). Hence, we introduce an effective training-free complementary regional diffusion model, grounded by MLLMs, to progressively refine image compositions with more precise control in the sampling process.

Multimodal LLMs for Image Generation

Large Language Models (LLMs) (ChatGPT, 2022; Chung et al., 2022; Zhang et al., 2022; Iyer et al., 2022; Workshop et al., 2022; Muennighoff et al., 2022; Zeng et al., 2022; Taylor et al., 2022; Chowdhery et al., 2023; Chen et al., 2023b; Zhu et al., 2023; Touvron et al., 2023a; Yang et al., 2023a; Li et al., 2023a) have profoundly impacted the AI community. Leading examples like ChatGPT (ChatGPT, 2022) have showcased the advanced language comprehension and reasoning skills through techniques such as instruction tuning (Ouyang et al., 2022; Li et al., 2023c; Zhang et al., 2023c; Liu et al., 2023). Further, Multimodal Large language Models (MLLMs), (Koh et al., 2023; Yu et al., 2023; Sun et al., 2023b; Dong et al., 2023; Fu et al., 2023; Pan et al., 2023; Wu et al., 2023a; Zou et al., 2023; Yang et al., 2023d; Gupta & Kembhavi, 2023; Surís et al., 2023) integrate LLMs with vision models to extend their impressive abilities from language tasks to vision tasks, including image understanding, reasoning and synthesis. The collaboration between LLMs (ChatGPT, 2022; OpenAI, 2023) and diffusion models (Ramesh et al., 2022; Betker et al., 2023) can significantly improve the text-image alignment as well as the quality of generated images (Yu et al., 2023; Chen et al., 2023b; Dong et al., 2023; Wu et al., 2023b; Feng et al., 2023; Pan et al., 2023). For instance, GILL (Koh et al., 2023) can condition on arbitrarily interleaved image and text inputs to synthesize coherent image outputs, and Emu (Sun et al., 2023b) stands out as a generalist multimodal interface for both image-to-text and text-to-image tasks. Recently, LMD (Lian et al., 2023) utilizes LLMs to enhance the compositional generation of diffusion models by generating images grounded on bounding box layouts from the LLM (Li et al., 2023b). However, existing works mainly incorporate the LLM as a simple plug-in component into diffusion models, or simply take the LLM as a layout generator to control image compositions. In contrast, we utilize MLLMs to plan out image compositions for diffusion models where MLLMs serves as a global task planner in both region-based generation and editing process.

Conclusion

In this paper, aiming to address the challenges of complex or compositional text-to-image generation, we propose a SOTA training-free framework RPG, harnessing MLLMs to master diffusion models. In RPG, we propose complementary regional diffusion models to collaborate with our designed MLLM-based recaptioner and planner. Furthermore, our RPG can unify text-guided imgae generation and editing in a closed-loop approach, and is capable of generalizing to any MLLM architectures and diffusion backbones. For future work, we will continue to improve this new framework for incorporating more complex modalities as input condition, and extend it to more realistic applications.

References