Self-correcting LLM-controlled Diffusion Models
Tsung-Han Wu, Long Lian, Joseph E. Gonzalez, Boyi Li, Trevor Darrell
Introduction
Text-to-image generation has made remarkable advancements, especially with the advent of diffusion models. However, these models often struggle with interpreting complex input text prompts, particularly those that require skills such as understanding the concept of numeracy, spatial relationships, and attribute binding with multiple objects. Despite the astonishing scaling of model sizes and training data, these challenges, as illustrated in Fig. 1, are still present in state-of-the-art open-source and proprietary diffusion models.
Several research and engineering efforts aim to overcome these limitations. For instance, methods such as DALL-E 3 focus on the diffusion training process and incorporate high-quality captions into the training data at a massive scale. However, this approach not only incurs substantial costs but also frequently falls short in generating accurate images from complicated user prompts, as shown in Fig. 1. Other work harnesses the power of external models for a better understanding of the prompt in the inference process before the actual image generation. For example, leverages Large Language Models (LLMs) to pre-process textual prompts into structured image layouts and thus ensures the preliminary design aligns with the user’s directives. However, such integration does not resolve the inaccuracies produced by the downstream diffusion models, particularly in images with complex scenarios like multiple objects, cluttered arrangements, or detailed attributes.
Drawing inspiration from the from the process of a human painting and a diffusion model in generating images, we observe a key distinction in their approach to creation. Consider a human artist tasked with painting a scene featuring two cats. Throughout the painting process, the artist remains cognizant of this requirement, ensuring that two cats are indeed present before considering the work complete. Should the artist find only one cat depicted, an additional one would be added to meet the prompt’s criteria. This contrasts sharply with current text-to-image diffusion models, which operate on an open-loop basis. These models generate images through a predetermined number of diffusion steps and present the output to the user, regardless of its alignment with the initial user prompt. Such a process, irrespective of scaling training data or LLM pre-generation conditioning, lacks a robust mechanism to ensure the final image aligns with the user’s expectations.
In light of this, we propose our method Self-correcting LLM-controlled Diffusion (SLD) that performs self-checks to confidently offer users guarantees of the alignment between the prompt and the generated images. Departing from conventional single-round generation methods, SLD is a novel closed-loop approach that equips diffusion models with the ability to iteratively identify and rectify errors. Our SLD framework, illustrated in Fig. 2, contains two main components: LLM-driven object detection as well as LLM-controlled assessment and correction.
The SLD pipeline follows a standard text-to-image generation setting. Given a textual prompt that outlines the desired image, SLD begins with calling an image generation module (e.g., the aforementioned open-loop text-to-image diffusion models) to generate an image in a best-effort fashion. Given that these open-loop generators do not guarantee an output that aligns perfectly with the prompt, SLD then conducts a thorough evaluation of the produced image against the prompt, with an LLM parsing key phrases for an open-vocabulary detector to check. Subsequently, an LLM controller takes the detected bounding boxes and the initial prompt as input, checks for potential mismatches between the detection results and the prompt requirements, suggesting appropriate self-correction operations, such as adding, moving, and removing objects. Finally, utilizing a base diffusion model (e.g., Stable Diffusion ), SLD employs latent space composition to implement these adjustments, thereby ensuring that the final image accurately reflects the user’s initial text prompt.
Notably, our pipeline does not pose restrictions on the source of the initial generation and thus is applicable to images generated by proprietary models, such as DALL-E 3 , through APIs, as shown in Fig. 1. Furthermore, none of the self-correction operations require any additional training on our base diffusion model , which easily allows our method to be applied to various diffusion models without the costs of external human annotation or training.
We demonstrate that our SLD framework can achieve significant improvement over current diffusion-based methods on cpmplex prompts with the LMD benchmark . The results show that our method is able to surpass LMD+, which is a strong baseline that already leverages LLM in the image process generation, by . More importantly, with DALL-E 3 for initial generation, the generated images from our method achieve performance gains compared to ones before self-correction.
Finally, since the SLD pipeline is agnostic to the initially generated image, it can easily be transformed into an image editing pipeline by simply changing the prompts to the LLM. While text-to-image generation and image editing are often treated as distinct tasks by the generative modeling community, our SLD is able to perform these two tasks with a unified pipeline. We list our key contributions below:
SLD is the first to integrate a detector and an LLM to self-correct generative models, ensuring accurate generation without extra training or external data.
SLD offers a unified solution for both image generation and editing, enabling enhanced text-to-image alignment for any image generator (e.g., DALL-E 3) and object-level editing on any images.
Experimental results show that our approach can correct a majority of incorrect generations, particularly in aspects of numeracy, attribute binding, and spatial relationships.
We will release our code for future research and applications.
Related Work
Diffusion-based text-to-image generation has advanced significantly. Initial studies showed diffusion models’ ability to create high-quality images, but they struggle with complex prompts. Subsequent research has incorporated additional inputs such as keypoints and bounding boxes to control the diffusion generation process.
Recent advancements have incorporated LLMs to control the generation of diffusion models, bypassing the need for additional complementary information as inputs . In these approaches, LLMs play a central role in directly interpreting user textual prompts and managing the initial layout configuration. Despite some progress, these models often operate in an open-loop fashion, producing images in one iteration that cannot guarantee the generated images align with user prompts.
Unlike prior work, SLD is the first closed-loop diffusion-based generation method. Integrating advanced object detectors and LLMs, SLD performs iterative self-checking and correction, significantly enhancing text-to-image alignment. This improvement spans numeracy, attribute binding to multiple objects, and spatial reasoning, applicable to various models, including those like DALL-E 3.
2 Diffusion-based Image Editing
Recent advancements in text-to-image diffusion models have significantly expanded their applications in image editing, encompassing both global and local editing tasks. Techniques like Prompt-2-prompt and InstructPix2Pix specialize in global edits, such as style transformations. Conversely, methods like SDEdit , DiffEdit , and Plug-and-Play focus on local edits, targeting specific areas within images. Despite their progress, these methods often struggle with precise object-level manipulation and tasks that require spatial reasoning, such as resizing or repositioning objects. While recent approaches like Self-Guidance offer fine-grained operations, they still necessitate user inputs for specific coordinates when moving or repositioning objects.
Unlike these methods only focusing on diffusion models, SLD introduces the combination of detectors and LLMs in the loop of editng, enabling fine-grained editing with only user prompts. Also, SLD excels in a variety of object-level editing tasks, including adding, replacing, moving, and modifying attributes, swapping, and so on, demonstrating a notable improvement in both ease of use and editing capabilities.
Self-correcting LLM-controlled Diffusion
In this section, we introduce our Self-correcting LLM-controlled Diffusion (SLD) framework. SLD consists of two main components: LLM-driven object detection (Sec. 3.1) as well as LLM-controlled assessment and correction (Sec. 3.2). Moreover, with a simple change of the LLM instructions, we show that SLD is applicable to image editing, unifying text-to-image generation and editing as discussed in Sec. 3.3. The complete pipeline is shown in Algorithm 1.
Our SLD framework starts with LLM-driven object detection, which extracts information necessary for downstream assessment and correction. As shown with green arrows in Fig. 2, the LLM-driven object detection includes two steps: 1) We leverage an LLM as a parser that parses the user prompt and outputs key phrases that are potentially relevant to image assessment. 2) The phrases are then passed into an open-vocabulary object detector. The detected boxes are supposed to contain information that supports the assessment of whether the image is aligned with the specifications in the user prompt.
In the initial step, an LLM parser is directed to extract a list of key object details, denoted as , from the user-provided text prompt . This parser, aided by text instructions and in-context examples, can easily accomplish this as shown in Fig. 3 (a). For a user prompt that includes phrases like “a green motorcycle” and “a blue motorcycle,” the LLM is expected to identify and output “green” and “blue” as attributes associated with “motorcycle.” When the prompt references objects without specific quantities or attributes, such as “a monkey” and “a raccoon,” these descriptors are appropriately left blank. Importantly, the LLM’s role is not limited to merely identifying object nouns; it also entails identifying any associated quantities or attributes.
In the second step, an open-vocabulary detector processes the list of key object information, , parsed in the first step, to detect and localize objects within the image. We prompt the open-vocabulary object detector with queries formatted as image of a/an [attribute] [object name], where the “attribute” and “object name” are sourced from the parser’s output. The resulting bounding boxes, , are then organized into a list format like [("[attribute] [object name] [#object ID]", [x, y, w, h])] for further processing. A special case is when the prompt poses constraints on the object quantity. For cases where attributed objects (e.g., “blue dog”) fall short compared to the required quantities, a supplementary count of non-attributed objects (e.g., “dog”) is provided to provide context for the subsequent LLM controller deciding whether to add more “blue dogs” or simply alter the color of existing dogs to blue. We will explain these operations, including object addition and attribute modification, in greater detail in Sec. 3.2.1.
2 LLM-controlled Analysis and Correction
We use an LLM controller for image analysis and the subsequent correction. The controller, given the user prompt and detected boxes , is asked to analyze whether the image, represented by objects bounding boxes, aligns with the description of the user prompt and offer a list of corrected bounding boxes , as shown in Fig. 3 (b).
SLD then programmatically analyzes the inconsistencies between the refined and original bounding boxes to output a set of editing operations , which includes addition, deletion, repositioning, and attribute modification. However, a simple set-of-boxes representation does not carry correspondence information, which does not allow an easy way to compare the input and the output layout of the LLM controller when multiple boxes share the same object name. For example, when there are two cat boxes in both the model input and the model output, whether one cat box corresponds to which cat box in the output layout is unclear. Rather than introducing another algorithm to guess the correspondence, we propose to let the LLM output correspondence with a very simple edit: we give an object ID to each bounding box, with the number increasing within each object type, as a suffix added after the object name. In the in-context examples, we demonstrate to the LLM that the object should have the same name and object ID before and after the proposed correction.
The LLM controller outputs a list of correction operations to apply. For each operation, we first transform the original image into latent features. Our approach then executes a series of operations , such as addition, deletion, repositioning, and attribute modification, applied to these latent layers. We explain how each operation is performed below.
Addition. Inspired by , the addition process entails two phases: pre-generating an object and integrating its latent representation into the original image’s latent space. Initially, we use a diffusion model to create an object within a designated bounding box, followed by precise segmentation using models (e.g., SAM ). This object is then processed through a backward diffusion sequence with our base diffusion model, yielding masked latent layers corresponding to the object, which are later merged with the original canvas.
Deletion operation begins with SAM refining the boundary of the object within its bounding box. The latent layers associated with these specified regions are then removed and reset with Gaussian noise. This necessitates a complete regeneration of these areas during the following forward diffusion process.
Repositioning involves modifying the original image to align objects with new bounding boxes, taking care to preserve their original aspect ratios. The initial steps include shifting and resizing the bounding box in the image space. Following this, SAM refines the object boundary, succeeded by a backward diffusion process to generate its relevant latent layers, similar to the approach in the addition operation. Latent layers corresponding to the excised parts are replaced with Gaussian noise, while the newly added sections are integrated into the final image composition. An important consideration in repositioning is conducting object resizing in the image space rather than the latent space to maintain high-quality results.
Attribute modification starts with SAM refining the object boundary within the bounding box, followed by applying attribute modifications such as DiffEdit . The base diffusion model then reverses the image, producing a series of masked latent layers ready for final composition.
After editing operations on each object, we proceed to the recomposition phase as shown in Fig. 4. In this phase, while latents for removed or repositioned regions are reinitialized with Gaussian noise, the latents for added or modified latents are updated accordingly. For regions with multiple overlapping objects, we place the larger masks first to ensure the visibility of the smaller objects.
The stack of modified latent then undergoes a final forward diffusion process, which begins with steps in which regions not reinitialized with Gaussian noise are frozen (i.e., forced to align with the unmodified latent at the same step). This is crucial for the accurate formation of updated objects while maintaining background consistency at the same time. The procedure finishes with several steps where everything is allowed to change, resulting in a visually coherent and correct image.
2.2 Termination of the Self-Correction Process
Even though we observe that one round of generation is often enough for a majority of the cases that we encountered, subsequent rounds could still benefit the performance in terms of correctness further, making our self-correction an iterative process.
Determining the optimal number of self-correction rounds is critical for balancing efficiency and accuracy. As outlined in Algorithm 1, our method sets a maximum number of attempts on the correction rounds to ensure the process finishes within a reasonable amount of time.
The process completes when the LLM outputs the same layout as the input (i.e., if the bounding boxes suggested by the LLM controller () align with the current detected bounding boxes ()), or when the maximum rounds of generation are reached, which indicates that the method is unable to make a correct generation for the prompt. This iterative process provides guarantees on the correctness of the image, up to the accuracy of the detector and the LLM controller, ensuring it aligns closely with the initial text prompt. We explore the efficacy of multi-round corrections in Sec. 4.3.
3 Unified text-to-image generation and editing
In addition to self-correcting image generation models, our SLD framework is readily adaptable for image editing applications, requiring only minimal modifications. A key distinction is in the format of user input prompts. Unlike image generation, where users provide scene descriptions, image editing requires users to detail both the original image and the desired changes. For instance, to edit an image with two apples and a banana by replacing the banana with an orange, the prompt could be: “Replace the banana with an orange, while keeping the two apples unchanged.”
The editing process is similar to our self-correction mechanism. The LLM parser extracts key objects from the user’s prompt. These objects are then identified by the open-vocabulary detector, establishing a list of current bounding boxes. The editing-focused LLM controller, equipped with specific task objectives, guidelines, and in-context examples, analyzes these inputs. It proposes updated bounding boxes and corresponding latent space operations for precise image manipulation.
SLD’s ability to perform detailed, object-level editing distinguishes it from existing diffusion-based methods like InstructPix2Pix and prompt2prompt , which mainly address global image style changes. Also, SLD outperforms tools like DiffEdit and SDEdit , which are restricted to object replacement or attribute adjustment, by enabling comprehensive object repositioning, addition, and deletion with exact control. Our comparative analysis in Sec. 4.2 will further highlight SLD’s superior editing capabilities over existing methods.
Experiments
Setup. We evaluate the performance of the SLD framework with the LMD benchmark , which is specifically designed to evaluate generation methods on complex tasks such as handling negation, numeracy, accurate attribute binding to multiple objects, and spatial reasoning. For each task, 100 programmatically generated prompts are fed into various text-to-image generation methods to produce corresponding images. We evaluate the images generated by our method and the baselines with open-vocabulary detector OWL-ViT v2 for a robust quantitative evaluation of the alignment between the input prompts and the generated images. We compared SLD with several leading text-to-image diffusion methods, such as Multidiffusion , BoxDiff , LayoutGPT , LMD+ , and DALL-E 3 . To ensure fair comparisons, all models incorporating LLMs used the same GPT-4 model. For our SLD implementation, we utilized LMD+ as the base model for latent space operations and OWL-ViT v2 for the open-vocabulary object detector.
Results. As shown in Tab. 1, applying the SLD method to both open-source (LMD+) and proprietary models (DALL-E 3) significantly enhances their performance in terms of generation correctness. For negation tasks, as LMD+ converts user prompts containing “without” information into negative prompts, which already achieves a remarkable 100% accuracy without SLD integration. In contrast, even though DALL-E 3 also uses an LLM to rewrite the prompt, it still fails for some negation cases, likely because the LLM simply puts the negation keyword (e.g., “without”) into the rewritten prompt. In this case, our SLD method can automatically rectify most of these errors. For numeracy tasks, integrating SLD with LMD+ results in a significant improvement, with up to 98% accuracy. We noted that DALL-E 3 often struggles to generate an image with the correct number of objects. However, this issue is substantially mitigated by SLD, which enhances performance by over 20%. For attribute binding tasks, SLD improves the performance of both DALL-E 3 and LMD+ by 6% and 14%, respectively. Notably, DALL-E 3 initially outperforms LMD+ in this task, likely due to its training on high-quality image caption datasets. Finally, for spatial reasoning tasks, the integration of SLD with both LMD+ and DALL-E 3 demonstrates enhanced performance by 12% and 6%, respectively.
2 Application to Image Editing
As discussed in Sec. 3.3, SLD excels in fine-grained image editing over existing methods. As demonstrated in Fig. 6, our integration of an open-language detector with LLMs enables precise modifications within localized latent space regions. SLD adeptly performs specific edits, like seamlessly replacing an apple with a pumpkin, while preserving the integrity of surrounding objects. In contrast, methods like InstructPix2Pix are confined to global transformations, and DiffEdit often fails to accurately locate objects for modification, leading to undesired results.
Furthermore, as exemplified in Fig. 7, SLD supports a wide array of editing instructions, including counting control (such as adding, deleting, or replacing objects), attribute modification (like altering colors or materials), and intricate location control (encompassing object swapping, resizing, and moving). A standout example is featured in the “Object Resize” column of the first row, where SLD precisely enlarges the cup on the table by an exact factor of 1.25. We encourage readers to verify this with a ruler for a clear demonstration of our method’s precision. This level of precision stems from the detector’s exact object localization coupled with the LLMs’ ability in reasoning and suggestions for new placements. Such detailed control over spatial adjustments is unmatched by any previous method, highlighting SLD’s contributions to fine-grained image editing.
3 Discussion
Multi-round self-correction. Our analysis in Tab. 2 highlights the benefits of multi-round self-correction and the fact that the first round correction is always the most effective one and has a marginal effect. The first round of corrections substantially mitigates issues inherent in Stable Diffusion . Then, a second round of correction still yields significant improvements across all four tasks.
Limitations and future work. A limitation of our method is illustrated in Fig. 8, where SLD fails to accurately remove a person’s hair. In this instance, despite the successful identification and localization of the hair, the complex nature of its shape poses a challenge to the SAM module used for region selection, resulting in the unintended removal of the person’s face in addition to the hair. However, since the person’s cloth is not removed, the base diffusion model fails to generate a natural composition. This suggests that a better region selection method is needed for further improvements in the generation and editing quality.
Conclusion
We introduce the Self-correcting Language-Driven (SLD) framework, a pioneering self-correction system using detectors and LLMs to significantly enhance text-to-image alignment. This method not only sets a new SOTA in image generation benchmark but is also compatible with various generative models, including DALL-E 3. Also, SLD extends its utility to image editing applications, offering fine-grained object-level manipulation that surpasses existing methods.
References
Appendix A More Comparisons with Image Editing Methods
In Fig. 6 of the main paper, we showcased an example of object replacement contrasting our SLD method with previous approaches. In this section, we provide more visual examples in Fig. 9 that highlight the differences between SLD and prior diffusion-based image editing methods. While InstructPix2Pix is confined to pixel-aligned changes (e.g., style changes), DiffEdit often fails with precise object-level edits. Our SLD framework excels in these detailed, fine-grained editing tasks.
Appendix B Comparison between LMMs and Detector-LLM Combinations
In our main paper, we introduce a self-correction framework that utilizes both LLMs and object detectors for image assessment, followed by the provision of correction suggestions. With the rapid evolution of Large Multimodal Models (LMMs) such as GPT-4V and LLaVA , we have investigated the feasibility of using an LMM to conduct image assessment. In this setup, the LMM evaluates images generated by the open-loop generator alongside user prompts, aiming to provide precise, object-level editing recommendations.
However, GPT-4V’s analysis of the “princess and dwarfs” image from DALL-E 3 (refer to Fig. 1 in our paper) reveals inaccuracies, as the model miscounts the characters, identifying seven dwarfs rather than the actual five, and struggles to define precise bounding box coordinates. In contrast, as highlighted in Fig. 10, the latest generation open-vocabulary detector, OWL-ViT v2, demonstrates remarkable proficiency in detecting minute objects (seagulls in the sky), which is vital for the correction process. These limitation underscores our current approach, which combines detectors with LLMs for more accurate assessment and editing suggestions. Nevertheless, the potential of integrating advanced LMMs for streamlined image generation and editing remains a compelling and promising direction for future research and development.
Appendix C Comprehensive Image Generation Results
In the main paper, we demonstrate the significant performance gain achieved by integrating the OWL-ViT v2 open-vocabulary detector into our SLD pipeline. Since this detector is also used in the benchmark, we also explore using other detectors in our setting, which makes our self-correction detector distinct from the detector used in the evaluation benchmark.
Tab. 3 presents a comparison between the results obtained using OWL-ViT v1 and OWL-ViT v2 detectors within our framework. The results indicate that SLD consistently enhances overall accuracy compared to the baseline text-to-image generators, irrespective of the detector used. The marginal reduction in performance gains when substituting the v2 detector can be attributed to two main factors: 1) the relatively inferior detection capabilities of the alternate detector employed in the SLD, and 2) the variations in object recognition between the two detectors. This situation parallels real-world experiences, where individual perceptions and recognition of objects or attributes can vary significantly. For instance, a bowl perceived as distinctly blue by one might be seen as less blue by another. These perceptual variances contribute to the marginally lower attribute binding scores of SLD compared to original DALL-E 3 results. Despite these discrepancies, the overall accuracy of our method, especially in areas such as numeracy and spatial reasoning, confirms the effectiveness of SLD, regardless of the specific detector used in our self-correction pipeline.
Appendix D Our Prompts and Instructions to the LLM
As outlined in the main paper, we leverage two LLMs to steer the self-correction process: we employ one LLM parser to identify key objects from user prompts and another LLM controller to propose bounding box adjustments. The specific prompts for both the parser and the controller are detailed in Tab. 4 and Tab. 5, respectively. We also provide in-context examples for the LLM controller, tailored for self-correcting generation and image editing, in Tab. 6 and Tab. 7, respectively.
In crafting our LLM prompts, we emphasized clarity in defining the roles and guidelines for the LLM, drawing inspiration from prior work . Notably, we discovered that GPT-4 possesses the capability to manipulate bounding box coordinates, a task that involves mathematical reasoning. We achieved improved results by guiding the model to employ chain-of-thought reasoning , where the model explicates its reasoning process during generation. This approach yielded more accurate suggestions compared to instances where the model’s reasoning was not explicitly stated.