LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image Generation
Leigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie, Tat-Seng Chua
Introduction
In the latest days, the topic of AI-Generated Content (AIGC) has made thrilling progress, such as DELL-E 2 (Ramesh et al., 2022a), Stable Diffusion (SD) (Rombach et al., 2022a), and ChatGPT (Ouyang et al., 2022). As one of the representative generative AI themes, text-to-image generation (T2I) has received extensive attention from both academia and industry. Given input language prompts, T2I aims to produce images that accurately reflect the desired contents as well as their semantic correlations. Currently, the diffusion-based models have become the state-of-the-art (SoTA) T2I method, due to the preferable distribution coverage, a stationary training objective, and easy scalability (Ho et al., 2020; Dhariwal and Nichol, 2021; Rombach et al., 2022a). Despite the satisfactory performance achieved by recent SD-based models, synthesizing high-faithful images in complex scenes is still challenging (Saharia et al., 2022; Ramesh et al., 2022b). In Figure 1 we showcase several representative issues in current SD-based T2I,Here we generate images using the official SD model with v1-4 checkpoint weights from https://github.com/CompVis/stable-diffusion. such as problematic spatial relation understanding and numeration failure.
Diffusion models are competent in accurately rendering the visual objects by recognizing the explicit entity mentions of interest from prompt texts. However, we argue that the key to high-faithfulness image synthesis, especially for complex scenes, also lies in the rigorous understanding of the underlying layout and delicate interactions between objects. We note that some recent SD-based methods, e.g., ControlNet (Zhang and Agrawala, 2023a), in combination with additional human guidance can promisingly handle complex-scene T2I, while this work mainly considers fully automatic solutions without human efforts. Intuitively, whenever we humans create a fine picture by the prompt instruction, we often follow a two-stage drawing process. First, we pin down the general layout of the overall picture, i.e., sketching out all the objects as well as their relative semantic relations. With the top-level design of the picture, we then complete all the necessary details. In a nutshell, high-faithful image synthesis further requires the capability of high-level planning. Inspired by such coarse-to-fine drawing intuition, in this work, we investigate endowing T2I models with scene layout planning abilities, such that the model is able to validly devise the coarse-grained architecture and the semantic structure before rendering the fine-grained details.
However, it is non-trivial to achieve high-faithfulness image synthesis via the above-mentioned coarse-to-fine framework, due to the following challenges. 1) Layout Planning requires abstract spatial imagination and analysis capabilities. The limited annotated layout data and intrinsic inductive bias make it difficult for existing diffusion methods (Rombach et al., 2022a; Nichol et al., 2022) to accurately and aesthetically generate layouts. Although notable efforts (Li et al., 2023b; Mou et al., 2023; Zhang and Agrawala, 2023b) have been dedicated to synthesizing complex scenes by manually providing guidance information, these strategies suffer from weak flexibility and low efficiency since they heavily rely on extra labor-intensive guidance. And 2) Relation Modeling, e.g., expressing high-level spatial and semantic relations, plays a pivotal role in understanding, imagining, and depicting complex scenes for T2I models, but it is still under-explored owing to the complex environments in real life.
Facing these two challenges, we propose an effective model by eliciting layout guidance from LLM for high-faithful T2I generation (LayoutLLM-T2I). As shown in Figure 2, our framework comprises two main modules, including the text-to-layout induction and the layout-guided text-to-image generation. In the first stage, we explore the scene understanding ability of large language models (LLMs), e.g., ChatGPT, for layout planning. To fully stimulate this ability, we design a feedback-based sampler learning mechanism, which is able to adaptively select informative examples for in-context learning, guided by layout-level and image-level feedback. During the second layout-guided text-to-image generation stage, based on the parameter-frozen SD model we devise a layout-aware adapter, in which the mentioned entities with well-organized layouts and their semantic relation information are injected into the backbone SD with fine-grained adequate interaction. We perform extensive experiments on T2I benchmarks, where the proposed model achieves new SoTA results over existing methods, demonstrating the effectiveness of the layout guidance for diffusion-based T2I. We also show that the proposed feedback-based sampler learning mechanism is beneficial in activating high-quality layout planning; and the layout-guided adapter helps maintain effective layout feature integration for T2I synthesis. In-depth experiments and analysis demonstrate that our method improves T2I generation, especially in complex-scene cases and zero-shot settings.
To summarize, our contributions are three-fold:
To the best of our knowledge, it is the first work to investigate layout planning under complex natural scenes in the context of LLMs and Diffusion models.
We propose a feedback-based sampler learning paradigm for layout generation and a layout-guided object interaction scheme for conditioned image synthesis.
The proposed framework empirically pushes the current SoTA T2I performance, achieving high-faithfulness image synthesis in complex scenes.
Related Work
Text-to-image generation, a.k.a., text-conditional image synthesis, has been the key research topic in the multimodal learning community. There have been a number of efforts devoted to generating realistic and natural-looking images. The generative adversarial networks (GANs) (Goodfellow et al., 2014; Reed et al., 2016b) are a popular class of generative models that use a two-part network: a generator and a discriminator, while variational autoencoders (VAEs) (Kingma and Welling, 2014) apply a probabilistic encoder-decoder architecture. Recently, inspired by the application of auto-regressive models (ARMs) in text generation, numerous work has adopted ARMs to achieve impressive results for text-to-image generation, such as DALL-E (Ramesh et al., 2021), CogView (Ding et al., 2021), and Pariti (Yu et al., 2022). Despite their success, existing T2I generation models still suffer from some weaknesses, such as training instability in GANs (Rombach et al., 2022b) and unidirectional bias in ARMs (Gu et al., 2022). The diffusion models (DM) have currently emerged as the SoTA T2I approaches (Rombach et al., 2022b; Nichol et al., 2022; Saharia et al., 2022), due to the natural fit to inductive biases of image data, leading to remarkable synthesis quality. For example, Rombach et al. (2022b) proposed a latent diffusion model that enables DM training on limited computational resources while retaining their quality and flexibility. Nichol et al. (2022) presented GLIDE, an effective text-guidance strategy leading to photorealistic image generation and editing.
While most of the existing methods have secured satisfactory performance for the T2I (Reed et al., 2016b; Bao et al., 2017), generating high-fidelity images in complex scenes that faithfully reflect the original text prompt is still challenging (Nie et al., 2022; Qu et al., 2021). In the realistic world, it is ubiquitous that user prompts come with complicated descriptions, i.e., multiple various objects in complex interrelation (such as spatial relations, action-based semantic relations, and numeric relations). Correspondingly, prior efforts have been paid to model complex scenes, e.g., stacking multiple GANs (Zhang et al., 2017), conditioning on scene graphs (Johnson et al., 2018), and introducing attentional generative networks (Xu et al., 2018). However, few attempts focused on enhancing the faithfulness for T2I. Feng et al. (2022) proposed to integrate the syntactic structure of the prompt sentence such that the key objects and the corresponding relations can be learned more correctly. To strengthen the modeling of object spatial relations, additional segmentation features (Avrahami et al., 2022; Gafni et al., 2022), or spatial conditioning (Ruiz et al., 2022; Voynov et al., 2022; Bar-Tal et al., 2023; Mou et al., 2023; Reed et al., 2016a; Hinz et al., 2019) are integrated into the visual synthesis process for achieving higher faithfulness. In this work, we argue that the key to high-faithfulness T2I generation lies in the layout planning and comprehensive understanding of the underlying interactions between objects. We draw inspiration from human intuition and consider strengthening the image generation in complex scenes by taking advantage of the high-level layout features as guidance for high-fidelity diffusion-based T2I.
Previous research has demonstrated that modeling the high-level object layout information helps to capture the underlying abstract semantic relations and results in better vision generation (Johnson et al., 2018; Hong et al., 2018; Vo and Sugimoto, 2020). Some works study the task of image synthesis from layout input (i.e., layout-to-image), where GANs (Sun and Wu, 2019; Ma et al., 2020; He et al., 2021) are employed. Recently, diffusion models are adopted for layout-to-image and achieve more reliable image generation (Li et al., 2023b; Cheng et al., 2023; Zheng et al., 2023). Different from these works, in this study, we focus on the T2I setting without giving any extra layout information, i.e., only textual prompts as input. By eliciting the layout generation abilities from LLMs by a feedback-based sampler, we achieve high-quality layout label acquisition without relying on any human effort. Besides, we devise an effective strategy to integrate layouts into the diffusion process.
Preliminary on Latent Diffusion
In this paper, we apply our method based on the open-sourced SD model (Rombach et al., 2022b). SD employs a hierarchical VAE to operate the diffusion process in low-dimensional latent space, instead of operating in the image space, improving the computational efficiency. Technically, an encoder of VAE maps a given image into a spatial latent code , i.e., . A diffusion model (Ho et al., 2020) operates over the learned latent space to produce a denoised version of an input latent at each timestep conditioned on addition input. In a text-to-image scenario, this additional input is typically a text encoded by a pre-trained CLIP text encoder (Radford et al., 2021). During the training process, at each timestep , the denoising network is optimized to remove the noise added to the latent code , given the noised latent , the timestep , and the conditioning text :
Here, is often implemented with a UNet (Ronneberger et al., 2015) consisting of convolution, self-attention, and cross-attention layers.
At inference time, a sampling process is performed to iteratively denoise with as the start. Specifically, at each denoising step , is obtained by denoise conditioned on the text prompt . After the final denoising step, will be mapped back to the original image space, generating an image by a decoder of VAE, .
Methodology
In Figure 2, we illustrate the overall architecture of the proposed layout-guided diffusion model, consisting of two modules. First, the text-to-layout induction module (Section 4.1) infers a coarse-grained layout via an LLM conditioned on the given textual prompt. Combining the prompt and the generated layout, the layout-guided image generation module (Section 4.2) synthesizes the final image. In what follows, we will delve into these two modules.
Recent years have witnessed the tremendous potential of LLMs (Touvron et al., 2023; OpenAI, 2023; Chowdhery et al., 2022). Benefiting from the large corpus and ample computing resources, they achieve outstanding performance in most natural language processing (NLP) tasks, especially under the challenging zero-shot or few-shot settings (OpenAI, 2023). The impressive success of LLMs in NLP illustrates the multifaceted abilities of LLMs. Inspired by it, we aim to excavate the spatial imagination, semantic relation, and numeration understanding abilities of LLMs toward layout planning and facilitate the text-to-image generation task.
Concretely, we resort to in-context learning (ICL) (Wei et al., 2022) to activate LLMs for layout generation. Typically, ICL employs a natural language prompt that includes a task description (Instruction), a few examples (in-context examples) selected from the training dataset as demonstrations, and a test instance (Test), as depicted in Figure 3. Previous studies have shown that the effectiveness of ICL is highly influenced by the design of demonstrations (Min et al., 2022; Lu et al., 2022; Zhao et al., 2021). Therefore, it is essential to select a subset of examples that can effectively leverage the ICL capability of LLMs. To tackle this issue, we devise an adaptive sampler based on layout-level and image-level feedback to select examples in a reinforcement learning framework. This framework mainly consists of three parts, i.e., policy network, reward, and optimization.
Policy Network. We randomly sample instances from the training set to form a candidate set . Given a text , we aim to select suitable in-context examples . The selection is modeled by a policy network parameterized by :
where is independently sampled from the candidate set. In practice, the policy is implemented as,
where denotes the text with respect to the candidate . acts as a mapping function that transforms a text into a latent layout embedding. In this latent space, two sentences describing similar layouts will be mapped close to each other.
Combining the given text, the selected in-context examples, and the instruction as the prompt, we obtain the layout from an LLM:
Reward. As discussed in Section 1, layouts play a key role in text-to-image generation without other fine-grained guidance. Meanwhile, the final aim is to generate a reasonable and aesthetic image to satisfy user intention. With these two aspects in consideration, based on the generated layout , we define the total reward as:
where and denote the layout reward and image reward, respectively. Specifically, they are calculated by:
where refers to the maximum intersect over union (Kikuchi et al., 2021) between the induced layout and the ground-truth layout , measuring the layout similarity in the spatial dimension. Besides, denotes the CLIP (Radford et al., 2021) similarity of the generated image from to the ground-truth one and the given text , respevtively. Concretely, we employ both intra-modal (image-to-image) and cross-modal (image-to-text) similarities, i.e., = . In addition to the semantic alignment, we also consider another aspect, i.e., aesthetics, to measure the image generation quality. In detail, we adopt the aesthetic predictorhttps://github.com/christophschuhmann/improved-aesthetic-predictor trained on the LAION dataset (Schuhmann et al., 2022) to calculate the aesthetic score .
Optimization. To optimize the policy network, we first carry out Monte Carlo Sampling (Shapiro, 2003) to estimate the expected reward:
in which denotes the batch size. And then we perform optimization using the REINFORCE policy gradient algorithm (Williams, 1992):
By maximizing the expected reward, the policy network learns to select those in-context examples which motivate the LLM to generate a reasonable and aesthetic layout. Meanwhile, the induced layout could guide the image generation model to synthesize a high-quality and high-faithfulness image.
2. Layout-guided Image Generation
In the above coarse-grained layout planning process, we activate an LLM to generate reasonable and aesthetic layouts. However, an accurate layout does not guarantee high-faithfulness image generation, since the same layout can induce multiple images with different semantics. For example, given the two prompts, “A man walks towards a traffic light” and “A man looks at the traffic lights”, two similar spatial arrangements could be obtained, where a man is on the left side of the image and a traffic light is on the right side. In light of this, it is essential to consider semantic relation modeling and scene understanding during the image generation process. Toward this end, we endow such capabilities to the diffusion model via relation-aware object interaction.
Condition Encoder. To encode the text prompt , we leverage the pre-trained CLIP (Radford et al., 2021) to yield a feature sequence . Furthermore, resorting to scene graph parserhttps://github.com/vacancy/SceneGraphParser, we capture explicit semantic relations by extracting object-predicate-object phrases , and then represent them as:
After the layout induction presented in Section 4.1, we obtain the layout in which represent the object textual label of the bounding box . Afterwards, we encode a bounding box coordinated with Fourier (Tancik et al., 2020) mapping. As shown in Figure 2, we concatenate label features and bounding box features, and feed them into a multi-layer perception (MLP):
where denotes the layout feature sequence.
Relation-aware Image Generation. Existing models (Li et al., 2023b; Mou et al., 2023; Zhang and Agrawala, 2023b) have demonstrated the potential of SD to generate high-quality images based on layout information offered by users. In this paper, based on GLIGEN (Li et al., 2023b), we present the relation-aware image generation module. In GLIGEN, two attention layers are frozen in the original Transformer block of SD, and an extra gated self-attention layer is added as an adapter to model the cross-modal interaction between intermediate visual features and layout features :
where is a token selection operation that considers visual tokens only and is a learnable scalar. acts as a hyperparameter to balance quality and controllability.
Though the self-attention operation over the combination of and encourages the interaction between layout, text, and image tokens, the intact visual object, as well as relations, are not considered. Therefore, we first select visual objectsNote that denotes visual tokens. Consequently, an intact object may be divided into multiple tokens, and one token may also consist of several objects. and obtain their feature maps according to the bounding box:
where denotes the mask induced from the bounding box . After obtaining all the object features , we apply the across-attention to integrate relation information into the model:
Note that Eq.(13) is injected in between the gated self-attention layer and the cross-attention layer as shown in Figure 2.
3. Optimization
We adopt the pre-trained diffusion model such that layout information can be injected while all the original components remain intact. By denoting the new parameters as , we use the original denoising objective as in Eq.(1) for the model’s continual learning, based on the text prompt and layout instructions . Finally, the generation process can be optimized via:
Experiments
In this section, we carried out extensive experiments on COCO2014, the widely used benchmark dataset in vision understanding and generation, to answer the following research questions:
RQ1: How does the proposed method perform in the challenging layout planning and high-faithfulness image synthesis compared with state-of-the-art baselines? RQ2: How does each component of the proposed method affect the performance of layout generation and image synthesis? RQ3: How are the authenticity and rationality of the generated layouts and images?
We conduct experiments on COCO (Lin et al., 2014), which contains 82,783 training images and 40,504 test images over 80 semantic classes, where each image is associated with instance-wise annotations (i.e., object bounding boxes and segmentation masks) and 5 text descriptions. We split the training data into 95% for training and 5% for validation.
To thoroughly evaluate the layout planning and relation understanding abilities, we re-organize the raw test set and construct a new one. Concretely, we first pre-processed captions by means of NLP tools (Bird et al., 2009) and then select those samples which require specific layout planning capabilities. Finally, we obtain a new test set including five categories, i.e., numerical, spatial, semantic, mixed, and null. Appendix §B.2 gives more details of this part.
1.2. Evaluation Metrics
For quantitative evaluation, we employ the following metrics with respect to layout generation and image generation. 1) Layout Evaluation: Following prior work (Inoue et al., 2023; Kong et al., 2022), we adopt layout-level Fréchet Inception Distance (FID) (Heusel et al., 2017), Maximum IoU (mIoU) (Kikuchi et al., 2021), and Layout Similarity (LaySim) (Inoue et al., 2023) to assess the layout induction performance. 2) Image Evaluation: We use image-level FID, cross-modal similarities (Sim(I-T)) and intra-modal ones (Sim(I-I)) to evaluate image generation quality. Refer to the Appendix §B.3 section for more details.
1.3. Baselines
To evaluate the effectiveness of the proposed method, we compare it with the following layout generation baselines: LayoutTrans (Gupta et al., 2021) is a self-attention framework capturing the contextual relationships and generating layouts of graphical elements. BLT (Kong et al., 2022) introduce a bidirectional layout transformer to empower the transformer-based models. MaskGIT (Chang et al., 2022) propose to learn a bidirectional transformer by masked visual token prediction. VQDiffusion (Gu et al., 2022) is based on a VQ-VAE whose latent space is modeled by a conditional variant of the recently developed discrete diffusion model. LayoutDM (Inoue et al., 2023) adopt the VQDiffusion to handle the structure layout data in a discrete representation.
1.4. Implementation Details
Based on the pre-trained GLIGEN (Li et al., 2023b), we add extra relation-aware layers to model semantic relations and perform continual learning. We take the gpt-3.5-turbo model via OpenAI APIhttps://platform.openai.com/docs/models/gpt-3-5 as our LLM. Under the few-shot setting, we randomly sample 64 instances for training, and 32 instances to form the candidate set. Besides, we set 2 as the shot number by default. During the optimization phase, the total number of epochs, the batch size, and the initial learning rate are set to 80, 8, and , respectively. One can refer to Appendix §B.1 for more details.
2. Performance Comparison (RQ1)
To justify the overall effectiveness of the proposed model, we carry out extensive experiments to evaluate the quality of generated layouts and images. As shown in Table 1, we can see that the proposed method substantially outperforms the compared baselines, achieving state-of-the-art results, especially on the pair-wise relevance metrics. Next, to further assess the validity and superiority of the proposed model, we evaluate the proposed method from five aspects with respect to layout planning and image generation.
First, we evaluate the text-to-layout generation capability in terms of numerical, spatial, and semantic modeling, as shown in Table 2. The results show that the proposed method achieves the best performance under most evaluation metrics, e.g., mIoU and LaySim, substantially surpassing the compared baselines. To further explore how the proposed approach performs on complex scenes and abstract prompts, we perform another two groups of experiments, i.e., “Mixed” and “Null”. As for complex scenarios with mixed relations and abstract prompts without any explicit relations, the proposed model remarkably surpasses all the existing baselines. These results demonstrate the superiority of the proposed layout-guided text-to-image generation model.
2.2. Layout-guided Text-to-Image Generation
Based on the layouts generated by different methods, we employ the proposed layout-guided T2I generation model to synthesize images on the constructed test set of COCO 2014 in real-world scenes. From the results shown in Table 3, we have the following observations:
The auto-regressive model LayoutTrans performs worst compared with other methods in all the evaluation metrics for image generation, indicating the limitation of the traditional auto-regressive paradigm for the layout-guided image generation task.
LayoutDM, VQDiffusion, BLT, and MaskedGIT gain similar performance in text-to-image generation, and this similarity is a direct reflection of their comparable layout generation capabilities.
The proposed method exhibits a substantial performance advantage over the existing baselines, as evidenced by a remarkable improvement observed in layout induction. This outcome further suggests that LLMs possess spatial and relational reasoning capabilities, which can be effectively harnessed for the demanding task of layout-based image generation.
3. In-depth Analysis (RQ2 & RQ3)
Here we present model ablations to ascertain the efficacy of each part of the proposed method, including the feedback-based sampling strategy, the shot number of in-context examples, and the relation-aware image generation module, as elaborated subsequently.
Impact of Feedback-based Sampling. Previous studies (Wei et al., 2022; Zhang et al., 2022) have indicated that the activation of certain abilities of LLMs necessitates appropriate examples combined with corresponding questions for in-context learning. To facilitate the spatial comprehension, language-layout alignment, and layout planning abilities of LLMs, we propose the feedback-based sampling strategy. To assess its effectiveness and investigate the impact of various sampling strategies on LLMs, we design two additional variants: 1) Random Samp., wherein examples are randomly sampled from a predefined candidate set and combined with the prompt template; and 2) NN Samp. denotes that in-context examples are chosen through nearest neighbor search using textual branch-based similarities derived from CLIP (Radford et al., 2021). The experimental results on layout generation are illustrated in Figure 5(a). Compared to random sampling, NN sampling generates more accurate layouts, indicating that semantic similarities are helpful for the layout planning of LLMs. However, closeness in semantics does not mean all of that in layout, i.e., the abilities of language understanding and layout planning may be not the same, and thus different internal mechanisms of LLMs may be triggered. In contrast, the proposed feedback-based sampling strategy achieves the best performance regarding mIoU and LaySim metrics, demonstrating its effectiveness.
Impact of Shot Number. To investigate the impact of the number of in-context examples in activating the layout planning of LLMs, we conduct experiments under zero-shot and few-shot settings (2, 3, 4, and 5). As seen in Figure 5(b), we first observe that the performance of layout planning exhibits improvement with the increase of the shot number from 0 to 4, signifying that a larger number of in-context examples provide more informative clues, thereby enhancing the performance of LLMs. When reaching the 3-shot, the performance improvement seems to be saturated, after which LLMs improve slightly. Significantly, even under the zero-shot setting, LLMs demonstrate competitive performance, outperforming recent baseline models as indicated in Table 1, underscoring the generalization capability of LLMs.
Impact of Relation-aware Image Generation. In Section 4.2, we introduce the interaction-based relation-aware image generation module to enhance generation quality. To delve into how the generation can be affected by this module, we conduct the ablation study by removing it from the full framework, i.e., the original GLIGEN framework, as shown in Figure 6. The experimental results on five categories of the re-organized test set on COCO 2014 manifest that the cross-modal interaction among local relation-aware concepts contributes substantively to the relation modeling for text-to-image diffusion models guided by layout information. Particularly, considerable performance improvements are observed within “Semantic” and “Mixed” relation categories, which may be attributable to the high requirement for cross-modal semantic understanding and modeling.
3.2. Case Study
To gain an intuitive efficacy of the proposed method, we display some cases from 5 test subsets, as shown in Figure 7. By comparing different methods with the ground truth samples. We have the following discussions: 1) A layout of an image plays a key role in the generation process since the prior layout determines logic and overall semantics for the target image. 2) The layout-to-image generation has achieved impressive performance since the generated image given the ground-truth layout is comparable to the real image except for some details. 3) Although the recently proposed LayoutDM (Inoue et al., 2023) achieves promising performance on user interface and research paper layout design, it fails to generate satisfying layouts in real-world scenes. 4) Our proposed method is able to legitimately reason the distribution of objects and precisely depict their relations in the generated images, which demonstrates the effective elicitation of the layout planning capabilities from LLMs.
Conclusion
In this work, we aim to explore the cross-modal text-guided image generation problem. We find existing generative models are weak in layout planning, and propose to tackle this issue from five aspects, including numerical reasoning, spatial relation modeling, semantic relation understanding, complex layout planning, and abstract imagination. Inspired by the recent remarkable success of LLMs, we probe the above abilities via prompting and then further motivate LLMs to achieve layout planning. Concretely, we propose a feedback-based learning strategy to perform in-context learning for LLMs and a relation-aware interaction module to promote image generation. Extensive experiments on the constructed test set validate the effectiveness and superiority of the proposed model.
References
Appendix A Extended Technical Details
Each bounding box is represented as with its top-left and bottom-right coordinate quadruple. Recent work (Rahaman et al., 2019) shows that deep networks are biased toward learning lower frequency functions, resulting in performing poorly at representing high-frequency variation in coordinates. Thus, following (Mildenhall et al., 2020; Tancik et al., 2020), we encode bounding box coordinates with the Fourier embedding before feeding them to the network :
where function is applied separately to each of the four coordinate values, and we set .
A.2. The Analysis of Text Prompt
When synthesizing images conditioned on text prompts, the pivotal thing is that the model should have a comprehensive understanding of the latent intention behind the text prompt, which involves identifying the objects to be generated, their properties, and the relationships between them. Based on observations, we divide the content in the text prompt into the following components:
Objects: The specific entities or elements that need to be present in the image, such as human, animal, plant, transportation, building, etc.
Attributes: The specific properties or characteristics of the objects that need to be accurately represented in the image, such as color, size, shape, texture, or quantity.
Relationships: These describe the connections or interactions between the objects, such as spatial relationships (e.g., next to, above, below, left, inside, or contain), semantic relationships (e.g., belonging to, interacting with), or action-based relationships (e.g., holding, pushing, driving, sitting, lying, or driving).
Scene Context: This refers to the overall context or environment of the scene, including background elements, lighting, style, and other contextual factors.
Appendix B Experiment Settings
We adopt a two-stage strategy to optimize the proposed framework. In the first stage, we use a Scene Graph parser (https://github.com/vacancy/SceneGraphParser) to extract the ‘subject-predicate-object’ triplets for each caption, the maximum number of triplets is 10, and then obtain their embeddings using CLIP textual branch with the “clip-vit-large-patch14” version. Take as input the triplet embeddings and intermediate representations in the UNet of Latent Diffusion model, multi-head cross-attention layers are plugged to perform relation-aware interaction. Then, we perform continual learning based on the pre-trained GLIGEN model, with an initial learning of 3e-5, and a batch size of 1.
In the second stage, the key is to learn an optimized policy to select informative in-context examples. Concretely, we implement in Eq.(3) with a linear layer with 128 hidden neurons which is optimized to learn layout-level similarities on the top of semantic embeddings induced from the CLIP textual branch. we optimize the feedback-based sampler via Reinforcement Learning. Two-fold feedback is considered for policy gradient, including layout-level reward (implemented with mIoU) and image-level reward (consisting of image-to-text similarities, image-to-image similarities, and aesthetic scores). Considering the different numerical scales and distributions, we apply balancing factors to reweight each term: . We use the gpt-3.5-turbo model via OpenAI API (https://platform.openai.com/docs/models/gpt-3-5), considering its powerful language understanding and reasoning abilities. Under the few-shot setting, we randomly sample 64 instances for training, and 32 instances to form the candidate set. Besides, we set 2 as the shot number by default. During the optimization phase, the total number of epochs, the batch size, and the initial learning rate are set to 80, 8, and , respectively.
As for baseline methods, considering some layout generation models are unconditional or other types of conditions (e.g., partial labels and simple phrases) instead of a complete free-form natural language, we add extra cross-attention layers and train them on the COCO 2014 dataset again.
B.2. Detailed Test Set Construction
To thoroughly assess the layout planning and relation understanding abilities, we construct a new test set from the raw COCO 2014 validation set . Concretely, we build the test set in four steps:
1. Pre-define Filtering Rules. Specifically, we choose data samples whose captions include specific keywords to construct numeral, spatial subset and use the NLP toolkit spacy (https://spacy.io/models/en#en_core_web_sm) to parse captions and build semanitic subset according to Part-of-Speech (POS) tagging:
The keywords list for filtering the captions containing the numeral is: ”two”, ”three”, ”four”, ”five”, ”six”, ”seven”, ”eight”, ”nine”, ”ten”, ”many”, ”bunch”, ”some”, ”several”, ”various”, ”group”.
The keywords list for filtering the captions containing the spatial relationship is: ”left”, ”right”, ”top”, ”down”, ”near”, ”next”, ”side”, ”above”, ”inside”, ”outside”, ”below”, ”front”, ”back”, ”under”, ”around”, ”bottom”, ”up”, ”beside”, ”beneath”, ”underneath”.
Generally, a caption including notional verbs is highly possible to depict some semantic relations. To decide whether a caption contains any notional verbs, we first find words with “VERB” POS. If a word is neither detected as an auxiliary or model verb using dependency labels nor detected as a linking verb, then it is viewed as a notional verb.
2. Primary Screening. We screen the valid dataset according to the keywords, constructing three primary screening datasets, i.e., the numerical dataset (), the spatial dataset (), and the semantic dataset ().
To construct the only numerical subset that only contains the numeral in the captions, we exclude the instances from the numerical dataset that also appears in the spatial and semantic datasets, i.e., . Similarly, we build the only semantic dataset (i.e., ) and the only spatial dataset (i.e., ).
We take the intersection of numeral, spatial, and semantic relationship datasets as the Mixed dataset, i.e., .
To construct the Null datasets that do not contain any explicit relation keywords in prompts, we filter the instances included in the numeral, spatial, and semantic dataset from the total dataset, i.e., .
4. Sampling. For a dataset with more than 200 instances, we randomly select a subset of 200 instances as the final dataset. Finally, the statistics of the constructed test dataset are shown in Table 4.
B.3. Detailed Layout and Image Evaluation
For quantitative experiments, we consider various metrics from different aspects to evaluate our method on layout generation and image generation. We now introduce these metrics as follows.
Fréchet Inception Distance (FID) (Heusel et al., 2017). Following (Inoue et al., 2023), we first train a Transformer-based model that can extract discriminative layout features, which is then utilized to compute the FID.
Maximum IoU (mIoU). This score evaluates the overlap of the ground layout and predicted layout.
Layout Similarity (LaySim). LaySim proposed by (Patil et al., 2020) aims to measure the similarity between the generated layout and the given layout. Specifically, given the generated layout and gold layout , we first assign a weighted edge between any pair of bounding boxes and , indicating how similar and are in terms of shape. Then, we calculate as the final score the aggregated weight of the maximum (weighted) matching between the layouts and .
Note that mIoU and LaySim are calculated based on close-set labels, while our method generates free labels for each bounding box. To obtain mIoU and LaySim of our method, we employ CLIP textual branch to compute semantic similarities between the predicted labels and the pre-defined 80 classes, and then map each free label to the closest pre-defined one.
Fréchet Inception Distance (FID) (Heusel et al., 2017). FID measures the Fréchet Distance between the distribution of the synthetic images and real-world images in the feature space of a pre-trained Inception v3 network. Practically, we employ torch-fidelity (https://github.com/mseitzer/pytorch-fid) to calculate the FID score.
Cross-modal (Sim(I-T)) and Intra-model Similarity (Sim(I-I)). In order to perform the quantitative analysis on the semantic correctness of the text-to-image generation, we take into account the similarity scores via BLIP-2 (Li et al., 2023a). Concretely, we calculate the two types of semantic similarities between generated images and given text prompt, and gold images, respectively.
B.4. Detailed Prompts
Here, we show the detailed prompt examples used during in-context learning under 0/1/2/3-shot settings:
[Instruction]. Now you are an assistant to help me design a layout given a description. Concretely, a layout denotes a set of ”object: bounding box” items. ”object” means any object name in the world, while ”bounding box” is formulated as [x, y, w, h], where ”x, y” denotes the top left coordinate of the bounding box, ”w” denotes the width, and ”h” denotes the height. The six values ”x, y, w, h, x+w, y+h” are all larger than 0 and smaller than 1. Next, I will give you an input that describes an image, and then you should give me an output with the format ” output: object: [x, y, w, h], object: [x, y, w, h], … ”. [Test]. input: An open refrigerator with food and condiments inside of it. [Layout]. output: food: [0.1, 0.2, 0.4, 0.3], food: [0.6, 0.2, 0.3, 0.4], condiments: [0.1, 0.6, 0.4, 0.2], condiments: [0.6, 0.6, 0.3, 0.3], refrigerator: [Instruction]. Now you are an assistant to help me design a layout given a description. Concretely, a layout denotes a set of ”object: bounding box” items. ”object” means any object name in the world, while ”bounding box” is formulated as [x, y, w, h], where ”x, y” denotes the top left coordinate of the bounding box, ”w” denotes the width, and ”h” denotes the height. The six values ”x, y, w, h, x+w, y+h” are all larger than 0 and smaller than 1. Next, I will give you an input that describes an image, and then you should give me an output with the format ” output: object: [x, y, w, h], object: [x, y, w, h], … ”. [Test]. input: a person with a hat and some ski poles. [Layout]. output: person: [0.2, 0.1, 0.4, 0.8], hat: [0.3, 0.05, 0.3, 0.2], ski pole 1: [0.6, 0.3, 0.05, 0.6], ski pole 2: [0.35, 0.3, 0.05, 0.6] 1-shot prompt.
[Instruction]. Now you are an assistant to help me design a layout given a description. Concretely, a layout denotes a set of ”object: bounding box” items. ”object” means any object name in the world, while ”bounding box” is formulated as [x, y, w, h], where ”x, y” denotes the top left coordinate of the bounding box, ”w” denotes the width, and ”h” denotes the height. The six values ”x, y, w, h, x+w, y+h” are all larger than 0 and smaller than 1. Next, I will give you several examples for you to understand this task. [In-context Examples]. input: a kitchen with low lights and allot on the counters. output: knife: [0.22, 0.48, 0.02, 0.02] knife: [0.2, 0.45, 0.02, 0.02] knife: [0.22, 0.45, 0.02, 0.03] knife: [0.21, 0.47, 0.02, 0.02] sink: [0.34, 0.51, 0.42, 0.05] knife: [0.19, 0.45, 0.03, 0.03] spoon: [0.03, 0.47, 0.04, 0.04] oven: [0.01, 0.61, 0.25, 0.39] knife: [0.17, 0.45, 0.04, 0.03] knife: [0.21, 0.48, 0.02, 0.03] knife: [0.2, 0.49, 0.02, 0.02] knife: [0.19, 0.48, 0.02, 0.02] [Test]. input: An open refrigerator with food and condiments inside of it. [Layout]. output: refrigerator: [0.1, 0.1, 0.4, 0.8] milk: [0.15, 0.2, 0.1, 0.1] eggs: [0.25, 0.3, 0.1, 0.1] cheese: [0.35, 0.2, 0.1, 0.1] mayonnaise: [0.15, 0.5, 0.1, 0.1] ketchup: [0.25, 0.6, 0.1, 0.1] lettuce: [0.35, 0.5, 0.1, 0.1] [Instruction]. Now you are an assistant to help me design a layout given a description. Concretely, a layout denotes a set of ”object: bounding box” items. ”object” means any object name in the world, while ”bounding box” is formulated as [x, y, w, h], where ”x, y” denotes the top left coordinate of the bounding box, ”w” denotes the width, and ”h” denotes the height. The six values ”x, y, w, h, x+w, y+h” are all larger than 0 and smaller than 1. Next, I will give you several examples for you to understand this task. [In-context Examples]. input: A yield sign followed by a stop sign on a deserted road. output: stop sign: [0.23, 0.48, 0.05, 0.07] [Test]. input: a person with a hat and some ski poles. [Layout]. output: person: [0.4, 0.1, 0.2, 0.8] hat: [0.45, 0.05, 0.1, 0.1] ski pole 1: [0.35, 0.6, 0.05, 0.3] ski pole 2: [0.55, 0.6, 0.05, 0.3] 2-shot prompt.
[Instruction]. Now you are an assistant to help me design a layout given a description. Concretely, a layout denotes a set of ”object: bounding box” items. ”object” means any object name in the world, while ”bounding box” is formulated as [x, y, w, h], where ”x, y” denotes the top left coordinate of the bounding box, ”w” denotes the width, and ”h” denotes the height. The six values ”x, y, w, h, x+w, y+h” are all larger than 0 and smaller than 1. Next, I will give you several examples for you to understand this task. [In-context Examples]. input: Guy walking an elephant down a dirt path. output: person: [0.37, 0.43, 0.19, 0.56] elephant: [0.47, 0.0, 0.41, 0.98] input: a black cow looking over an iron fence. output: cow: [0.09, 0.23, 0.77, 0.66] cow: [0.74, 0.7, 0.24, 0.14] [Test]. input: Three zebra and four giraffes inside a fenced area. [Layout]. output: zebra1: [0.1, 0.2, 0.2, 0.6] zebra2: [0.3, 0.3, 0.2, 0.6] zebra3: [0.5, 0.2, 0.2, 0.6] giraffe1: [0.1, 0.8, 0.3, 0.2] giraffe2: [0.4, 0.8, 0.3, 0.2] giraffe3: [0.7, 0.8, 0.3, 0.2] giraffe4: [0.4, 0.6, 0.3, 0.2] fence: [0.0, 0.0, 1.0, 1.0] [Instruction]. Now you are an assistant to help me design a layout given a description. Concretely, a layout denotes a set of ”object: bounding box” items. ”object” means any object name in the world, while ”bounding box” is formulated as [x, y, w, h], where ”x, y” denotes the top left coordinate of the bounding box, ”w” denotes the width, and ”h” denotes the height. The six values ”x, y, w, h, x+w, y+h” are all larger than 0 and smaller than 1. Next, I will give you several examples for you to understand this task. [In-context Examples]. input: Guy walking an elephant down a dirt path. output: person: [0.37, 0.43, 0.19, 0.56] elephant: [0.47, 0.0, 0.41, 0.98] input: A man holding a tennis racquet on top of a tennis court. output: person: [0.5, 0.09, 0.45, 0.89] tennis racket: [0.65, 0.1, 0.17, 0.12] [Test]. input: a person with a hat and some ski poles. [Layout]. output: person: [0.35, 0.4, 0.3, 0.6] hat: [0.4, 0.35, 0.2, 0.15] ski pole 1: [0.6, 0.4, 0.05, 0.5] ski pole 2: [0.35, 0.4, 0.05, 0.5] 3-shot prompt.
Appendix C Experimental Results
Impact of In-context Example Sampling. We report the experimental results and performance comparison of Feedback Sampling (Ours), Nearest Neighbor Sampling, and Random Sampling on the full test set and five categories, as shown in Figure 8. Based on these results, we have the following discussions: 1) In general, the proposed feedback-based sampling performs better than the other two variants in most categories and evaluation metrics. It validates the effectiveness of the proposed sampling strategy. 2) The layouts generated by Random Sampling are the worst in most cases, especially in the numerical subset. The comparison results show that NN Sampling is able to provide informative in-context examples to some extent and performs better than Random Sampling. Meanwhile, compared with other categories, the numerical subset depends more heavily on the selection of in-context examples. 3) As for the image evaluation metric Sim (I-T) shown in Figure 8(c), all the three variants in the “mixed” relation category perform best, while worst in the “null” category. It may be attributable to more contributions of abundant relations in the “mixed” category to the textual faithfulness. And 4) NN Sampling performs best in the Null category according to both mIoU and LaySim metrics. The reason may be that this category does not rely heavily on layout planning abilities and semantic closeness measured by CLIP in NN Sampling is more helpful for the selection of in-context examples.
Impact of Shot Number. As shown in Figure 9, we carry out extensive experiments to explore the influence of shot numbers on the layout planning process across five test subsets. All the experiments consistently show that the layout generation performance is sensitive to the shot number, verifying the necessity of using sufficient in-context examples to activate certain abilities of LLMs. Despite this, striving for a balance between the shot number and inference cost should also be considered in practice.
Stable Diffusion vs. Ours. To compare Stable Diffusion (Rombach et al., 2022b) and our method in terms of textual faithfulness, we design ten representative prompts and based on which we run Stable Diffusion and our method to generate corresponding images, as shown in Figure 10. These examples demonstrate that the proposed layout planning and relation-aware interaction methods are able to improve the generation quality, especially in textual faithfulness.
Layout-guided Generation Baselines vs. Ours. We provide more example images synthesized by our method and baselines in Figure 11 and Figure 12. The results are consistent with Figure 7, where our method generates images with high numerical, semantic, and spatial fidelities.