CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Hui Zhang, Dexiang Hong, Yitong Wang, Jie Shao, Xinglong Wu, Zuxuan Wu, Yu-Gang Jiang

Introduction

Text-to-image (T2I) generation has been widely applied and deeply ingrained in various fields thanks to the rapid advancement of diffusion models . To achieve more controllable generation, Layout-to-image (L2I) has been proposed to generate images based on layout conditions consisting of spatial location and description of entities.

Recently, multimodal diffusion transformers (MM-DiTs) have taken text-to-image generation to the next level. These models treat text as an independent modality equally important as the image and utilize MM-Attention instead of cross-attention for interaction between modalities, thus enhancing prompt following. However, previous layout-to-image methods mainly fall into UNet-based architectures and achieve layout control by introducing extra image-layout fusion modules between the image’s self-attention and the image-text cross-attention. While enabling MM-DiT for layout-to-image generation seems straightforward, it is challenging due to the complexity of how layout is introduced, integrated, and balanced among multiple modalities. To this end, there is a pressing need to tailor a layout integration network for MM-DiTs, fully unleashing their capabilities for high-quality and precisely controllable generation.

To address this issue, we explore various network variants and ultimately propose SiamLayout. Firstly, we treat layout as an independent modality, equally important as the image and text modalities. More specifically, we employ a separate set of transformer parameters to process the layout modality. During the forward process, the layout modality interacts with other modalities via MM-Attention and maintains self-updates. Secondly, we decouple the interactions among the three modalities into two siamese branches: image-layout and image-text MM-Attentions. By independently guiding the image with text and layout and then fusing them at a later stage, we alleviate competition among modalities and strengthen the guidance from the layout.

To this end, a high-quality layout dataset composed of image-text pairs and entity annotations is crucial for training the layout-to-image model. As shown in Tab. 1, the closed-set and coarse-grained nature of existing layout datasets may limit the model’s ability to generate complex attributes (e.g. color, shape, texture). Thus, we construct an automated annotation pipeline and contribute a large-scale layout dataset derived from the SAM dataset , named LayoutSAM. It includes 2.7M image-text pairs and 10.7M entities. Each entity includes a spatial position (i.e. bounding box) and a region description. The descriptions of images and entities are fine-grained, with an average length of 95.41 tokens and 15.07 tokens, respectively. We further introduce the LayoutSAM-Eval benchmark to provide a comprehensive tool for evaluating layout-to-image generation quality.

To support diverse user inputs rather than just bounding boxes of entities, we turn a large language model into a layout planner named LayoutDesigner. This model can convert and optimize various user inputs such as center points, masks, scribbles, or even a rough idea, into a harmonious and aesthetically pleasing layout.

Through comprehensive evaluations on LayoutSAM-Eval and COCO benchmarks, SiamLayout outperforms other variants and previous SOTA models by a clear margin, especially in generating entities with complex attributes, as illustrated in Fig. 1. For layout planning, LayoutDesigner shows more comprehensive and specialized capabilities compared to baseline LLMs.

Related Work

Text-to-image generation has emerged as a promising application due to its impressive capabilities. Recently, studies such as SD3 , SD3.5 , FLUX.1 , and Playground-v3 have advanced the multimodal Diffusion Transformer architecture (MM-DiT), elevating text-to-image generation to the next level. MM-DiT significantly enhances text understanding by treating text as an independent modality, equally important as the image, and replacing traditional cross-attention with MM-Attention for modal interaction.

Layout-to-Image Generation.

To achieve more precise and controllable generation, layout-to-image generation has been proposed to generate images based on layout guidance, which includes several entities. Each entity comprises a spatial location and a region description. However, previous methods primarily focused on UNet-based architectures, and enabling MM-DiT for layout-to-image generation is challenging due to the complexity of introducing, integrating, and balancing the layout among multiple modalities. In this paper, we focus on exploring winner solutions for incorporating layout into MM-DiT, thereby fully unleashing its capabilities and achieving precisely controllable generation.

Layout Datasets.

Layout datasets typically consist of image-text pairs with entity annotations. A common type originates from COCO , featuring images with global descriptions and entities marked by bounding boxes and brief descriptions. Although some effort expands descriptions using Large Language Models (LLMs) or Vision-Language Models (VLMs), they are still limited due to the close-set nature. Ranni collects large-scale text-image pairs from LAION and WebVision , moving towards an open-set layout dataset. However, entity descriptions remain coarse-grained and lack complex attributes. In this paper, we present an annotation pipeline and introduce a large-scale layout dataset containing 2.7M image-text pairs and 10.7M detailed entity annotations.

Large Language Model for Layout Generation.

Layout generation refers to the multimodal task of creating layouts for flyers, magazines, UI interfaces, or natural images. Some studies have explored using LLMs to generate layouts based on textual descriptions, which then guide the generation of images. In this paper, we further enhance the capabilities of LLMs for generation and optimization as well as supporting user input of different granularities.

Methodology

Latent diffusion models perform the diffusion process in the latent space, which consists of a VAE , text encoders, and either a UNet-based or transformer-based noise prediction model ϵθ\epsilon_{\theta}. The VAE encoder E\mathcal{E} encodes images x\mathbf{x} into the latent space z\mathbf{z}, while the VAE decoder D\mathcal{D} reconstructs the latent back into images. The text encoders τ\bm{\tau}, such as CLIP and T5 , project tokenized text prompts into text embeddings y\mathbf{y}. The training objective is to minimize the following LDM loss:

where tt is time step uniformly sampled from {1,…,T}\{1,\ldots,T\}. The latent zt\mathbf{z}_{t} is obtained by adding noise to z0\mathbf{z}_{0}, with the noise ϵ\epsilon sampled from the standard normal distribution N(0,I)\mathbf{\mathcal{N}}(\mathbf{0},\mathbf{I}).

Multimodal Diffusion Transformer.

SD3/3.5 , FLUX.1 , and Playground-v3 instantiate noise prediction using MM-DiT, which uses two independent transformers to handle text and image embeddings separately. Unlike previous diffusion models that process different modalities through cross-attention, MM-DiT concatenates the embeddings of the image and text for the self-attention operation, referred to as MM-Attention:

MM-DiT treats image and text as equally important modalities to improve prompt following . In this paper, we explore incorporating layout into MM-DiTs, unleashing their potential for high-quality and precise L2I generation.

2 Layout-to-Image Generation

Layout-to-image generation aims at precise and controllable image generation based on the instruction I\bm{I}, which consists of a global-wise prompt condition p\bm{p} and a region-wise layout condition l\bm{l}, denoted as:

The layout condition includes information for NN entities e\bm{e}, each consisting of two parts: region caption c\bm{c} and spatial location b\bm{b}, denoted as:

In this work, we use bounding boxes to represent spatial locations, consisting of the coordinates of the top-left and bottom-right corners.

Tokenize Different Modalities.

Image tokens hz\bm{h}^{z} are derived by patchifying the latent z\mathbf{z}, and tokens hp\bm{h}^{p} of global caption p\bm{p} are obtained from the text encoder τ\bm{\tau}, denoted as hp=τ(p)\bm{h}^{p}=\bm{\tau}(\bm{p}). We denote the layout tokens as hl=[h1l,⋯ ,hNl]\bm{h}^{l}=[h_{1}^{l},\cdots,h_{N}^{l}]. Inspired by GLIGEN , each hilh^{l}_{i} is obtained from the layout encoder in Fig. 2:

Fourier refers to the Fourier embedding , [·, ·] denotes concatenation across the feature dimension, and MLP is a multi-layer perception.

Layout Integration.

MM-DiT allows the two modalities to interact through the following MM-Attention:

where [·, ·] denotes concatenation across the tokens dimension, Qz=hzWqz\mathbf{Q}^{z}=\bm{h}^{z}\mathbf{W}_{q}^{z}, Kp=hzWkz\mathbf{K}^{p}=\bm{h}^{z}\mathbf{W}_{k}^{z}, Vz=hzWvz\mathbf{V}^{z}=\bm{h}^{z}\mathbf{W}_{v}^{z}; and Wqz\mathbf{W}_{q}^{z}, Wkz\mathbf{W}_{k}^{z}, Wvz\mathbf{W}_{v}^{z} are the weight matrices for the query, key, and value linear projection layers for image tokens, respectively. Tokens hp\bm{h}^{p} of the global caption p\bm{p} are handled by the same paradigm as hz\bm{h}^{z} but with their own weights. hz′{\bm{h}^{z}}^{\prime} and hp′{\bm{h}^{p}}^{\prime} are the image and caption tokens after interaction. To this end, the next critical step is to incorporate the layout tokens hl\bm{h}^{l}. We explore three variants of network designs that incorporate layout tokens, as shown in Fig. 2.

Layout Adapter. Based on the fundamental idea of previous L2I methods introducing layout conditions in UNet-based architectures, we design extra image-layout cross-attention to incorporate layout into MM-DiT, defined as the Layout Adapter. Formally, it is represented as:

In this paper, SiamLayout is chosen as the network that incorporates layout into MM-DiT, and our primary experiments are conducted on it.

Training and Inference.

We freeze the pre-trained model and only train the newly introduced parameters θ′{\theta}^{\prime} using the following loss function:

Here, we employ two strategies to accelerate the convergence of the model: \@slowromancapi@) Biased sampling of time steps: Since layout pertains to the structural content of images, which is primarily generated during the larger time steps, we sample time steps with a 70% probability from a normal distribution N(0.7∗T,T)\mathcal{N}(0.7*T,T) and with a 30% probability from N(0,T)\mathcal{N}(0,T). \@slowromancapii@) Region-aware loss: We enhance the model’s focus on areas specified by the layout by assigning greater weight to the region loss Lregion\mathcal{L}_{region} associated with the regions localized by the bounding boxes in the latent space. The updated loss is:

where λregion{\lambda}_{region} modulates the importance of Lregion\mathcal{L}_{region}. During the inference phase, we perform layout-conditioned denoising only in the first 30% of the steps .

3 Layout Dataset and Benchmark

As there is no large-scale and fine-grained layout dataset explicitly designed for layout-to-image generation, we collect 2.7 million image-text pairs with 10.7 million regional spatial-caption pairs derived from the SAM Dataset , named LayoutSAM. We design automatic schemes and strict filtering rules to annotate layout and clean noisy data, with the following five parts:

i@) Image Filtering: We employ the LAION-Aesthetics predictor to curate a high visual quality subset from SAM, selecting images in the top 50% of aesthetic scores.

ii@) Global Caption Annotation: As the SAM dataset does not provide descriptions for each image, we generate detailed descriptions using a Vision-Language Model (VLM) . The average length of the captions is 95.41 tokens.

iii@) Entity Extraction: Existing SoTA open-set grounding models prefer to detect entities through a list of short phrases rather than dense captions. Thus, we utilize a Large Language Model to derive brief descriptions of main entities from dense captions via in-context learning. The average length of the brief descriptions is 2.08 tokens.

iv@) Entity Spatial Annotation: We use Grounding DINO to annotate bounding boxes of entities and design filtering rules to clean noisy data. Following previous work , we first filter out bounding boxes that occupy less than 2% of the total image area, then only retain images with 3 to 10 bounding boxes.

v@) Region Caption Recaptioning: At this point, we have the spatial locations and brief captions for each entity. We use a VLM to generate fine-grained descriptions with complex attributes for each entity based on its visual content and brief description. The average length of these detailed descriptions is 15.07 tokens.

Layout-to-Image Benchmark.

The LayouSAM-Eval benchmark serves as a comprehensive tool for evaluating L2I generation quality collected from a subset of LayouSAM. It comprises a total of 5,000 layout data. We evaluate L2I generation quality using LayouSAM-Eval from two aspects:

Region-wise quality. This aspect is evaluated for adherence to spatial and attribute accuracy via VLM’s Visual Question Answering (VQA). For each entity, spatially, the VLM evaluates whether the entity exists within the bounding box; for attributes, the VLM assesses whether the entity matches the color, text, and shape mentioned in the detailed descriptions.

Global-wise quality. This aspect scores based on visual quality and global caption following, across multiple metrics including recently proposed scoring models like IR score and Pick score , as well as traditional metrics such as CLIP , FID and IS scores. For more details on the proposed dataset and benchmark, please refer to the supplementary materials.

4 Layout Designer

Experiments

Text-to-Image Generation.

We conduct experiments on the T2I-CompBench and evaluate image quality from five aspects: spatial, color, shape, texture, and numeracy.

Layout Generation and Optimization.

We construct 180,000 training sets based on the LayoutSAM training set for three types of user input and layout pairs: caption-layout pairs, center point-layout pairs, and suboptimal layout-layout pairs. Similarly, we construct 1,000 validation sets for each of these tasks from LayoutSAM-Eval.

Implementation Details.

We employ experiments on SD3-medium, and the extra parameters introduced by SiamLayout amount to 1.28B. The training resolution for the LayoutSAM dataset is 1024×10241024\times 1024, and 512×512512\times 512 for COCO. We utilize the AdamW optimizer with a fixed learning rate of 5e-5 and train the model for 600,000 iterations with a batch size of 16. We train SiamLayout with 8 A800-40G GPUs for 7 days. The value of λregion\lambda_{region} is set to 2. LayoutDesigner is fine-tuned on Llama-3.1-8B-Instruct for one day using one A800-40G GPU.

2 Evaluation on Layout-to-Image Generation

Tab. 2 presents the quantitative results of SiamLayout on the fine-grained open-set LayoutSAM-Eval, including metrics of region-wise quality and global-wise quality. SiamLayout not only surpasses the current SOTA in terms of spatial response but also exhibits more precise responses in attributes such as color, texture, and shape. By fully unleashing the power of MM-DiT, SiamLayout also demonstrates a dominant advantage in overall image quality. This is further confirmed by the qualitative results in Fig. 5, showing that SiamLayout achieves more accurate and aesthetically appealing attribute rendering in the regions localized by the bounding boxes, including the rendering of shapes, colors, textures, text, and portraits.

Coarse-Grained Closed-set L2I.

We train and evaluate SiamLayout on COCO to confirm its generalization in coarse-grained closed-set layout-to-image generation, as shown in Tab. 3. In terms of image quality, SiamLayout outperforms previous methods by a clear margin on CLIP, FID, and IS metrics, thanks to the tailored framework that unleashes the capabilities of MM-DiT. In terms of spatial positioning response, SiamLayout is slightly inferior to InstanceDiff. We attribute this to two factors: \@slowromancapi@) The training dataset of InstanceDiff is a more fine-grained COCO dataset, with per-entity fine-grained attribute annotations; \@slowromancapii@) InstanceDiff generates each entity separately and then combines them, achieving more precise control at the cost of increased time and computational resources.

3 Evaluation on Text-to-Image Generation

To further validate the impact of incorporating layout on text-to-image generation, we conduct experiments on the T2I-CompBench . For the prompts, we first use LLM to plan the layout, which is then used to generate images via SiamLayout. Tab. 4 reveals that, by introducing layout to provide further guidance signals for image generation, SD3 has seen a significant improvement in spatial adherence (from 32.00 to 47.36). Additionally, the benefits of layout are also reflected in the improved adherence to prompts regarding color, shape, texture, and the number of objects.

4 Evaluation on Layout Planning

To validate the capability of layout generation and optimization, we compare LayoutDesigner fine-tuned on Llama3.1 with the latest LLMs, as presented in Tab. 5. Accuracy measures the correctness of the generated bounding boxes, including ensuring that the coordinates of the top-left corner are less than those of the bottom-right corner and that the bounding box does not exceed the image boundaries. Quality refers to the IR score of the image generated according to the planned layout, which is used to reflect the rationality and harmony of the layout. In terms of format accuracy, LayoutDesigner shows significant improvement over the vanilla Llama and clearly outperforms the previous SOTA. Additionally, as the LayoutDesigner contributes more aesthetically pleasing layouts, the generated images possess higher quality. Fig. 7 further confirms this, showing that layouts generated by Llama often fail to meet formatting standards or miss key elements, while those generated by GPT4-Turbo often violate fundamental physical laws (e.g. overly small objects). In contrast, images generated from layouts designed by LayoutDesigner exhibit better quality as the layouts are more harmonious and aesthetically pleasing.

5 Ablation Study

The Impact of Training Strategies.

In Fig. 8, we explore the impact of two strategies introduced during training on SiamLayout: biased time step sampling and region-aware loss. With the region-aware loss Lregion\mathcal{L}_{region}, the model focuses more on the areas localized by the layout, accelerating the model’s convergence. In addition, layout predominantly guides structural content, which is mainly generated at larger time steps. Thus, sampling larger time steps with a higher probability (i.e. biased time step sampling) also effectively speeds up the model’s convergence.

Conclusion

We presented SiamLayout, which treats layout as an independent modality and guides image generation through an image-layout branch that is siamese to the image-text branch. SiamLayout unleashes the power of MM-DiT to achieve high-quality, accurate layout-to-image generation, and significantly outperforms previous work in generating complex attributes such as color, shape, and text. Additionally, we introduced LayoutSAM, a large-scale layout dataset with 2.7M image-text pairs and 10.7M entities, along with LayoutSAM-Eval to assess generation quality. Finally, we proposed LayoutDesigner, which tames a large language model into a professional layout planner capable of handling user inputs of varying granularity.

We introduce a large language model for layout planning, which brings extra computation costs. Integrating layout planning with layout-to-image generation into an end-to-end model is an important direction for future research. Additionally, the automatic annotation pipeline introduces noisy data mainly caused by the object detection model, so its impact on the performance of the layout-to-image model needs to be further studied.

References

A More details on Datasets and Benchmarks

We design a mechanism to automatically annotate the layout for any given image, as shown in Fig. 9.

i@) Image Filtering: We employ the LAION-Aesthetics predictor to assign aesthetic scores to images and filter out those with low scores. For SAM , we analyze the aesthetic scores shown in Fig. 10 and curate a high visual quality subset consisting of images in the top 50% of aesthetic scores.

ii@) Global Caption Annotation: We generate the detailed descriptions for the query image using the Qwen-VL-Chat-Dense-Captioner , which is a vision-language model fine-tuned on generative and human-annotated data using LoRA. This model supports accurate and detailed image descriptions. The average length of the captions is 95.41 tokens, measured by the CLIP tokenizer.

iii@) Entity Extraction: Existing state-of-the-art open-set grounding models perform better at detecting entities using a list of short phrases than directly using dense captions. Thus, we utilize the large language model Llama3.1-8b-it to extract the main entities from dense captions via in-context learning. The brief descriptions include simple attribute descriptions with an average length of 2.08 tokens.

iv@) Entity Spatial Annotation: We use Grounding DINO to annotate bounding boxes of entities. To clean noisy data, we design the following filtering rules. We first filter out bounding boxes that occupy less than 2% of the total image area, then only retain images with 3 to 10 bounding boxes. The average number of entities per image is 3.96.

v@) Region Caption Recaptioning: We use the vision language model MiniCPM-V-2.6 to generate fine-grained descriptions with complex attributes for each entity based on its visual content and brief description. The generated detailed descriptions generally cover attributes such as color, shape, texture, and some text details, with an average length of 15.07 tokens.

Finally, we contribute the large-scale layout dataset LayoutSAM, which includes 2.7 million image-text pairs and 10.7 million entities. Each entity is annotated with a bounding box and a detailed description. Fig. 11 shows some examples from LayoutSAM.

LayoutSAM-Eval Benchmark.

The LayoutSAM-Eval benchmark, constructed from LayoutSAM, serves as a comprehensive tool for evaluating layout-to-image generation quality. It consists of 5,000 layout data points. We evaluate layout-to-image generation quality using LayoutSAM-Eval from two aspects: Spatial and Attribute, both of which are evaluated via the vision language model MiniCPM-V-2.6 in a visual question-answering manner.

Spatial Accuracy. To measure spatial adherence, for each bounding box, we ask the VLM whether the given entity exists within the bounding box, with the answer being either “Yes” or “No.” Finally, we divide the number of entities with a “Yes” answer by the total number of entities to obtain the spatial score.

Attribute Accuracy. To measure attribute adherence, we ask the VLM whether the entity within the bounding box matches the attributes in the detailed description. For attributes like color, shape, and texture, each attribute is evaluated independently through visual question answering, and the score is obtained in the same manner as the spatial score.

A.2 Layout Planning Dataset and Benchmark

To train LayoutDesigner, we construct a layout planning dataset derived from LayoutSAM. It consists of a total of 180,000 data points, covering the following three tasks, each with 60,000 data points:

Caption-to-layout generation. We randomly select data from LayoutSAM to construct pairs of global captions and ground truth layouts of entities. Each entity includes a bounding box and a description. This portion of the data is used to train the generation of layouts based on global captions.

Center point-to-layout generation. For each entity, we calculate the center point from its bounding box to construct pairs of center points and ground truth layouts of entities. This portion of the data is used to train the generation of layouts based on the center points of entities.

Suboptimal layout-to-layout optimization. For each entity, we create suboptimal layouts by performing operations such as deletion, duplication, movement, and resizing with a certain probability. This portion of the data consists of pairs of suboptimal layouts and GT layouts and is used to train the optimization of layouts from suboptimal to better.

LayoutDesigner is a unified layout planning model that supports all three tasks simultaneously through joint training of a large language model on these three types of data.

Layout Planning Benchmark.

We construct 1,000 data points for each layout planning task using the same method from LayoutSAM-Eval and conduct experiments to evaluate layout planning capabilities under a 3-shot in-context learning setting. First, we evaluate the formatting accuracy of the generated layouts, including the coordinates of the top-left corner being smaller than those of the bottom-right corner and the bounding box not exceeding the image boundaries. To assess the harmony and aesthetics of the layouts, we did not use metrics like AP to measure the adherence of the generated bounding boxes to the ground truth. This is because layout generation is an open-ended problem, and there can be multiple optimal solutions for the input. Even if a solution does not resemble the GT, it can still be an excellent layout. Therefore, we generate images based on the designed layouts and reflect the quality of the layouts through the quality of the images. Additionally, we evaluate the layout planning capabilities through qualitative results.

B More analysis on modal competition

C More qualitative results

We present more qualitative results in Fig. 13. Experimental results show that our proposed method empowers MM-DiT for layout-to-image generation, achieving visually appealing and precisely controllable generation, as demonstrated by the high adherence to complex attributes such as color, texture, and shape.