LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation

Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, Xi Li

Introduction

Recently, the diffusion model has achieved encouraging progress in conditional image generation, especially in text-to-image generation such as GLIDE , Imagen , and Stable Diffusion . However, text-guided diffusion models may still fail in the following situations. As shown in Fig. 1 (a), when aiming to generate a complex image with multiple objects, it is hard to design a prompt properly and comprehensively. Even input with well-designed prompts, problems such as missing objects and incorrectly generating objects’ positions, shapes, and categories still occur in the state-of-the-art text-guided diffusion model . This is mainly due to the ambiguity of the text and its weakness in precisely expressing the position of the image space. Fortunately, this is not a problem when using the coarse layout as guidance, which is a set of objects with the annotation of the bounding box (bbox) and object category. With both spatial and high-level semantic information, the diffusion model can obtain more powerful controllability while maintaining the high quality.

However, early studies on layout-to-image generation are almost limited to generative adversarial networks (GANs) and often suffer from unstable convergence and mode collapse . Despite the advantages of diffusion models in easy training and significant quality improvement , few studies have considered applying diffusion in the layout-to-image generation task. To our knowledge, only LDM supports the condition of layout and has shown encouraging progress in this field.

In this paper, different from LDM that applies the simple multimodal fusion method (e.g., the cross attention) or direct input concatenation for all conditional input, we aim to specifically design the fusion mechanism between layout and image. Moreover, instead of conditioning only in the second stage like LDM, we propose an end-to-end one-stage model that considers the condition for the whole process, which may have the potential to help mitigate loss in the task that requires fine-grained accuracy in pixel space . The fusion between image and layout is a difficult multimodal fusion problem. Compared to the fusion of text and image, the layout has more restrictions on the position, size, and category of objects. This requires a higher controllability of the model and often leads to a decrease in the naturalness and diversity of the generated image. Furthermore, the layout is more sensitive to each token and the loss in token of layout will directly lead to the missing objects.

To address the problems mentioned above, we propose treating the patched image and the input layout in a unified form. Specifically, we construct a structural image patch at multi-resolution by adding the concept of region that contains information of position and size. As a result, each patch of the image is transformed into a special type of object, and the entire patched image will also be regarded as a layout. Finally, the difficult problem of multimodal fusion between image and layout will be transformed into a simple fusion with a unified form in the same spatial space of the image. We name our model LayoutDiffuison, a layout-conditional diffusion model with Layout Fusion Module (LFM), object-aware Cross Attention Mechanism (OaCA), and corresponding classifier-free training and sampling scheme. In detail, LFM fuses the information of each object and models the relationship among multiple objects, providing a latent representation of the entire layout. To make the model pay more attention to the information related to the object, we propose an object-aware fusion module named OaCA. Cross-attention is made between the image patch feature and layout in a unified coordinate space by representing the positions of both of them as bounding boxes. To further improve the user experience of LayoutDiffuison, we also make several optimizations on the speed of the classifier-free sampling process and could significantly outperform the SOTA models in 25 iterations.

Experiments are conducted on COCO-stuff and Visual Genome (VG) . Various metrics ranging from quality, diversity, and controllability show that LayoutDiffusion significantly outperforms both state-of-the-art GAN-based and diffusion-based methods.

Instead of using the dominated GAN-based methods, we propose a diffusion model named LayoutDiffusion for layout-to-image generations, which can generate images with both high-quality and diversity while maintaining precise control over the position and size of multiple objects.

We propose to treat each patch of the image as a special object and accomplish the difficult multimodal fusion of layout and image in a unified form. LFM and OaCA are then proposed to fuse the multi-resolution image patches with user’s input layout.

LayoutDiffuison outperforms the SOTA layout-to-image generation method on FID, DS, CAS by relatively around 46.35%\%, 9.61%\%, 26.70%\% on COCO-stuff and 44.29%\%, 11.30%\%, 41.82%\% on VG.

Related work

The related works are mainly from layout-to-image generation and diffusion models.

Layout-to-Image Generation. Before the layout-to-image generation is formally proposed, the layout is usually used as as a complementary feature or an intermediate representation in text-to-image , scene-to-image generation . The first image generation directly from the layout appears in Layout2Im and is defined as a set of objects annotated with category and bbox. Models that work well with fine-grained semantic maps at the pixel level can also be easily transformed to this setting . Inspired by StyleGAN , LostGAN-v1 , LostGAN-v2 used a reconfigurable layout to obtain better control over individual objects. For interactive image synthesis, PLGAN employed panoptic theory by constructing stuff and instance layouts into separate branches and proposed Instance- and Stuff-Aware Normalization to fuse into panoptic layouts. Despite encouraging progress in this field, almost all approaches are limited to the generative adversarial network (GAN) and may suffer from unstable convergence and mode collapse . As a multimodal diffusion model, LDM supports the condition of coarse layout and has shown great potential in layout-guided image generation.

Diffusion models are being recognized as a promising family of generative models that have proven to be state-of-the-art sample quality for a variety of image generation benchmarks , including class-conditional image generation , text-to-image generation , and image-to-image translation . Classifier guidance was introduced in ADM-G to allow diffusion models to condition the class label. The gradient of the classifier trained on noised images could be added to the image during the sampling process. Then Ho et al. proposed a classifier-free training and sampling strategy by interpolating between predictions of a diffusion model with and without condition input. For the acceleration of training and sampling speed, LDM proposed to first compress the image into smaller resolution and then apply denoising training in the latent space.

Method

In this section, we propose our LayoutDiffusion, as shown in Fig. 2. The whole framework consists mainly of four parts: (a) layout embedding that preprocesses the layout input, (b) layout fusion module that encourages more interaction between objects of layout, (c) image-layout fusion module that constructs the structal image patch and object-aware cross attention developed with the specific design for layout and image fusion, (d) the layout-conditional diffusion model with training and accelerated sampling methods.

A layout l ⁣= ⁣{o1,o2,⋯ ,on}l\!=\!\{o_{1},o_{2},\cdots,o_{n}\} is a set of nn objects. Each object oio_{i} is represented as oi ⁣= ⁣{bi,ci}o_{i}\!=\!\{b_{i},c_{i}\}, where bi ⁣= ⁣(x0i,y0i,x1i,y1i) ⁣∈ ⁣4b_{i}\!=\!(x_{0}^{i},y_{0}^{i},x_{1}^{i},y_{1}^{i})\!\in\!^{4} denotes a bounding box (bbox) and ci∈[0,C+1]c_{i}\in[0,\mathcal{C}+1] is its category id.

To support the input of a variable length sequence, we need to pad ll to a fixed length kk by adding one olo_{l} in the front and some padding opo_{p} in the end, where olo_{l} represents the entire layout and opo_{p} represents no object. Specifically, bl ⁣= ⁣(0,0,1,1)b_{l}\!=\!(0,0,1,1), cl ⁣= ⁣0c_{l}\!=\!0 denotes a object that covers the whole image and bp ⁣= ⁣(0,0,0,0)b_{p}\!=\!(0,0,0,0), cp ⁣= ⁣C+1c_{p}\!=\!\mathcal{C}+1 denotes a empty object that has no shape or does not appear in the image.

2 Layout Fusion Module

Currently, each object in layout has no relationship with other objects. This leads to a low understanding of the whole scene, especially when multiple objects overlap and block each other. Therefore, to encourage more interaction between multiple objects of the layout to better understand the entire layout before inputting the layout embedding, we propose Layout Fusion Module (LFM), a transformer encoder that uses multiple layers of self-attention to fuse the layout embedding and can be denoted as

3 Image-Layout Fusion Module

Structural Image Patch. The fusion of image and layout is a difficult multimodal fusion problem, and one of the most important parts lies in the fusion of position and size. However, the image patch is limited to the semantic information of the whole feature and lacks the spatial information. Therefore, we construct a structural image patch by adding the concept of region that contains the information of position and size.

The bounding box sets of a patched image II is defined as bI ⁣= ⁣{bIu,v∣u∈[0,h),v∈[0,w)}b_{\mathcal{I}}\!=\!\{b_{\mathcal{I}_{u,v}}|u\in[0,h),v\in[0,w)\}. As a result, the positional information of image patch and layout object is contained in the unified bounding box defined in the same spatial space, leading to better fusion of image and layout.

Positional Embedding in Unified Space. We define the positional embedding of the image and layout as PIP_{\mathcal{I}} and PLP_{\mathcal{L}} as follows:

With the help of LFM in Eq. 4, O1′O_{1}^{\prime} can be considered as a global information of the entire layout, and Oi′(i∈[2,k])O_{i}^{\prime}(i\in[2,k]) is considered as the local information embedding of single object along with the other related objects. One of the easiest ways to condition the layout in the image is to directly add O1′O_{1}^{\prime}, the global information of the layout, to the multiple resolution of image features. Specifically, the condition process can be defined as

Object-aware Cross Attention for Local Conditioning.

Cross attention is successfully applied in to condition text into image feature, where the sequence of the image patch is used as the query and the concatenated sequence of the image patch and text is applied as key and value. The equation of cross-attention is defined as

In text-to-image generation, each token in the text sequence is a word. The aggregation of these words constitutes the semantics of a sentence. After the transformer encoder, the first token in text sequence is well-semantic information that generalizes the whole text but may not reverse the semantic meaning of each word. However, the loss in information of one token is relatively serious in layout rather than in text. Each token in the layout sequence is a single object with a specific category, size, and position. The loss of information on a layout token will directly lead to a missing or wrong object in the generated image pixel space.

Therefore, we take into account the fusion of locations, size, and category of objects and define our object-aware cross-attention (OaCA) as

We first construct the key and value of the layout:

We construct the query, key, and value of the image feature as follows:

4 Layout-conditional Diffusion Model

Here, we follow the Gaussian diffusion models improved by . Given a data point sampled from a real data distribution x0∼q(x0)x_{0}\sim q(x_{0}), a forward diffusion process is defined by adding small amount of Gaussian noise to the x0x_{0} in TT steps:

If the total noise added throughout the Markov chain is large enough, the xTx_{T} will be well approximated by N(0,I)\mathcal{N}(0,\mathbf{I}). If we add noise at each step with a sufficiently small magnitude 1−αt1-\alpha_{t}, the posterior q(xt−1∣xt)q(x_{t-1}|x_{t}) will be well approximated by a diagonal Gaussian. This nice property ensures that we can reverse the above forward process and sample from xT∼N(0,I)x_{T}\sim\mathcal{N}(0,\mathbf{I}), which is a Gaussian noise. However, since the entire dataset is needed, we are unable to easily estimate the posterior. Instead, we have to learn a model pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) to approximate it:

Instead of using the tractable variational lower bound (VLB) in log⁡pθ(x0)\log p_{\theta}(x_{0}), Ho et al. proposed to reweight the terms of the VLB to optimize a surrogate objective. Specifically, we first add tt steps of Gaussian noise to a clean sample x0x_{0} to generate a noised sample xt∼q(xt∣x0)x_{t}\sim q(x_{t}|x_{0}). Then train a model ϵθ\epsilon_{\theta} to predict the added noise using the following loss:

, which is a standard mean-squared error loss.

To support the layout condition, we apply classifier-free guidance, a technique proposed by Ho et al. for conditional generation that requires no additional training of the classifier. It is accomplished by interpolating between predictions of a diffusion model with and without condition input. For the condition of layout, we first construct a padding layout lϕ={ol,op,⋯ ,op}l_{\phi}=\{o_{l},o_{p},\cdots,o_{p}\}. During training, the condition of layout ll of diffusion model will be replaced with lϕl_{\phi} with a fixed probability. When sampling, the following equation is used to sample a layout-condional image:

, where the scale ss can be used to increase the gap between ϵθ(xt,t∣lϕ)\epsilon_{\theta}(x_{t},t|l_{\phi}) and ϵθ(xt,t∣l)\epsilon_{\theta}(x_{t},t|l) to enhance the strength of conditional guidance.

To further improve the user experience of LayoutDiffuison, we also make several optimizations on the speed of the classifier-free sampling process and could significantly outperform the SOTA models in 25 iterations. Specifically, we adapt DPM-solver for the conditional classifier-free sampling, a fast dedicated high-order solver for diffusion ODEs with the convergence order guarantee, to accelerate the conditional sampling speed.

Experiments

In this section, we evaluate our LayoutDiffusion on different benchmarks in terms of various metrics. First, we introduce the datasets and evaluation metrics. Second, we show the qualitative and quantitative results compared with other strategies. Finally, some ablation studies and analysis are also mentioned. More details can be found in Appendix, including model architecture, training hyperparameters, reproduction results, more experimental results and visualizations.

We conduct our experiments on two popular datasets, COCO-Stuff and Visual Genome .

COCO-Stuff has 164K images from COCO 2017, of which the images contain bounding boxes and pixel-level segmentation masks for 80 categories of thing and 91 categories of stuff, respectively. Following the settings of LostGAN-v2 , we use the COCO 2017 Stuff Segmentation Challenge subset that contains 40K / 5k / 5k images for train / val / test-dev set, respectively. We use images in the train and val set with 3 to 8 objects that cover more than 2%2\% of the image and not belong to crowd. Finally, there are 25,210 train and 3,097 val images.

Visual Genome collects 108,077 images with dense annotations of objects, attributes, and relationships. Following the setting of SG2Im , we divide the data into 80%\%, 10%\%, 10%\% for the train, val, test set, respectively. We select the object and relationship categories occurring at least 2000 and 500 times in the train set, respectively, and select the images with 3 to 30 bounding boxes and ignoring all small objects. Finally, the training / validation / test set will have 62565 / 5062 / 5096 images, respectively.

2 Evaluation Metrics & Protocols

We use five metrics to evaluate the quality, diversity, and controllability of generation.

Fr‘echet Inception Distance (FID) shows the overall visual quality of the generated image by measuring the difference in the distribution of features between the real images and the generated images on an ImageNet-pretrained Inception-V3 network.

Inception Score (IS) uses an Inception-V3 pretrained on ImageNet network to compute the statistical score of the output of the generated images.

Diversity Score (DS) calculates the diversity between two generated images of the same layout by comparing the LPIPS metric in a DNN feature space between them.

Classification Score (CAS) first crops the ground truth box area of images and resizing them at a resolution of 32×\times32 with their class. A ResNet-101 classifier is trained with generated images and tested on real images.

YOLOScore evaluates 80 thing categories bbox mAP on generated images using a pretrained YOLOv4 model, and shows the precision of control in one generated model.

In summary, FID and IS show the generation quality, DS shows the diversity, CAS and YOLOScore represent the controllability. We follow the architecture of ADM , which is mainly a UNet. All experiments are conducted on 32 NVIDIA 3090s with mixed precision training . We set batch size 24, learning rate 1e-5. We adopt the fixed linear variance schedule. More details can be found in the Appendix.

3 Qualitative results

Comparison of generated 256 ×\times 256 images on the COCO-Stuff with our method and previous works is shown in Fig. 3.

LayoutDiffusion generates more accurate high quality images, which has more recognizable and accurate objects corresponding to their layouts. Grid2Im , LostGAN-v2 and PLGAN generate images with distorted and unreal objects.

Especially when input a set of multiple objects with complex relationships, previous work can hardly generate recognizable objects in the position corresponding to layouts. For example, in Fig. 3 (a), (c), and (e), the main objects (e.g. train, zebra, bus) in images are poorly generated in previous work, while our LayoutDiffusion generates well. In Fig. 3 (b), only our LayoutDiffusion generates the laptop in the right place. The images generated by our LayoutDiffusion are more sensorially similar to the real ones.

We show the diversity of LayoutDiffusion in Fig. 4. Images from the same layouts have high quality and diversity (different lighting, textures, colors, and details).

We continuously add an additional layout from the initial layout, the one in the upper left corner, as shown in Fig. 5. In each step, LayoutDiffusion adds the new object in very precise locations with consistent image quality, showing user-friendly interactivity.

4 Quantitative results

Tab. 1 provides the comparison among previous works and our method in FID, IS, DS, CAS and YOLOScore. Compared to the SOTA method, the proposed method achieves the best performance in comparison.

In overall generation quality, our LayoutDiffusion outperforms the SOTA model by 46.35% and 29.29% at most in FID and IS, respectively. While maintaining high overall image quality, we also show precise and accurate controllability, LayoutDiffusion outperforms the SOTA model by 122.22% and 41.82% at most on YOLOScore and CAS, respectively. As for diversity, our LayoutDiffusion still achieves 11.30% imporvement at most accroding to the DS. Experiments on these metrics show that our methods can successfully generate the higher-quality images with better location and quantity control.

In particular, we conduct experiments compared to LDM in Tab. 3. “Ours-small” uses comparable GPU resources to have better FID performance with much fewer parameters and better throughout compared to LDM-8 when “Ours-small” outperforms LDM-4 in all respects. The results of “Ours” indicate that LayoutDiffusion can have better FID performance, 31.6, at a higher cost. From these results, LayoutDiffusion always achieves better performance at different cost levels compared with LDM .

5 Ablation studies

We validate the effectiveness of LFM and OaCA in Tab. 2, using the evaluation metrics in Sec. 4.2. The significant improvement on FID, IS, CAS, and YOLOScore proves that the application of LFM and OaCA allows for higher generation quality and diversity, along with more controllability. Furthermore, when applying both, considerable performance, 13.37 / 6.58 / 39.77 / 27.00 on FID / IS / CAS / YOLOScore, is gained.

An interesting phenomenon is that the change of the Diversity Score (DS) is in the opposite direction of other metrics. This is because DS, which stands for diversity, is physically the opposite of the controllability represented by other metrics such as CAS and YOLOScore. The precise control offered on generated image leads to more constraints on diversity. As a result, the Diversity Score (DS) has a slight drop compared to the baseline.

Limitations & Societal Impacts

Limitations. Despite the significant improvements in various metrics, it is still difficult to generate a realistic image with no distortion and overlap, especially for a complex multi-object layout. Moreover, the model is trained from scratch in the specific dataset that requires detection labels. How to combine text-guided diffusion models and inherit parameters pre-trained on massive text-image datasets remains a future research.

Societal Impacts. Trained on the real-world datasets such as COCO and VG , LayoutDiffusion has the powerful ability to learn the distribution of data and we should pay attention to some potential copyright infringement issues.

Conclusion

In this paper, we have proposed a one-stage end-to-end diffusion model named LayoutDiffuison, which is novel for the task of layout-to-image generation. With the guidance of layout, the diffusion model allows more control over the individual objects while maintaining higher quality than the prevailing GAN-based methods. By constructing a structural image patch with region information, we regrad each patch as a special object and accomplish the difficult multimodal image-layout fusion in a unified form. Specifically, Layout Fusion Module and Object-aware Cross Attention are proposed to model the relationship among multiple objects and fuse the patched image feature with layout at multiple resolutions, respectively. Experiments in challenging COCO-stuff and Visual Genome (VG) show that our proposed method significantly outperforms both state-of-the-art GAN-based and diffusion-based methods in various evaluation metrics.

Acknowledgements. This work is supported in part by National Natural Science Foundation of China under Grant U20A20222, National Science Foundation for Distinguished Young Scholars under Grant 62225605, National Key Research and Development Program of China under Grant 2020AAA0107400, Research Fund of ARC Lab, Tencent PCG, Zhejiang – Singapore Innovation and AI Joint Research Lab, Ant Group through CCF-Ant Research Fund, and sponsored by CCF-AFSG Research Fund, CAAI-HUAWEI MindSpore Open Fund as well as CCF-Zhipu AI Large Model Fund(CCF-Zhipu202302).

Appendix A More Visualizations

A.2 More visualizations on COCO-stuff

A.3 More visualizations on Visual Genome

A.4 More comparision with previous methods

A.5 Different Scales

Appendix B Details of Diffusion Models

In this section, we will review the formulation of Gaussian diffusion models introduced by DDPM . A data point is defined as x0∼q(x0)x_{0}\sim q(x_{0}). By gradually adding noise to the clean data x0x_{0}, we can obtain the noised samples from x1x_{1} to xTx_{T}, where TT denotes the maximum steps. Specifically, Gaussian noise according to some variance schedule given by βt\beta_{t} is added to xt−1x_{t-1} in step tt of the Markovian noising process qq:

Due to the convenient nature of Gaussian noise, we do not need to apply tt times of q(xt∣xt−1)q(x_{t}|x_{t-1}) repeatedly to sample from xt∼q(xt∣x0)x_{t}\sim q(x_{t}|x_{0}). Instead, q(xt∣x0)q(x_{t}|x_{0}) can be directly sampled from a Gaussian distribution:

If the total noise added throughout the markov chain is large enough when T→∞T\to\infty and correspondingly βt→0\beta_{t}\to 0, the xTx_{T} will be well approximated by N(0,I)\mathcal{N}(0,\mathbf{I}). This nice property ensures that we can reverse the above forward process and sample from xT∼N(0,I)x_{T}\sim\mathcal{N}(0,\mathbf{I}), which is a Gaussian noise. However, since the entire dataset is needed, we cannot easily estimate the posterior q(xt−1∣xt,x0)q(x_{t-1}|x_{t},x_{0}). Instead, we have to learn a model pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) to approximate it:

To ensure that the estimated pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) can learrn the true data distribution q(x0)q(x_{0}), we can optimize the following variational lower bound LvlbL_{\text{vlb}} for pθ(x0)p_{\theta}(x_{0}):

Although the above objective is well justified, DDPM applied a different objective that produces better samples in practice. Specifically, they do not directly predict μθ(xt,t)\mu_{\theta}(x_{t},t) as the output of a neural network, but instead train a model ϵθ(xt,t)\epsilon_{\theta}(x_{t},t) to predict ϵ\epsilon from Equation 22. This simplified objective is defined as follows:

Then, we can derive μθ(xt,t)\mu_{\theta}(x_{t},t) from ϵθ(xt,t)\epsilon_{\theta}(x_{t},t) using the following substitution:

B.2 Classifier-free Method for Layout-conditional Training and Sampling

Instead of training a separate classifier model, Ho & Salimans choose to train a diffusion model that allows for both conditional and unconditional sampling, where the unconditional diffusion model pθ(xt,t)p_{\theta}(x_{t},t) is parameterized through a score estimator ϵθ(xt,t)\epsilon_{\theta}(x_{t},t) and the conditional diffusion model pθ(xt,t∣c)p_{\theta}(x_{t},t|c) is parameterized through ϵθ(xt,t,c)\epsilon_{\theta}(x_{t},t,c).

which can be considered as a linear combination of conditional and unconditional score estimates. ss is the scale and Since Eq. 33 has no classifier gradient, no more gradient calculation in classifier guidance is needed during sampling. Furthermore, ss can be changed to modify the effect of the condition.

In the layout-to-image generation, cc is the layout ll defined in Sec. 3.1. Layout Embedding and ∅\varnothing is the emtpy layout lpadl_{\text{pad}} defined in Sec. 3.4. Layout-conditional Diffusion Model.

Appendix C Implementation details

C.2 Analysis of Training Resources and Sampling Speed

Appendix D Evaluation

COCO-Stuff . COCO is a large-scale object detection, segmentation, and captioning dataset. COCO-Stuff augments all 164K images of the COCO dataset with pixel-level stuff annotations. These annotations can be used for scene understanding tasks like semantic segmentation, object detection and image captioning. COCO-Stuff contains 80 categories of thing and 91 categories of stuff, respectively. Following the settings of LostGAN-v2, we use the COCO 2017 Stuff Segmentation Challenge subset containing 40K / 5k / 5k images for train / val / test-dev set. Segmengtation annotation is not used. We use images in the train and val set with 3 to 8 objects that cover more than 2% of the image and not belong to ’crowd’. Finally, there are 25,210 train and 3,097 val images. Some previous works don’t filter objects belong to ’crowd’, causing different numbers of images. Specificly, there are 24,972 train and 3,074 val images. LAMA uses the full COCO-Stuff 2017 dataset, leaving 74,777 train and 3,097 val images. Tab. 7 summarizes the difference in the number of images by different filtering methods.

Visual Genome . Following the settings of Sg2Im, we experiment on Visual Genome version 1.4 (VG) which comprises 108,077 images annotated with scene graphs. Visual Genome collects images with dense annotations of objects, attributes, and relationships. Here, we only use bounding boxes. We divide the data into 80% / 10% / 10% for the train / val / test set. We select the object / relationship categories occurring at least 2000 / 500 times in the train set, respectively, and select the images with 3 to 30 bounding boxes and ignoring all small objects. Finally, the train / val / test set has 62,565 / 5,062 / 5,096 images.

D.2 Evaluation Metrics

Comprehensive evaluations of generated images remains a challenge. We use six metrics, from image-level to layout-level, to evaluate the quality of the generated images and the layout control from different aspects.

Fr‘echet Inception Distance (FID) shows the difference between the real images and the generated images by using an ImageNet-pretrained Inception-V3 network and computing the Fr‘echet distance between two Gaussian distributions fitted to generated images and real images respectively. We save the GT images as real images when sampling the generated images. The real images and the generated images are saved to two folder respectively. Then compute the FID score, based on the official code of FIDthe official code of FID:https://github.com/bioinf-jku/TTUR. For FID, the lower the score, the smaller the difference between the generated images and the real images, meaning the generator model is better.

Inception Score (IS) shows the overall quality of the generated images by using an Inception-V3 network pretrained on the ImageNet-1000 classification benchmark and computing a score(statistics) of the network’s outputs with generated images of a generator model. IS measures the quality of images on two aspects: clarity and diversity. We only need the generated images to compute the IS score, based on the official code of ISthe official code of IS:https://github.com/openai/improved-gan. For IS, the higher the score, the better the quality of generated images, meaning the generator model is better.

Diversity Score (DS) measures the diversity between the generated images from the same layout by comparing the perceptual similarity in a DNN feature space between them. Here, we adopt the LPIPS metric. For each sample, we repeat two times, use the two images from the same layout to compute the DS score, then calculate mean and std of these scores as the reported DS score, based on the official code of DSthe official code of DS :https://github.com/richzhang/PerceptualSimilarity. For DS, the higher the score, the better the diversity between the generated images, meaning the generator model is better.

YOLO Score uses a pretrained YOLOv4 model to evaluate bbox mAP on 80 thing categories based on the official code of LAMAthe official code of LAMA :https://github.com/ZejianLi/LAMA and the official code of YOLOv4the official code of YOLOv4 :https://github.com/AlexeyAB/darknet. YOLO Score is proposed to evaluate the alignmentand fidelity of generated objects, measuring how generated objects are recognizable when even the layout is unknow. YOLO is a well-known series of object detector, inferring layouts from the given images. Before send images to detector, they are upsampled to 512×\times512. And different from LAMA, since we filter the objects and images in datasets, we think it is better to evaluate bbox mAP only on filtered annotations.

Classification Score (CAS) measures classification accuracy of layout areas on generated images. We crop the GT box area of images and resize objects at a resolution of 32×\times32 with their class. Then train a ResNet101 classifier with cropped images on generated images and test it on cropped images on real images, based on a widely used codebase of image classificationthe widely used codebase of image classification: https://github.com/hysts/pytorch_image_classification. For CAS, the higher the score, the better the quality of layout control, meaning the generator model is better.

References