Diffusion-Based Scene Graph to Image Generation with Masked Contrastive Pre-Training
Ling Yang, Zhilin Huang, Yang Song, Shenda Hong, Guohao Li, Wentao Zhang, Bin Cui, Bernard Ghanem, Ming-Hsuan Yang
Introduction
Image generation has made remarkable progress in the past few years brock2018large; razavi2019generating; crowson2022vqgan; gafni2022make, largely due to the success of diffusion and score-based generative models sohl2015deep; song2019generative; ho2020denoising; song2020score. These methods allow for the creation of realistic and diverse image samples dhariwal2021diffusion; ramesh2022hierarchical; saharia2022photorealistic, which users can specify through various forms—labels song2020score; dhariwal2021diffusion; ho2022classifier; kim2022diffusionclip, captions rombach2022high; saharia2022photorealistic, segmentation masks couairon2022diffedit, sketches wang2022pretraining, stroke paintings meng2021sdedit, and more Yang2022DiffusionMA. However, these types of specifications often fall short when it comes to complex relations between multiple objects in images. Instead, scene graphs provide a concise and accurate way of depicting objects and their relations to one another johnson2015image; krishna2017visual. It is therefore crucial to investigate image generation based on scene graphs as a means of synthesizing complex scenes johnson2018image.
A key challenge in generating images from scene graphs is ensuring that the resulting image closely aligns with the input scene graph. To this end, generative models must be able to understand the correspondence between the two vastly different data domains: images and graphs. Existing methods johnson2018image; herzig2020learning; ashual2019specifying; li2019pastegan mainly address this challenge by using an image-like representation of scene graphs, often in the form of scene layouts, to create coarse sketches for guiding the image generation process. These sketches are then refined by generative models to produce realistic images that follow the specifications given by the scene graph.
While intermediate representations such as scene layouts can be useful, they are often crafted manually and are not specifically designed to facilitate the alignment between images and graphs. For instance, in the case of scene layouts, nodes in scene graphs are usually mapped to bounding boxes and connections are mapped to their spatial layouts. However, not all connections within scene graphs can be accurately translated to spatial layouts, such as eating and looking at. Additionally, some relations, such as behind, inside, and in front of, all correspond to similar spatial relations in scene layouts, creating ambiguity. These intermediate scene layout representations may also contain extraneous information that complicates the training of downstream generative models.
To overcome these limitations, we propose learning intermediate representations that explicitly maximize the alignment between scene graphs and images. We provide an illustration of our approach in Fig. 1. Specifically, we pre-train a scene graph encoder on graph-image pair datasets to produce embeddings that extract both local and global information from scene graphs, while maximizing their alignment with images. To extract local information, we introduce a masked autoencoding loss, randomly masking out objects in an image and reconstructing the missing portion using unmasked regions and embeddings acquired from the encoder. To gain global information, we leverage contrastive learning to train our encoder to discern between images that do and do not adhere to scene graphs. By combining embeddings obtained from both approaches, we obtain compact intermediate representations of scene graphs that facilitate the alignment between graphs and images.
We showcase the effectiveness of our scene graph embeddings by building a latent diffusion model vahdat2021score; rombach2022high that generates images from scene graphs with the aid of our pre-trained embeddings. We evaluate the importance of local and global embeddings through ablation studies, and demonstrate the clear advantages our approach has over traditional intermediate layout representations. Our model, dubbed SGDiff, successfully generates images that capture accurate local and global structures of scene graphs. Additionally, our model enables the semantic manipulation of images through scene graph surgery. We evaluate SGDiff on standard datasets such as Visual Genome (VG) krishna2017visual and COCO-Stuff caesar2018coco, and find that it performs better than current state-of-the-art approaches in both qualitative comparison and quantitative measurements.
Related Work
Diffusion Models for Conditional Image Synthesis. Diffusion models Yang2022DiffusionMA generate data samples by learning to reverse a prescribed diffusion process that converts data to noise. First introduced by Sohl-Dickstein et al. sohl2015deep and later improved by Song & Ermon song2019generative and Ho et al. ho2020denoising, they are now able to generate image samples with unprecedented quality and diversity gu2022vector; saharia2022photorealistic; saharia2022palette. Diffusion models excel at conditional image synthesis using various forms of user guidance, such as text and images. Existing conditional diffusion models often leverage auxiliary classifiers song2020score; dhariwal2021diffusion; ho2022classifier; pmlr-v162-nichol22a; kim2022diffusionclip to incorporate conditional information into the data generating process, using methods of classifier guidance song2020score; dhariwal2021diffusion or classifier-free guidance ho2022classifier. Latent Diffusion Models (LDMs) vahdat2021score; rombach2022high reduce the training cost for high resolution images by learning the diffusion model in a low-dimensional latent space. They also incorporate conditional information into the sampling process via cross attention vaswani2017attention. Similar techniques are employed in DALLE-2 ramesh2022hierarchical for image generation from text, where the diffusion model is conditioned on text embeddings obtained from CLIP latent codes radford2021learning. Alternatively, Imagen saharia2022photorealistic implements text-to-image generation by conditioning on text embeddings acquired from large language models (e.g., T5 raffel2020exploring). Despite all this progress on diffusion-based conditional image synthesis, generating images from graph-structured data is under-explored. We fill this gap by designing the first diffusion model for image generation from scene graphs, leveraging new scene graph embeddings constructed from self-supervised learning.
Image Generation from Scene Graphs. Scene graphs are graph-structured data for describing multiple objects and their complex relationships in scene images, wherein nodes represent objects and edges represent relations johnson2015image; krishna2017visual. Image generation from scene graphs johnson2018image; tripathi2019using requires the generative model to reason over both objects and their relations. The first generative model of this kind, Sg2Im johnson2018image, proposes a two-stage generation pipeline. First, an embedding model is trained to map scene graphs to scene layouts, which are image-like representations that capture the coarse structure of images to generate. Second, a generative model is trained to refine scene layouts into realistic images. Most subsequent work on this task follows the same pipeline. For example, WSGC herzig2020learning accounts for semantic equivalence in graph representations by canonicalizing scene graphs before mapping them to scene layouts. Li et al. li2019pastegan and Ashual & Wolf ashual2019specifying leverage a repository of external reference images to improve the quality of scene layouts. Other works that rely on scene layout style representations include zhao2020layout2image; he2021context; sylvain2021object; sun2019image; li2021image. In lieu of manually crafted scene layouts, we propose learning scene graph embeddings that are both concise and predictive of graph-image alignment. We then use these embeddings to build a latent diffusion model for scene graph to image generation, avoiding the limitations of scene layouts.
Method
In what follows, we first discuss how to learn effective embeddings of scene graphs via self-supervised learning, then focus on using these embeddings to build diffusion models for scalable scene graph to image generation.
Scene layouts are manually constructed representations of scene graphs (SGs). While effective in many cases, they are not specifically optimized to capture all the necessary information from SGs for generating the corresponding scene images. This can lead to suboptimal alignment between SG inputs and generated images. To overcome this limitation, we propose to directly learn such SG representations via self-supervised learning. Specifically, given a dataset that contains SG-image pairs, we learn an SG encoder to produce embeddings that capture both local and global information of the input SG, while maximizing their alignment with the corresponding scene images.
Given a set of objects and a set of relations , we denote a scene graph with a tuple . Here represents the set of objects in the scene and denotes the set of relations that exist between these objects. We use a triplet to denote a directed connection from to , representing the relation tuple . We denote by (resp. ) the set of children (resp. parents) for node . To generate an effective representation for , we embed both objects and relations by iterating the following:
Below, we introduce two self-supervised techniques for learning object and relation embeddings, focusing on extracting local and global information from SG inputs.
Pretraining with Masked Autoencoding.
Masked pre-training is a preeminent technique in image/text representation learning devlin2019bert; bao2021beit; he2022masked; chang2022maskgit; xie2022simmim and visual-language modeling su2019vl; lu2019vilbert; zhang2021vinvl. Inspired by its success in these applications, we propose to use a similar technique for learning our SG embeddings.
In particular, we randomly choose a triplet in the scene graph and mask out objects and in the corresponding scene image . We denote this masked area as , and the remaining image as . To train the SG encoder, we consider the task of predicting from and embeddings obtained from the SG encoder. Specifically, we concatenate object and relation embeddings from the SG encoder to form , defined as:
where is sampled uniformly at random from , a dataset of graph-image pairs. By training the SG embeddings to reconstruct randomly masked areas in scene images, we explicitly encourage the graph embeddings to encode local structural information that focuses on predicting fine-grained image details from scene graphs.
Pretraining with Contrastive Learning.
Contrastive pre-training chen2020simple; he2020momentum is a widely adopted technique for learning shared representations across multiple data modalities. They have found success in many visual-language modeling applications radford2021learning; patashnik2021styleclip. We propose to leverage contrastive learning as a second way to train our SG embeddings, with a focus on capturing global structural information. Specifically, we compute a graph-level embedding by concatenating all object and relation embeddings obtained from the SG encoder, that is,
where is a learnable multiplicative scalar that acts as a temperature parameter, denotes the embeddings of an image produced by a trainable image encoder, and denotes the set of images in dataset that do not comply with the SG . This contrastive training objective optimizes our SG embeddings to capture global structures that can identify whether images are in line with the SGs or not.
We combine both the masked autoencoding loss and the contrastive loss to form our training objective for SG embeddings:
where is a hyperparameter. With this objective function, we can train an effective SG encoder that converts SGs to embeddings without losing predictive information of the matching scene images. Such embeddings provide strong conditioning signals that facilitate downstream diffusion models to generate images from SGs.
2 Diffusion-Based SG to Image Generation
In diffusion-based conditional image synthesis, we train diffusion models to sample from , where is an image, and is a conditioning signal, typically taking the form of texts in text-to-image generation saharia2022photorealistic; rombach2022high and image editing avrahami2022blended; kawar2022imagic; valevski2022unitune, or reference images in the context of image translation meng2021sdedit. We consider the task of image generation from SGs in this work. With the SG encoder obtained from the previous section, we can build diffusion models to generate images from SGs by setting to the SG embeddings. For scalable image modeling, we build upon the framework of latent diffusion vahdat2021score; rombach2022high, where the diffusion model is trained in a low-dimensional latent space obtained from a pre-trained autoencoder.
After training both the encoder and the decoder, we use the encoder to generate latent codes for all images in the training dataset, then train a diffusion model on these latent codes separately. In particular, given the latent code for a randomly sampled training image , we convert it to noise with a Markov process defined by the transition kernel , where , , and is a hyper-parameter that controls the rate of noise injection. When the amount of noise is sufficiently large, becomes approximately distributed according to . In order to convert noise back to data for sample generation, we have to estimate the reverse diffusion process by learning the reverse transition kernel as an approximation to . Following Ho et al. ho2020denoising, we parameterize with a neural network (called the score model song2019generative; song2020score) and fix to be a constant. The score model can be optimized with denoising score matching Hyvrinen2005EstimationON; vincent2011connection. For sample generation, we first generate a latent code with the diffusion model, then produce an image sample through the pre-trained decoder.
Conditioning on SG Embeddings.
With the method in Section 3.1, we can learn an SG encoder to produce two types of scene graph embeddings, and , defined according to Eq. 3 and Eq. 5. To combine these embeddings, we first merge and by summing with each triplet of the form (cf., Eq. 3), which has been generated in the process of computing . This gives us a new embedding for each relation:
Afterwards, we concatenate and transform them to get our final embedding, given by
where is a trainable model called the SG conditioner. We then use the resulting embedding to guide the generation of our latent diffusion model.
Here is defined as:
Experimental Results
The Visual Genome (VG) krishna2017visual dataset contains 108,077 scene graph & image pairs, with additional annotations such as bounding boxes and object attributes. Scene graphs in this dataset have 179 object identities, 80 attributes, and 49 relations. For a fair comparison, we follow previous works johnson2018image; herzig2020learning to use the standard dataset splits: 62,565 pairs for training, 5,506 for validation, and 5,088 for test. COCO-Stuff caesar2018coco contains pixel-wise annotations with 40,000 training images and 5,000 validation images with corresponding bounding boxes and segmentation masks. It has 80 item categories and 91 stuff categories. Following johnson2018image, we use synthesized scene graphs and standard dataset splits: 25,000 for training, 1,024 for validation, and 2,048 for test.
We report results with two widely used evaluation metrics. One is Inception Score (IS) salimans2016improved, which measures both the quality and diversity of synthesized images. Same as previous work, we employ a pre-trained inception network szegedy2016rethinking to obtain network activations for computing IS. The IS is better when larger. The other evaluation metric is Fréchet Inception Distance (FID) heusel2017gans, which is reported to align well with human evaluation. It measures the distance between the distribution of the generated images and that of the real test images, both modeled as multivariate Gaussians. The FID values are better when lower.
Baselines and Implementation Details.
In our experiments, we choose four previous methods on scene graph to image generation as our baselines: Sg2Im johnson2018image, WSGC herzig2020learning, SOAP ashual2019specifying, and PasteGAN li2019pastegan. We follow their evaluation settings in all experiments. For masked contrastive pretraining, we train models using the Adam optimizer kingma2015adam with a learning rate of 5e-4 and a batch size of 64 for 100,000 iterations. For latent diffusion training, we train models from scratch using the same optimizer but with a learning rate of 1e-6 and a batch size of 16 for 700,000 iterations. Please consult the Appendix A for more details.
1 Quantitative Comparisons
Similar to prior works johnson2018image; herzig2020learning; ashual2019specifying; li2019pastegan, we perform quantitative comparisons on the COCO-Stuff and VG datasets for image resolution 6464, 128128, and 256256. All IS and FID results are provided in Table 1. We observe that our SGDiff consistently outperforms existing methods on both evaluation metrics by a significant margin, demonstrating superiority in both generation fidelity and diversity. One possible explanation is that conventional intermediate representations of scene graphs used in previous methods, as exemplified by scene layouts, do not specifically optimize the semantic alignment between scene graphs and images. By contrast, SGDiff directly learns scene graph embeddings, maximizing both local and global semantic compliance with images in a self-supervised manner. Since our latent diffusion model incorporates these optimized embeddings to guide the generation process, we naturally achieve improved quality in image generation. We verify this hypothesis with rigorous ablation studies in Section 4.3. We also test the two methods for conditioning on scene graph embeddings, and find that concatentation with the noisy latent code consistently outperforms concatenation with time embedding .
2 Qualitative Evaluations
To explore SGDiff’s ability of generating images that contain multiple objects and complex relations, we visualize some typical examples (with resolution 128128) in Fig. 3, where images are placed in ascending order of scene graph complexity. We compare SGDiff with a classical algorithm Sg2Im johnson2018image and the state-of-the-art method PasteGAN li2019pastegan. It is clear that images synthesized by SGDiff are not only more realistic, but also comply with the corresponding scene graphs better than both Sg2Im and PasteGAN. In contrast, images generated by Sg2Im are fuzzy and lack fine-grained details. Although PasteGAN can generate images with more details compared to Sg2Im, it tends to miss some important relations specified by the scene graphs, leading to obfuscated results. Both methods are prone to generating images with wrong objects or relations. In contrast, our SGDiff consistently generates realistic images that match scene graphs well, since we rely on scene graph embeddings that are explicitly optimized for local and global semantic alignment. The generated objects have clear object boundaries and more visual details. We place more image samples in Appendix B.
To demonstrate the semantic consistency between generated images and scene graphs, we apply our SGDiff to manipulate image samples by modifying objects and relations in scene graph inputs. As shown in Fig. 4, SGDiff can not only produce compliant manipulation results (of resolution 256256) with respect to objects and relations, but also synthesize perceptually diverse images when conditioned on the same scene graph. The results demonstrate that our model can effectively leverage the scene graph embeddings learned through masked contrastive pre-training.
3 Ablation Studies
Key to our approach is learning scene graph embeddings via masked contrastive pre-training. Here we perform ablation studies to understand the importance of masked autoencoding loss and contrastive loss in training our scene graph encoder. We also empirically verify that our scene graph embeddings outperform manually crafted scene layout representations when combined with the same latent diffusion model on scene graph to image generation.
We first evaluate the impacts of masked autoencoding loss and contrastive loss on cross-modal semantic alignment through graph-to-image and image-to-graph retrieval experiments. In graph-to-image retrieval, we use scene graph embeddings and image embeddings to search the most semantically similar image for a given scene graph, then report the accuracy of finding the correctly paired image from the given dataset. The image-to-graph retrieval task is defined analogously. All results are provided in Table 2. Here “obj.” stands for experiments where we only train object embeddings, whereas “Obj. + Rel.” represents settings where we train both object and relation embeddings. We observe that using both embeddings boost the performance on all retrieval tasks. With only the contrastive loss, we can already obtain over 70% accuracy in graph-to-image and image-to-graph retrieval. Adding masked auto-encoding loss further improves the accuracy, demonstrating better graph-image alignment.
Generation from Scene Graphs.
To understand the role of masked contrastive pre-training in scene graph to image generation, we train the same latent diffusion model with different scene graph embeddings: naïve one-hot embeddings, layout embeddings, masked embeddings, constrastive embeddings, and combination of the last two. As for one-hot embeddings, we fix object and relation embeddings as one-hot vectors of their categories, then concatenate these embeddings together to condition the latent diffusion model. For conventional scene layout embeddings, we follow the same settings in PasteGAN li2019pastegan. In Table 3, we report IS and FID scores for image samples of resolution 256256. We observe that embeddings obtained from either masked pretraining or contrastive pretraining outperforms one-hot embeddings or scene layout representations by a significant margin, and combining them together further improves sample quality.
For qualitative comparison, we provide image samples with different scene graph embeddings in Fig. 5. Compared with one-hot embeddings and scene layout representations, we observe that embeddings from masked pre-training and contrastive pre-training enable SGDiff to generate images that are more realistic and comply better with the scene graphs. Compared to contrastive pretraining, we observe that embeddings obtained from masked pretraining tend to produce images that capture local structures better with more fine-grained object details, whereas embeddings from contrastive pretraining tend to focus more on matching the global structures at the overall image level.
Conclusion
This paper proposes a new framework, SGDiff, for image generation from scene graphs. SGDiff uses a masked contrastive pre-training approach to obtain scene graph emebeddings that allow for improved alignment between scene graphs and images, while also leveraging latent diffusion for improved scalability and generation quality. As a result, SGDiff produces more realistic and compliant images than previous methods that rely on manually crafted scene graph representations, such as scene layouts. SGDiff also makes image generation semantically controllable, allowing for easier manipulation of images through scene graph editing. Evaluation on standard datasets such as Visual Genome and COCO-Stuff show that SGDiff outperforms state-of-the-art approaches both qualitatively and quantitatively.
References
Appendix A Implementation Details
As illustrated in main text, our proposed framework SGDiff consists of three (pre-)training stages, masked contrastive pre-training for SG encoder, variational autoencoder pre-training for latent embedding of images, and latent diffusion training for SG-based image generation. Here we introduce the concrete network and optimization details of these stages in Table 4 and Table 5. We use the same settings on both VG and COCO-Stuff datasets.
In masked contrastive pre-training, the nodes and edges of SGs are all preprocessed into 512-dimensional vectors for SG encoder. The ViT model tokenizes input images with patch size of 3232. Specifically, we set the ratio of random mask to 0.3 in masked autoencoding branch, and set the ratio between masked autoencoding loss and contrastive loss to 10:1 for facilitating the optimization. In variational autoencoder pre-training, we embed input images into compact latents with a downsampling factor of 8, and maximize the decoding ability by optimizing the MSE objective. And the ratio between KL divergence and MSE is set to 8:10. In latent diffusion training, we use cross-attention mechanism for conditional diffusion process in all experiments.
Appendix B More Synthesis Results
In previous qualitative evaluations, we have shown the synthesis comparison results on VG dataset. Here, we provide more results on COCO-Stuff dataset with the resolution of 256256 in Fig. 6 and Fig. 7. SGDiff exhibits the superiority of generation quality on COCO-Stuff dataset, and SGDiff can generate images that are more realistic and semantic-compliant than previous methods Sg2Im johnson2018image and PasteGAN li2019pastegan. The results not only reveal the SG embeddings learned by our masked contrastive pre-training are effective in graph-image semantic alignment, but also demonstrate the efficacy of our local-global conditional latent diffusion.