Improving Diffusion-Based Image Synthesis with Context Prediction

Ling Yang, Jingwei Liu, Shenda Hong, Zhilong Zhang, Zhilin Huang, Zheming Cai, Wentao Zhang, Bin Cui

Introduction

Recent diffusion models have made remarkable progress in image generation. They are first introduced by Sohl-Dickstein et al. and then improved by Song & Ermon and Ho et al. , and can now generate image samples with unprecedented quality and diversity . Numerous methods have been proposed to develop diffusion models by improving their empirical generation results or extending the capacity of diffusion models from a theoretical perspective . We revisit existing diffusion models for image generation and break them into two categories, pixel- and latent-based diffusion models, according to their diffusing spaces. Pixel-based diffusion models directly conduct continuous diffusion process in the pixel space, they incorporate various conditions (e.g., class, text, image, and semantic map) or auxiliary classifiers for conditional image generation.

On the other hand, latent-based diffusion models conduct continuous or discrete diffusion process on the semantic latent space. Such diffusion paradigm not only significantly reduces the computational complexity for both training and inference, but also facilitates the conditional image generation in complex semantic space . Some of them choose to pre-train an autoencoder to map the input from image space to the continuous latent space for continuous diffusion, while others utilize a vector quantized variational autoencoder to induce the token-based latent space for discrete diffusion .

Despite all these progress of pixel- and latent-based diffusion models in image generation, both of them mainly focus on utilizing a point-based reconstruction objective over the spatial axes to recover the entire image in diffusion training process. This point-wise reconstruction neglects to fully preserve local context and semantic distribution of each predicted pixel/feature, which may impair the fidelity of generated images. Traditional non-diffusion studies have designed different context-preserving terms for advancing image representation learning, but few researches have been done to constrain on context for diffusion-based image synthesis.

In this paper, we propose ConPreDiff to explicitly force each pixel/feature/token to predict its local neighborhood context (i.e., multi-stride features/tokens/pixels) in image diffusion generation with an extra context decoder near the end of diffusion denoising blocks. This explicit context prediction can be extended to existing discrete and continuous diffusion backbones without introducing additional parameters in inference stage. We further characterize the neighborhood context as a probability distribution defined over multi-stride neighbors for efficiently decoding large context, and adopt an optimal-transport loss based on Wasserstein distance to impose structural constraint between the decoded distribution and the ground truth. We evaluate the proposed ConPreDiff with the extensive experiments on three major visual tasks, unconditional image generation, text-to-image generation, and image inpainting. Notably, our ConPreDiff consistently outperforms previous diffusion models by a large margin regarding generation quality and diversity.

Our main contributions are summarized as follows: (i): To the best of our knowledge, we for the first time propose ConPreDiff to improve diffusion-based image generation with context prediction; (ii): We further propose an efficient approach to decode large context with an optimal-transport loss based on Wasserstein distance; (iii): ConPreDiff substantially outperforms existing diffusion models and achieves new SOTA image generation results, and we can generalize our model to existing discrete and continuous diffusion backbones, consistently improving their performance.

Related Work

Diffusion models are a new class of probabilistic generative models that progressively destruct data by injecting noise, then learn to reverse this process for sample generation. They can generate image samples with unprecedented quality and diversity , and have been applied in various applications . Existing pixel- and latent-based diffusion models mainly utilize the discrete diffusion or continuous diffusion for unconditional or conditional image generation . Discrete diffusion models were also first described in , and then applied to text generation in Argmax Flow . D3PMs applies discrete diffusion to image generation. VQ-Diffusion moves discrete diffusion from image pixel space to latent space with the discrete image tokens acquired from VQ-VAE . Latent Diffusion Models (LDMs) reduce the training cost for high resolution images by conducting continuous diffusion process in a low-dimensional latent space. They also incorporate conditional information into the sampling process via cross attention . Similar techniques are employed in DALLE-2 for image generation from text, where the continuous diffusion model is conditioned on text embeddings obtained from CLIP latent codes . Imagen implements text-to-image generation by conditioning on text embeddings acquired from large language models (e.g., T5 ). Despite all this progress, existing diffusion models neglect to exploit rich neighborhood context in the generation process, which is critical in many vision tasks for maintaining the local semantic continuity in image representations . In this paper, we firstly propose to explicitly preserve local neighborhood context for diffusion-based image generation.

Context-Enriched Representation Learning

Context has been well studied in learning representations, and is widely proved to be a powerful automatic supervisory signal in many tasks. For example, language models learn word embeddings by predicting their context, i.e., a few words before and/or after. More utilization of contextual information happens in visual tasks, where spatial context is vital for image domain. Many studies propose to leverage context for enriching learned image representations. Doersch et al. and Zhang et al. make predictions from visible patches to masked patches to enhance the self-supervised image representation learning. Hu et al. designs local relation layer to model the context of local pixel pairs for image classification, while Liu et al. preserves contextual structure to guarantee the local feature/pixel continuity for image inpainting. Inspired by these studies, in this work, we propose to incorporate neighborhood context prediction for improving diffusion-based generative modeling.

Preliminary

We briefly review a classical discrete diffusion model, namely Vector Quantized Diffusion (VQ-Diffusion) . VQ-Diffusion utilizes a VQ-VAE to convert images xx to discrete tokens x0∈{1,2,...,K,K+1}x_{0}\in\{1,2,...,K,K+1\}, KK is the size of codebook, and K+1K+1 denotes the [MASK][\text{\tt{MASK}}] token. Then the forward process of VQ-Diffusion is given by:

which is optimized by minimizing the following variational lower bound (VLB) :

Continuous Diffusion

A continuous diffusion model progressively perturbs input image or feature map x0\bm{x}_{0} by injecting noise, then learn to reverse this process starting from xT\bm{x}_{T} for image generation. The forward process can be formulated as a Gaussian process with Markovian structure:

where β1,…,βT\beta_{1},\ldots,\beta_{T} denotes fixed variance schedule with αt:=1−βt\alpha_{t}:=1-\beta_{t} and α‾t:=∏s=1tαs\overline{\alpha}_{t}:=\prod_{s=1}^{t}\alpha_{s}. This forward process progressively injects noise to data until all structures are lost, which is well approximated by N(0,I)\mathcal{N}(0,\mathbf{I}). The reverse diffusion process learns a model pθ(xt−1∣xt)p_{\theta}(\bm{x}_{t-1}|\bm{x}_{t}) that approximates the true posterior:

Fixing Σθ\Sigma_{\theta} to be untrained time dependent constants σt2I\sigma^{2}_{t}\bm{I}, Ho et al. improve the diffusion training process by optimizing following objective:

where CC is a constant that does not depend on θ\theta. μ^(xt,x0)\hat{\mu}(\bm{x}_{t},\bm{x}_{0}) is the mean of the posterior q(xt−1∣x0,xt)q(\bm{x}_{t-1}|\bm{x}_{0},\bm{x}_{t}), and μθ(xt,t)\mu_{\theta}(\bm{x}_{t},t) is the predicted mean of pθ(xt−1∣xt)p_{\theta}(\bm{x}_{t-1}\mid\bm{x}_{t}) computed by neural networks.

The Proposed ConPreDiff

In this section, we elucidate the proposed ConPreDiff as in Figure 1. In Sec. 4.1, we introduce our proposed context prediction term for explicitly preserving local neighborhood context in diffusion-based image generation. To efficiently decode large context in training process, we characterize the neighborhood information as the probability distribution defined over multi-stride neighbors in Sec. 4.2, and theoretically derive an optimal-transport loss function based on Wasserstein distance to optimize the decoding procedure. In Sec. 4.3, we generalize our ConPreDiff to both existing discrete and continuous diffusion models, and provide optimization objectives.

SS-Stride Neighborhood Reconstruction Previous diffusion models make point-wise reconstruction, i.e., reconstructing each pixel, thus their reverse learning processes can be formulated by pθ(xt−1i∣xt)p_{\theta}(\bm{x}^{i}_{t-1}|\bm{x}_{t}). In contrast, our context prediction aims to reconstruct xt−1i\bm{x}^{i}_{t-1} and further predict its ss-stride neighborhood contextual representations HNis\bm{H}_{\mathcal{N}^{s}_{i}} based on xt−1i\bm{x}^{i}_{t-1}: pθ(xt−1i,HNis∣xt),p_{\theta}(\bm{x}^{i}_{t-1},\bm{H}_{\mathcal{N}^{s}_{i}}|\bm{x}_{t}), where pθp_{\theta} is parameterized by two reconstruction networks (ψp\psi_{p},ψn\psi_{n}). ψp\psi_{p} is designed for the point-wise denoising of xt−1i\bm{x}^{i}_{t-1} in xt\bm{x}_{t}, and ψn\psi_{n} is designed for decoding HNis\bm{H}_{\mathcal{N}^{s}_{i}} from xt−1i\bm{x}^{i}_{t-1}. For denoising ii-th point in xt\bm{x}_{t}, we have:

where tt is the time embedding and ψp\psi_{p} is parameterized by a U-Net or transformer with an encoder-decoder architecture. For reconstructing the entire neighborhood information HNis\bm{H}_{\mathcal{N}^{s}_{i}} around each point xt−1i\bm{x}^{i}_{t-1}, we have:

where x,yx,y are the width and height on spatial axes. x^i\bm{\hat{x}}^{i} (x^0i\bm{\hat{x}}_{0}^{i}) and H^Nis\bm{\hat{H}}_{\mathcal{N}^{s}_{i}} are ground truths. Mp\mathcal{M}_{p} and Mn\mathcal{M}_{n} can be Euclidean distance. In this way, ConPreDiff is able to maximally preserve local context for better reconstructing each pixel/feature/token.

We let Mp,Mn\mathcal{M}_{p},\mathcal{M}_{n} be square loss, Mn(HNis,H^Nis)=∑j∈Ni(x0i,j−x^0i,j)2,\mathcal{M}_{n}(\bm{H}_{\mathcal{N}^{s}_{i}},\bm{\hat{H}}_{\mathcal{N}^{s}_{i}})=\sum_{j\in\mathcal{N}_{i}}(\bm{x}_{0}^{i,j}-\hat{\bm{x}}_{0}^{i,j})^{2}, where x^0i,j\hat{\bm{x}}_{0}^{i,j} is the j-th neighbor in the context of x^0i\hat{\bm{x}}_{0}^{i} and x0i,j\bm{x}_{0}^{i,j} is the prediction of x0i,j\bm{x}_{0}^{i,j} from a denoising neural network. Thus we have:

Compactly, we can write the denoising network as:

We will show that the DDPM loss is upper bounded by ConPreDiff loss, by reparameterizing x0(xt,t)\bm{x}_{0}(\bm{x}_{t},t). Specifically, for each unit ii in the feature map, we use the mean of predicted value in its neighborhood as the final prediction:

Now we can show the connection between the DDPM loss and ConPreDiff loss:

In the last equality, we assume that the feature is padded so that each unit ii has the same number of neighbors ∣N∣|\mathcal{N}|. As a result, the ConPreDiff loss is an upper bound of the negative log likelihood.

Complexity Problem

We solve the challenging problem by changing the direct prediction of entire neighborhoods to the prediction of neighborhood distribution. Specifically, for each xt−1i\bm{x}^{i}_{t-1}, the neighborhood information is represented as an empirical realization of i.i.d. sampling QQ elements from PNis\mathcal{P}_{\mathcal{N}^{s}_{i}}, where PNis≜1K∑u∈Nisδhu\mathcal{P}_{\mathcal{N}^{s}_{i}}\triangleq\frac{1}{K}\sum_{u\in\mathcal{N}^{s}_{i}}\delta_{h_{u}}. Based on this view, we are able to transform the neighborhood prediction Mn\mathcal{M}_{n} into the neighborhood distribution prediction. However, such sampling-based measurement loses original spatial orders of neighborhoods, and thus we use a permutation invariant loss (Wasserstein distance) for optimization. Wasserstein distance is an effective metric for measuring structural similarity between distributions, which is especially suitable for our neighborhood distribution prediction. And we rewrite the Equation 9 as:

where ψn(xt−1i,t)\psi_{n}(\bm{x}^{i}_{t-1},t) is designed to decode neighborhood distribution parameterized by feedforward neural networks (FNNs), and W2(⋅,⋅)\mathcal{W}_{2}(\cdot,\cdot) is the 2-Wasserstein distance. We provide a more explicit formulation of W22(ψn(xt−1i,t),PNis)\mathcal{W}_{2}^{2}(\psi_{n}(\bm{x}^{i}_{t-1},t),\mathcal{P}_{\mathcal{N}^{s}_{i}}) in Sec. 4.2.

2 Efficient Large Context Decoding

Our ConPreDiff essentially represents the node neighborhood H^Nis\bm{\hat{H}}_{\mathcal{N}^{s}_{i}} as a distribution of neighbors’ representations PNis\mathcal{P}_{\mathcal{N}^{s}_{i}} (Equation 14). In order to characterize the distribution reconstruction loss, we employ Wasserstein distance. This choice is motivated by the atomic non-zero measure supports of PNis\mathcal{P}_{\mathcal{N}^{s}_{i}} in a continuous space, rendering traditional ff-divergences like KL-divergence unsuitable. While Maximum Mean Discrepancy (MMD) could be an alternative, it requires the selection of a specific kernel function.

The decoded distribution ψn(xt−1i,t)\psi_{n}(\bm{x}^{i}_{t-1},t) is defined as an Feedforward Neural Network (FNN)-based transformation of a Gaussian distribution parameterized by xt−1i\bm{x}^{i}_{t-1} and tt. This selection is based on the universal approximation capability of FNNs, enabling the (approximate) reconstruction of any distributions within 1-Wasserstein distance, as formally stated in Theorem 4.1, proved in Lu & Lu . To enhance the empirical performance, our case adopts the 2-Wasserstein distance and an FNN with dd-dim output instead of the gradient of an FNN with 1-dim outout. Here, the reparameterization trick needs to be used:

Another challenge is that the Wasserstein distance between ψn(xt−1i,t)\psi_{n}(\bm{x}^{i}_{t-1},t) and PNis\mathcal{P}_{\mathcal{N}^{s}_{i}} does not have a closed form. Thus, we utilize the empirical Wasserstein distance that can provably approximate the population one as in Peyré et al. . For each forward pass, our ConPreDiff will get qq sampled target pixel/feature points {x(i,j)tar∣1≤j≤q}\{\bm{x}^{tar}_{(i,j)}|1\leq j\leq q\} from PNis\mathcal{P}_{\mathcal{N}^{s}_{i}}; Next, get qq samples from N(μi,Σi)\mathcal{N}(\mu_{i},\Sigma_{i}), denoted by ξ1,ξ2,...,ξq\xi_{1},\xi_{2},...,\xi_{q}, and thus {x(i,j)pred=FNNn(ξj)∣1≤j≤q}\{\bm{x}^{pred}_{(i,j)}=\text{FNN}_{n}(\xi_{j})|1\leq j\leq q\} are qq samples from the prediction ψn(xt−1i,t)\psi_{n}(\bm{x}^{i}_{t-1},t); Adopt the following empirical surrogated loss of W22(ψn(xt−1i,t),PNis)\mathcal{W}_{2}^{2}(\psi_{n}(\bm{x}^{i}_{t-1},t),\mathcal{P}_{\mathcal{N}^{s}_{i}}) in Equation 14:

The loss function is built upon solving a matching problem and requires the Hungarian algorithm with O(q3)O(q^{3}) complexity . A more efficient surrogate loss may be needed, such as Chamfer loss built upon greedy approximation or Sinkhorn loss built upon continuous relaxation , whose complexities are O(q2)O(q^{2}). In our study, as qq is set to a small constant, we use Equation 16 built upon a Hungarian matching and do not introduce much computational costs. The computational efficiency of design is empirically demonstrated in Sec. 5.3.

3 Discrete and Continuous ConPreDiff

In training process, given previously-estimated xt\bm{x}_{t}, our ConPreDiff simultaneously predict both xt−1\bm{x}_{t-1} and the neighborhood distribution PNis\mathcal{P}_{\mathcal{N}^{s}_{i}} around each pixel/feature. Because xt−1i\bm{x}^{i}_{t-1} can be pixel, feature or discrete token of input image, we can generalize the ConPreDiff to existing discrete and continuous backbones to form discrete and continuous ConPreDiff. More concretely, we can substitute the point denoising part in Equation 14 alternatively with the discrete diffusion term Lt−1dis\mathcal{L}^{dis}_{t-1} (Equation 3) or the continuous (Equation 6) diffusion term Lt−1con\mathcal{L}^{con}_{t-1} for generalization:

where λt∈\lambda_{t}\in is a time-dependent weight parameter. Note that our ConPreDiff only performs context prediction in training for optimizing the point denoising network ψp\psi_{p}, and thus does not introduce extra parameters to the inference stage, which is computationally efficient.

Equipped with our proposed context prediction term, existing diffusion models consistently gain performance promotion. Next, we use extensive experimental results to prove the effectiveness.

Experiments

Regarding unconditional image generation, we choose four popular datasets for evaluation: CelebA-HQ , FFHQ , LSUN-Church-outdoor , and LSUN-bedrooms . We evaluate the sample quality and their coverage of the data manifold using FID and Precision-and-Recall . For text-to-image generation, we train the model with LAION and some internal datasets, and conduct evaluations on MS-COCO dataset with zero-shot FID and CLIP score , which aim to assess the generation quality and resulting image-text alignment. For image inpainting, we choose CelebA-HQ and ImageNet for evaluations, and evaluate all 100 test images of the test datasets for the following masks: Wide, Narrow, Every Second Line, Half Image, Expand, and Super-Resolve. We report the commonly reported perceptual metric LPIPS , which is a learned distance metric based on the deep feature space.

Baselines

To demonstrate the effectiveness of ConPreDiff, we compare with the latest diffusion and non-diffusion models. Specifically, for unconditional image generation, we choose ImageBART, U-Net GAN (+aug) , UDM , StyleGAN , ProjectedGAN , DDPM and ADM for comparisons. As for text-to-image generation, we choose DM-GAN , DF-GAN , DM-GAN + CL , XMC-GAN LAFITE , Make-A-Scene , DALL-E , LDM , GLIDE , DALL-E 2 , Improved VQ-Diffusion , Imagen-3.4B , Parti , Muse , and eDiff-I for comparisons. For image inpainting, we choose autoregressive methods( DSI and ICT ), the GAN methods (DeepFillv2 , AOT , and LaMa ) and diffusion based model (RePaint ). All the reported results are collected from their published papers or reproduced by open source codes.

Implementation Details

For text-to-image generation, similar to Imagen , our continuous diffusion model \textscConPreDiffcon{\textsc{ConPreDiff}}_{con} consists of a base text-to-image diffusion model (64×\times64) , two super-resolution diffusion models to upsample the image, first 64×\times64 → 256×\times256, and then 256×\times256 → 1024×\times1024. The model is conditioned on both T5 and CLIP text embeddings. The T5 encoder is pre-trained on a C4 text-only corpus and the CLIP text encoder is trained on an image-text corpus with an image-text contrastive objective. We use the standard Adam optimizer with a learning rate of 0.0001, weight decay of 0.01, and a batch size of 1024 to optimize the base model and two super-resolution models on NVIDIA A100 GPUs, respectively, equipped with multi-scale training technique (6 image scales). We generalize our context prediction to discrete diffusion models to form our \textscConPreDiffdis{\textsc{ConPreDiff}}_{dis}. For image inpainting, we adopt a same pipeline as RePaint , and retrain its diffusion backbone with our context prediction loss. We use T = 250 time steps, and applied r = 10 times resampling with jumpy size j = 10. For unconditional generation tasks, we use the same denoising architecture like LDM for fair comparison. The max channels are 224, and we use T=2000 time steps, linear noise schedule and an initial learning rate of 0.000096. Our context prediction head contains two non-linear blocks (e.g., Conv-BN-ReLU, resnet block or transformer block), and its choice can be flexible according to specific task. The prediction head does not incur significant training costs, and can be removed in inference stage without introducing extra testing costs. We set the neighborhood stride to 3 for all experiments, and carefully choose the specific layer for adding context prediction head near the end of denoising networks.

2 Main Results

We conduct text-to-image generation on MS-COCO dataset, and quantitative comparison results are listed in Tab. 1. We observe that both discrete and continuous ConPreDiff substantially surpasses previous diffusion and non-diffusion models in terms of FID score, demonstrating the new state-of-the-art performance. Notably, our discrete and continuous ConPreDiff achieves an FID score of 6.67 and 6.21 which are better than the score of 8.44 and 7.27 achieved by previous SOTA discrete and continuous diffusion models. We visualize text-to-image generation results in Figure 2, and find that our ConPreDiff can synthesize images that are semantically better consistent with text prompts. It demonstrates our ConPreDiff can make promising cross-modal semantic understanding through preserving visual context information in diffusion model training. Moreover, we observe that ConPreDiff can synthesize complex objects and scenes consistent with text prompts as demonstrated by Figure 6 in Sec. A.3, proving the effectiveness of our designed neighborhood context prediction. Human evaluations are provided in Sec. A.4.

Image Inpainting

Our ConPreDiff naturally fits image inpainting task because we directly predict the neighborhood context of each pixel/feature in diffusion generation. We compare our ConPreDiff against state-of-the-art on standard mask distributions, commonly employed for benchmarking. As in Tab. 2, our ConPreDiff outperforms previous SOTA method for most kinds of masks. We also put some qualitative results in Figure 3, and observe that ConPreDiff produces a semantically meaningful filling, demonstrating the effectiveness of our context prediction.

Unconditional Image Synthesis

We list the quantitative results about unconditional image generation in Tab. 3 of Sec. A.2. We observe that our ConPreDiff significantly improves upon the state-of-the-art in FID and Precision-and-Recall scores on FFHQ and LSUN-Bedrooms datasets. The ConPreDiff obtains high perceptual quality superior to prior GANs and diffusion models, while maintaining a higher coverage of the data distribution as measured by recall.

3 The Impact and Efficiency of Context Prediction

In Sec. 4.2, we tackle the complexity problem by transforming the decoding target from entire neighborhood features to neighborhood distribution. Here we investigate both impact and efficiency of the proposed neighborhood context prediction. For fast experiment, we conduct ablation study with the diffusion backbone of LDM . As illustrated in Figure 5, the FID score of ConPreDiff is better with the neighbors of more strides and 1-stride neighbors contribute the most performance gain, revealing that preserving local context benefits the generation quality. Besides, we observe that increasing neighbor strides significantly increases the training cost when using feature decoding, while it has little impact on distribution decoding with comparable FID score. To demonstrate the generalization ability, we equip previous diffusion models with our context prediction head. From the results in Figure 5, we find that our context prediction can consistently and significantly improve the FID scores of these diffusion models, sufficiently demonstrating the effectiveness and extensibility of our method.

Conclusion

In this paper, we for the first time propose ConPreDiff to improve diffusion-based image synthesis with context prediction. We explicitly force each point to predict its neighborhood context with an efficient context decoder near the end of diffusion denoising blocks, and remove the decoder for inference. ConPreDiff can generalize to arbitrary discrete and continuous diffusion backbones and consistently improve them without extra parameters. We achieve new SOTA results on unconditional image generation, text-to-image generation and image inpainting tasks.

Acknowledgement

This work was supported by the National Natural Science Foundation of China (No.61832001 and U22B2037).

References

Appendix A Appendix

While our ConPreDiff boosts performance of both discrete and continuous diffusion models without introducing additional parameters in model inference, our models still have more trainable parameters than other types of generative models, e.g GANs. Furthermore, we note the long sampling times of both and compared to single step generative approaches like GANs or VAEs. However, this drawback is inherited from the underlying model class and is not a property of our context prediction approach. Neighborhood context decoding is fast and incurs negligible computational overhead in training stage. For future work, we will try to find more intrinsic information to preserve for improving existing point-wise denoising diffusion models, and extend to more challenging tasks like text-to-3D and text-to-video generation.

Broader Impact

Recent advancements in generative image models have opened up new avenues for creative applications and autonomous media creation. However, these technologies also pose dual-use concerns, raising the potential for negative implications. In the context of our research, we strictly utilize human face datasets solely for evaluating the image inpainting performance of our method. It is important to clarify that our approach is not designed to generate content for the purpose of misleading or deceiving individuals. Despite our intentions, similar to other image generation methods, there exists a risk of potential misuse, particularly in the realm of human impersonation. Notorious examples, such as "deep fakes," have been employed for inappropriate applications, such as creating pornographic "undressing" content. We vehemently disapprove of any actions aimed at producing deceptive or harmful content featuring real individuals. Moreover, generative methods, including ours, have the capacity to be exploited for malicious intentions, such as harassment and the dissemination of misinformation . These possibilities raise significant concerns related to societal and cultural exclusion, as well as biases in the generated content . In light of these considerations, we have chosen not to release the source code or a public demo at this point in time.

Furthermore, the immediate availability of mass-produced high-quality images carries the risk of spreading misinformation and spam, contributing to targeted manipulation in social media. Deep learning heavily relies on datasets as the primary source of information, with text-to-image models requiring large-scale data . Researchers often resort to large, mostly uncurated, web-scraped datasets to meet these demands, leading to rapid algorithmic advances. However, ethical concerns surround datasets of this nature, prompting a need for careful curation to exclude or explicitly contain potentially harmful source images. Consideration of the ability to curate databases is crucial, offering the potential to exclude or contain harmful content. Alternatively, providing a public API may offer a cost-effective solution to deploy a safe model without retraining on a filtered subset of the data or engaging in complex prompt engineering. It is essential to recognize that including only harmful content during training can easily result in the development of a toxic model.

A.2 More Quantitative Results

We list the unconditional generation results on FFHQ, CelebA-HQ, LSUN-Churches, and LSUN-Bedrooms in Tab. 3. We find ConPreDiff consistently outperforms previous methods, demonstrating the effectiveness of the ConPreDiff.

A.3 More Synthesis Results

We visualize more text-to-image synthesis results on MS-COCO dataset in Figure 6. We observe that compared with previous powerful LDM and DALL-E 2, our ConPreDiff generates more natural and smooth images that preserve local continuity.

A.4 Human Evaluations

As demonstrated in qualitative results, our ConPreDiff is able to synthesize realistic diverse, context-coherent images. However, using FID to estimate the sample quality is not always consistent with human judgment. Therefore, we follow the protocol of previous works , and conduct systematic human evaluations to better assess the generation capacities of our ConPreDiff from the aspects of image photorealism and image-text alignment. We conduct side-by-side human evaluations, in which well-trained users are presented with two generated images for the same prompt and need to choose which image is of higher quality and more realistic (image photorealism) and which image better matches the input prompt (image-text alignment). For evaluating the coherence of local context, we propose a new evaluation protocol, in which users are presented with 1000 pairs of images and must choose which image better preserves local pixel/semantic continuity. The evaluation results are in Tab. 4, ConPreDiff performs better in pairwise comparisons against both Improved VQ-Diffusion and Imagen. We find that ConPreDiff is preferred in terms of all three evaluations, and ConPreDiff is strongly preferred regarding context coherence, demonstrating that preserving local neighborhood context advances sample quality and semantic alignment.