SinDiffusion: Learning a Diffusion Model from a Single Natural Image

Weilun Wang, Jianmin Bao, Wengang Zhou, Dongdong Chen, Dong Chen, Lu Yuan, Houqiang Li

Introduction

Generating images from a single natural image has extracted more and more attention due to its various applications. This task aims to learn an unconditional generative model from a single natural image to generate diverse samples with similar visual content by capturing the internal statistics of patches. Once trained, the generative model can not only produce high-quality diverse images of arbitrary resolutions but also easily adapt to multiple applications, i.e., image editing, image harmonization, and image-to-image translation.

The groundbreaking method of this task is SinGAN , which builds multiple scales of the natural image and trains a series of GANs to learn the internal statistics of patches in a single image. The core idea of SinGAN is to train multiple models at progressive growing scales. This becomes the default setting of this direction. However, we may observe these methods generate unsatisfactory images. This is because these methods accumulate the small detail errors produced at small scales, which then lead to obvious characteristic artifacts in resulted images (see Figure 2).

In this paper, we propose a novel framework called Single-image Diffusion Model (SinDiffusion) for learning from a single natural image. SinDiffusion is based on the recently developed Denoising Diffusion Probabilistic Model (DDPM). We find that multiple models at progressive growing scales are not essential for learning from a single image, A single diffusion-based model trained on a single scale is well-suited for this task. Although the diffusion model is a multiple-step generation process, it doesn’t suffer from the issue of accumulated errors. The reason is that the diffusion model has systemically mathematical formulations and the error in intermediate steps can be regarded as noise which could be refined during the diffusion process.

The other core design of SinDiffusion is to restrict the receptive field of the diffusion model. We revisit the commonly-used network structure in the previous diffusion model and find it enjoys a strong capability with a deep structure. The network structure has a large receptive field that can cover the full image, which leads to the model tending to memorize the training image and thus generating exactly the same image as the training image. To encourage the model to learn the patch statistics rather than memorize the whole image, we carefully design the network structure and introduce a patch-wise denoising network. Compared with the previous diffusion structure, SinDiffusion reduces the downsampling times and the number of resblocks in the original denoising network structure. Equipped with this design, SinDiffusion produces high-quality and diverse images by learning from a single natural image (see Figure 2).

Our proposed SinDiffusion enjoys the advantages of flexibility for various applications (see Figure 1). It can be used for various applications without any re-training of the model. In SinGAN, the downstream applications are mainly implemented by feeding the condition into the different scales of pre-trained GANs. Therefore, the applications of SinGAN are limited to those given “spatially-aligned” conditions. Different from that, SinDiffusion is available for a wider range of applications by designing the sampling procedure. SinDiffusion learns to predict the gradient of the data distribution through unconditional training. Supposing a score function (i.e., L−pL-p distance or a pre-trained network like CLIP ) which describes the relevance between the generated image and condition, we utilize the gradient of the relevance score to guide the sampling procedure of SinDiffusion. In this way, SinDiffusion is able to generate images that both fit the data distribution and are corresponding to the given conditions.

To demonstrate the superiority of our proposed framework, we conduct experiments on a variety of natural images, including landscapes and famous art. Both quantitative and qualitative results validate that SinDiffusion can generate both high-fidelity and diverse results. Downstream applications further demonstrate the usefulness and flexibility of our SinDiffusion.

Overall, the contributions are summarized as follows,

We propose a novel diffusion-based framework named SinDiffusion for capturing the internal statistics of patches from a single natural image.

We introduce two key ingredients in SinDiffusion: single-scale training and a new network with patch-level receptive fields, techniques essential for generating high-quality and diverse images.

We explore more downstream applications by leveraging the advantages of SinDiffusion, including text-guided image generation, image outpainting, and etc.

Extensive experiments on various natural images, i.e., landscapes and famous arts, demonstrate the effectiveness and wide applicability of our framework.

Related Work

In this section, we will briefly revisit the related topic, including denoising diffusion probabilistic models and single image generation.

As a new class of generative model, denoising diffusion probabilistic models have achieved remarkable success on various tasks compared with generative adversarial nets (GANs) . Diffusion model is a parameterized Markov chain that optimizes the lower variational bound on the likelihood function to generate samples matching real distribution. Ho et. al. first propose diffusion model and Dhariwal and Nichol further show the potential of diffusion models, achieving better image sample quality compared with other generative models, i.e., GAN, on ImageNet dataset . After that, more and more researchers turn their attention to diffusion models . Saharia et. al. achieve success in super-resolution with diffusion models. Pattle explores diffusion models on four image-to-image translation problems, i.e., colorization, inpainting, uncropping, and JPEG decompression. And two concurrent works apply diffusion models for text-to-image generation problem.

The above approaches deal with conditional image generation by directly training on corresponding datasets. Diffusion models are also capable of solving conditional image generation problem by taking advantage of pre-trained unconditional models. Bahjat et. al. propose an unsupervised posterior sampling method, i.e., DDRM, to solve any linear inverse problem, e.g., image inpainting and colorization, with a pre-trained diffusion model. ILVR guides the generative process in a diffusion model to generate high-quality images based on a given reference image. DDIBs and EGSDE apply pre-trained diffusion models to unpaired image-to-image translation task. And introduces multi-modal information like text as the guidance of the generative process to generate text-related images. With the characteristic of diffusion models, our SinDiffusion can also generate condition-related images to solve a variety of image manipulation tasks.

2 Single Image Generation

Single image generation aims to generate diverse results by learning the internal patch distribution from a single image. The groundbreaking work on this problem is SinGAN which first explores this problem and proposes a series of PatchGAN based on an image pyramid to generate diverse results hierarchically. A concurrent work, i.e., InGAN , train a conditional GAN to solve the same problem based on a geometry transformation. After that, more and more researchers turn their attention to this emerging topic. ExSinGAN train three modular GANs to model the distributions about structure, semantics and texture for learning an explainable generative model. ConSinGAN improves SinGAN by concurrently training several stages in a sequential multi-stage manner. However, these methods are based on multi-scale GAN structure in SinGAN, which may accumulate errors and lead to unsatisfactory generation results with characteristic artifacts. To this end, we propose a framework based on diffusion model to generate photorealistc and diverse results from a single natural image.

Methodology

In this paper, we present a novel framework named SinDiffusion to learn internal distribution from a single natural image. Unlike the progressive growing design in prior work, SinDiffusion is trained with a single denoising model at a single scale, which prevents the accumulation of errors. Furthermore, we identify that a patch-level receptive field of the diffusion network plays an important role in capturing internal patch distribution and design a new denoising network structure. Based on these two core designs, SinDiffusion generates high-quality and diverse images from a single natural image. The rest of this section is organized as follows: We begin with revisiting SinGAN and presenting the motivation of SinDiffusion. After that, we introduce the structure design of SinDiffusion.

We begin with briefly revisiting SinGAN. Figure 3(a) presents the generation procedure of SinGAN. To generate diverse images from a single image, one key design of SinGAN is to build an image pyramid and progressively grow the resolution of the generated image. At each scale nn, the image from the previous scale x~n+1\widetilde{x}_{n+1} is upsampled and fed into PatchGAN along with the input noise map znz_{n} to generate the image at scale nn, which is formulated as follow,

where αn\alpha_{n} is a blending factor that decreases as nn decreases.

Suppose there are small detail errors in the output image at a small scale. With the output resolution growing, these detail errors will be accumulated in the output image, thus causing an unsatisfactory final output. Through this analysis, we find the recently developed denoising diffusion probabilistic models (DDPMs) do not suffer from this issue since they naturally process detail errors during the diffusion process.

To this end, we propose a novel framework named SinDiffusion in Figure 3(b). Different from SinGAN, SinDiffusion performs the multiple-step generation process with a single denoising network at a single scale. Although SinDiffusion also adopts the multiple-step generation process like SinGAN, the generated results are high-quality. This is because the diffusion model is based on a systematic derivation of mathematical equations, and errors arising from intermediate steps are repeatedly refined as noise during the diffusion process.

2 SinDiffusion

SinDiffusion is a diffusion-based model trained on a single scale to learn the internal patch statistics. We first revisit the commonly-used diffusion model and find that they tend to generate exactly the same images as the training image. This is because the denoising network possesses strong capability with a large receptive field that covers the full image, which leads to the network memorizing the training image instead of learning the internal patch distribution. Motivated by this, we suppose that the receptive field plays an important role in learning the internal patch statistics, and a patch-level receptive field of the diffusion network is crucial and effective for single image generation.

We study the relationship between generation diversity and the receptive field of the denoising network. The receptive field is varied by modifying the network structure of the denoising network. We design four network structures with different receptive fields but comparable capabilities and train these models on a single natural image. Figure 4 shows the results generated by the model with different receptive fields. It is observed that, with a smaller receptive field, SinDiffusion tends to generate more diverse generation results and vice versa. However, we find that the model of an extremely small receptive field can not preserve the reasonable structure of the image. Therefore, a suitable receptive field is important and necessary to capture reasonable patch statistics.

Based on the above analysis, we redesign the commonly-used diffusion model and introduce a patch-wise denoising network for single image generation. Figure 5 presents an overview of patch-wise denoising network in SinDiffusion and the main difference compared with the previous denoising network. First, we reduce the depth of the denoising network by lessening the downsample and upsample operation, which greatly expands the receptive field. Meanwhile, the attention layer originally used in the deep layers of the denoising network is naturally removed, which makes SinDiffusion a fully convolutional network applicable to the generation of arbitrary resolution. Second, we further limit the receptive field of SinDiffusion by reducing the time-embedded resblocks in each resolution. In this way, we draw a patch-wise denoising network with a proper receptive field, generating photorealistic and diverse results.

We train our SinDiffusion with the original denoising loss. In diffusion models, given a training image xx and a random timestep t∈{0,1,…,T}t\in\{0,1,\dots,T\}, a noisy version of the image x~\widetilde{x} is produced as follows,

where ϵ\epsilon is a noise sampled from the standard Gaussian distribution. αt\alpha_{t} is a noise scheduler at timestep tt. TT is set to 1000 in our SinDiffusion. SinDiffusion is trained to reconstruct the training image xx by predicting the involved noise ϵ\epsilon with the timestep tt, which is formulated as follows,

Once trained, SinDiffusion can generate diverse images by an iterative denoising process, which is formulated as follows,

where αt\alpha_{t} and βt\beta_{t} is the variance schedule factor in diffusion model. zt\mathbf{z}_{t} denotes a Gaussian noise involved at timestep tt.

Experiments

Datasets. To evaluate the effectiveness of our method, we conduct experiments on various natural images collected online, including landscapes and arts. Additionally, to systematically evaluate the quantitative performance, we also experiment on a dataset of natural landscapes, i.e., Places50. Places50 is the set of 50 landscapes image used in SinGAN (50 images from Places365 dataset ). We train each SinDiffusion model on each image in Places50 and evaluate the fidelity and diversity of generated results.

Implementation details. We train our SinDiffusion with AdamW optimizer . During training, we adopt an exponential moving average (EMA) with 0.9999 decay. The whole framework is implemented by Pytorch and the experiments are performed on NVIDIA Tesla V100.

Evaluation metric. We aim to assess both visual quality and diversity of generated images. For the visual quality, following SinGAN , we adopt the single-image Frechet Inception Distance (SIFID) metric. Similar to FID, SIFID measures the deviation between the distribution of patch-wise features from the generated images and the real images. To evaluate the generation diversity, we compute the average distance measured by the LPIPS metrics between multimodal generation results.

2 Qualitative Evaluation

Qualitative results of random generated images from SinDiffusion are shown in Figure 6, and more qualitative results are included in Supplementary Material. We train our SinDiffusion on the image of natural images and famous arts. For each training image, we first present a generated image under the same aspect ratio. Then, we generate images of different aspect ratios with the training image to demonstrate the generalization of our SinDiffusion at different resolutions. It is observed that, for different resolutions, our SinDiffusion can generate realistic images which have similar patterns to the training image.

Furthermore, we explore SinDiffusion for generating high-resolution images from a single image. Figure 13 presents the training image and generated result. The training image is a landscape image of 486×741486\times 741 resolution which contains rich components, i.e., clouds, mountains, grass, flowers, and a lake. To accommodate the high-resolution image generation, we extend the SinDiffusion to an enhanced version, which has larger receptive fields and network capability. Compared with the structure in Figure 3, the enhanced version has 4 downsample layers and an additional time-embedded resblock on each scale. With the enhanced SinDiffusion, we generate a high-resolution long-scroll image of 486×2048486\times 2048 resolution. From Figure 13, it is observed that our result maintains the internal layout of the training image and generalize new content.

3 Comparison with previous methods

We compare our SinDiffusion with several challenging methods, i.e., SinGAN , ExSinGAN , ConSinGAN and GPNN . The quantitative results in shown in Table 1. With the help of progressive refinement, SinDiffusion achieves state-of-the-art performance compared with previous GAN-based method. Notably, our method highly improves the diversity of generated images, surpassing the most challenging method by +0.082 LPIPS score on the average of 50 models trained on the Places50 dataset.

Besides the quantitative results, we also present the qualitative results on the Places50 dataset in Figure 8. The images generated by SinGAN, ExSinGAN and ConSinGAN show unreasonable structure and artifact in the details. This is because of the error accumulation in the multi-scale structure, enlarging the artifact produced at the initial several scales. GPNN is able to produce realistic images. However, since GPNN clones the nearest patches from the training image, its generated images lose patch-level diversity and tend to be similar to the training images. Unlike these methods, the images generated by SinDiffusion are with reasonable structures and sharp details and are capable of generalizing novel patterns from the training images.

Furthermore, we conduct a user study to evaluate the visual performance of generated images. There are 20 volunteers participating in this study. In the study, we present each volunteer with 10 pairs of generated results for each paired user study (40 pairs in total). Volunteers are asked to answer this question, i.e., which group of images shows more diversity and better quality? The voting results are reported in Table. 2. It is observed that our method is clearly preferred over competitors in more than 65% of the time.

4 Image Manipulation

We explore the application of SinDiffusion on various image manipulation tasks. We directly use our trained SinDiffusion model for all the applications, without architectural changes or further finetuning. Different from SinGAN, injecting the condition image into the generation pyramid at some scale, SinDiffusion is utilized for various applications by designing the sampling procedure. By virtue of this, besides the applications in SinGAN, i.e., image editing, image harmonization and image-to-image translation, SinDiffusion can be further applied to image manipulation tasks like text-guided image generation and image outpainting. We will present the details as follow.

Text-guided image generation. To generate the images from a single image corresponding to the given text, we guide the sampling procedure by a gradient from a pre-trained visual-linguistic model C(⋅,⋅)C(\cdot,\cdot), i.e., CLIP. Supposing a pre-trained diffusion model with estimated mean μθ(xt−1∣xt)\mu_{\theta}(x_{t-1}|x_{t}), an image that corresponds to a given text LL can be generated by perturbing the mean, which is formulated as follows,

where the hyperparameter ss is the guidance scale, which balances the fidelity and correspondence with the given text.

Figure 15 presents the text-guided image generation results of SinDiffusion and previous methods. We train SinDiffusion on various images and use text as the condition to generate images with a different number of objects or different shapes from a single image. From the figure, it is observed that, by changing the conditional text, SinDiffusion is able to controllably generate realistic images from the training image. By comparison, previous methods fail to generate corresponding images with text under the single-image setting. This demonstrates that we provide an effective approach to control the single-image model through high-level semantics.

Image outpainting. Image outpainting aims to generate content which resides beyond the edges of an image. With iterative image outpainting, ideally, we can extend a finite-sized image to an infinite size. Since our SinDiffusion model learns the internal distribution of the patches from the training image, it is inherently capable of imagining the content outside the given image. Supposing a pre-trained diffusion model with iterative latent xθ(zt)x_{\theta}(z_{t}), we outpaint a natural image xax^{a} by replacing the given region, which is formulated as follows,

where mam^{a} indicates the outpainting region. xt−1ax^{a}_{t-1} refers to the noisy version of the natural image xax^{a} at timestep t−1t-1.

In Figure 10, we compare SinDiffusion with some previous image outpainting methods, i.e., DeepFillv2 , Boundless and InfinityGAN . From the figure, it is observed that, by learning the patch distribution, SinDiffusion generates reasonable and realistic images with content that conforms to the intrinsic distribution. In contrast, previous methods produce unrealistic and blurry results, failing to predict what is outside the original image. This indicates that SinDiffusion is more efficient and flexible for image outpainting of a single natural image.

Other image manipulation task. SinDiffusion can also be applied to image manipulation tasks in prior methods, i.e., image editing, image harmonization and image-to-image translation. Inspired by , we generate from a reference image yy by designing the sampling procedure as follows,

where ϕN(⋅)\phi_{N}(\cdot) refers to a linear low-pass filtering operation. yt−1y_{t-1} refers to the noisy version of the reference image yy at timestep t−1t-1. Some image manipulation results from SinDiffusion are shown in Figure 1, and more results are included in Supplementary Material.

5 Ablation Study

We conduct ablative experiments to evaluate the effectiveness of several important designs in SinDiffusion, i.e., whether to utilize the multi-scale structure and the receptive field of SinDiffusion. We perform experiments on a subset of the Places50 dataset.

Multi-scale v.s. Single-scale. Different from previous multi-scale approaches, we introduce a diffusion model to solve the problem on a single scale. To verify the efficiency of single-scale design, we design a baseline variant as the comparison. As an alternative, we convert the GAN in each scale of SinGAN to a diffusion model, and thereby present a multi-scale diffusion model on single image generation. From Table 3, it can be seen that SinDiffusion achieves superior performance to the multi-scale diffusion model. Meanwhile, compared with the multi-scale diffusion model, SinDiffusion has much smaller network parameters and computational consumption, which also indicates the advantage of the single-scale design.

Receptive field. In Section 3.2, we have shown some examples on how the receptive field affects the generated results. We will further perform more complete analysis here. The quantitative results are reported in Table 4. From the table, it is observed that SinDiffusion generates more diverse images with a smaller receptive field. However, the fidelity of the generated image (SIFID) increases as the receptive field increases. Therefore, We take a suitable receptive field to trade off generation quality and diversity.

Conclusion

In this paper, we present the first attempt to explore the diffusion model on single image generation and propose a novel framework named Single-image Diffusion Model (SinDiffusion). In particular, we find that the receptive field plays an important role in diverse image generation and design a patch-wise denoising network for producing realistic and diverse images. Furthermore, with the trained SinDiffusion model, we study a variety of image manipulation tasks, i.e., text-guided image generation, and image outpainting. Extensive experiments on various natural images and the Places50 dataset demonstrate the effectiveness of our method. Our method achieves state-of-the-art performance in terms of SIFID and LPIPS metrics and shows a better visual quality of generated images compared with previous methods. The performance on image manipulation further demonstrates the usefulness and flexibility of SinDiffusion.

References

A. Implementation Details

In this section, we provide more implementation details on SinDiffusion, including details on the denoising network and diffusion procedures.

Denoising network. In Section 3.2, we introduce the patch-wise denoising network in SinDiffusion, which is a U-Net-based network estimating the noise in the input noisy image. The encoder and decoder of the denoising network are 3 stages and the spatial resolution of the feature on each layer is 1, 1/2, and 1/4 of the input resolution, respectively. The channel of each layer is set to 64, 128, and 256, respectively. Each encoder layer contains 1 time-embedded Resblocks while each decoder layer contains 2 time-embedded Resblocks. In high-resolution single image generation in Section 4.2, we utilize an enhanced denoising network, whose encoder and decoder are 4 layers. The channel of each layer is set to 64, 128, 256, and 512, respectively. Each encoder layer contains 2 time-embedded Resblocks while each decoder layer contains 3 time-embedded Resblocks. In addition, we use half-precision float computation to accelerate training and reduce memory consumption.

Diffusion procedure. As mentioned in Section 3, we propose a single-image diffusion model to learn the internal distribution from a single natural image. Following DDPM , we set the total diffusion timestep TT to 1000. In the forward process, the Gaussian noise is involved in the data according to a variance schedule β1,…,βT\beta_{1},\dots,\beta_{T}. In our implementation, the variance schedule is arranged linearly with respect to the timestep tt. We adopt standard diffusion sampling in the diffusion process, introducing noise at each step and refining the generated image iteratively in 1000 steps.

Evaluation metrics. As mentioned in Section 4.1, we apply SIFID and LPIPS metrics to evaluate the quality and diversity of generated images, respectively. For SIFID metrics, we follow SinGAN and use deep features at the output of the convolutional layer just before the second pooling layer in VGG19 network. SIFID measures the distance between the patch distribution of two images in the feature space using a pre-trained Inception model, which is formulated as follows,

where (μX\mu_{\mathbf{X}}, ΣX\Sigma_{\mathbf{X}}) and (μY\mu_{\mathbf{Y}}, ΣY\Sigma_{\mathbf{Y}}) refer to the mean value and covariance of patch distribution of the generated and training image, respectively.

For the LPIPS metric, we generate a set of images xN={xi,i=1,2,…,N}x^{N}=\{x_{i},i=1,2,\dots,N\} by multimodal generation. The diversity metric, i.e., LPIPS, is formulated as follows,

where d(⋅,⋅)d(\cdot,\cdot) is a weighted perceptual similarity between two images, computed by the features extracted from a pre-trained AlexNet. The number of images NN is set to 10 in our implementation.

B. Additional Experiment Results

In this section, we first present more qualitative results trained on single natural images. Then, we supplement more single image manipulation task which is not shown in the main paper due to space limitations.

We present more qualitative results trained on single natural images with SinDiffusion. Figure 11 shows generated images under the same resolution as the training image. Figure 12 shows generated images of arbitrary resolutions. Figure 13 shows high-resolution generated images. Figure 14 shows more comparison results with previous methods, i.e., SinGAN, ExSinGAN, ConSinGAN, and GPNN.

B.2 Image Manipulation

In this subsection, we present more cases on image manipulation task, i.e., text-guided image generation, image outpainting and paint-to-image translation.

Text-guided image generation. As shown in Section 4.4, we explore text-guided image generation. We train a SinDiffusion on a single natural image and generalize from the training image according to the given text. We further show some text-guided image generation results in Figure 15.

Image outpainting. As mentioned in Section 4.4, image outpainting aims to generate content beyond the edges of an image. We train a SinDiffusion on a single natural image and generate what is outside the training image by replacing the given region during the sampling procedure. We show more qualitative results in Figure 16.

Paint-to-Image Translation. The paint-to-image translation is an image-to-image translation task that aims to convert a roughly-drawing image into a photorealistic image. Figure 17 shows the paint-to-image generated results of different methods. It is observed that SinDiffusion generates more realistic and reasonable images from the paint images compared with previous methods.