Wavelet Diffusion Models are fast and scalable Image Generators

Hao Phung, Quan Dao, Anh Tran

Introduction

Despite being introduced recently, diffusion models have grown tremendously and drawn many research interests. Such models revert the diffusion process to generate clean, high-quality outputs from random noise inputs. These techniques are applied in various data domains and applications but show the most remarkable success in image-generation tasks. Diffusion models can beat the state-of-the-art generative adversarial networks (GANs) in generation quality on various datasets . More notably, diffusion models provide a better mode coverage and a flexible way to handle different types of conditional inputs such as semantic maps, text, representations, and images . Thanks to this capability, they offer various applications such as text-to-image generation, image-to-image translation, image inpainting, image restoration, and more. Recent diffusion-based text-to-image generative models allow users to generate unbelievably realistic images just by text inputs, opening a new era of AI-based digital art and promising applications to various other domains.

While showing great potential, diffusion models have a very slow running speed, a critical weakness blocking them from being widely adopted like GANs. The foundation work Denoising Diffusion Probabilistic Models (DDPMs) requires a thousand sampling steps to produce the desired output quality, taking minutes to generate a single image. Many techniques have been proposed to reduce the inference time , mainly via reducing the sampling steps. However, the fastest algorithm before DiffusionGAN still takes seconds to produce a 32×\times32 image, which is about 100 times slower than GAN. DiffusionGAN made a break-though in fastening inference speed by combining Diffusion and GANs in a single system, which ultimately reduces the sampling steps to 4 and the inference time to generate a 32×\times32 image as a fraction of a second. It makes DiffusionGAN the fastest existing diffusion model. Still, it is at least 4 times slower than the StyleGAN counterpart, and the speed gap consistently grows when increasing the output resolution. Moreover, DiffusionGAN still requires a long training time and a slow convergence, confirming that diffusion models are not yet ready for large-scale or real-time applications.

This paper aims to bridge the speed gap by introducing a novel wavelet-based diffusion scheme. Our solution relies on discrete wavelet transform, which decomposes each input into four sub-bands for low- (LL) and high-frequency (LH, HL, HH) components. We apply that transform on both image and feature levels. This allows us to significantly reduce both training and inference times while keeping the output quality relatively unchanged. On the image level, we obtain a high speed boost by reducing the spatial resolution four times. On the feature level, we stress the importance of wavelet information on different blocks of the generator. With such a design, we can obtain considerable performance improvement while inducing only a marginal computing overhead.

Our proposed Wavelet Diffusion provides state-of-the-art training and inference speed while maintaining high generative quality, thoroughly confirmed via experiments on standard benchmarks including CIFAR-10, STL-10, CelebA-HQ, and LSUN-Church. Our models significantly reduce the speed gap between diffusion models and GANs, targeting large-scale and real-time systems.

In summary, our contributions are as following:

We propose a novel Wavelet Diffusion framework that takes advantage of the dimensional reduction of Wavelet subbands to accelerate Diffusion Models while maintaining good visual quality of generated results through high-frequency components.

We employ wavelet decomposition in both image and feature space to improve generative models’ robustness and execution speed.

Our proposed Wavelet Diffusion provides state-of-the-art training and inference speed, which serves as a stepping-stone to facilitating real-time and high-fidelity diffusion models.

Related work

Diffusion models are inspired by non-equilibrium thermodynamics where the diffusion and reverse processes are Markov chains. Unlike GANs and VAEs , they sample using a large number of denoising steps with a fixed procedure where each latent variable shares the same dimensionality as the original inputs. A line of methods shares the same motives but relies on score matching for the reverse process . Later works focus on improving sample quality of diffusion models . Some have shifted the process to latent space , which enables the success of text-to-image generation . Similar to us, some approaches aim to improve the efficiency of sampling process for better convergence and faster sampling. As the Markov process is a sequential probability model with a finite step, some try to break it into non-Markov chains for faster sampling . Despite several efforts to accelerate the sampling process, there is an inevitable trade-off between sampling speed and quality.

Diffusion GAN was recently proposed as a novel approach to sidestep the unimodal Gaussian assumption of small step sizes via modeling the complex multimodal distribution of large step sizes with generative adversarial networks. This proposal greatly reduces the number of denoising steps to just a finger count (e.g., 2, 4), which leads the inference time to a fraction of a second. However, it is still largely slower than GAN competitors. Hence, we further maximize its full potential by introducing new wavelet components on top of this framework.

2 Wavelet-based approaches

Wavelet decomposition is a classical operator widely used in various computer vision tasks. It is the fundamental process behind a popular JPEG-2000 image compression format . In recent years, wavelet transform has started being incorporated into many deep-learning-based systems as they can take advantage of spatial and frequency information to aid the training. Some have applied it to the design of networks to improve visual representation learning . Some have extended it to image restoration , style transfer , and face problems . On the line of image generation, wavelet transform is plugged in recent generative adversarial networks to improve the visual quality of output images.

Regarding diffusion model families, several works start employing wavelet transform. Guth et al. accelerate score-based models by modeling conditional probabilities of multiscale wavelet coefficients, but still suffer from low-quality output images. Li et al. perform score matching on the wavelet domain for better image colorization.

However, none of the mentioned techniques can balance the diffusion models’ quality and speed. Here, we introduce a novel wavelet diffusion approach to not only emphasize the importance of frequency information but also to reduce the sampling time by a large margin. We take advantage of frequency sparsity and dimensionality reduction of wavelet transform to reduce high-dimensional space to the lower-dimensional manifolds, which are essential for faster and more efficient sampling. Unlike existing methods on diffusion families, we enrich the frequency awareness on both image and feature levels to facilitate high-fidelity image generation. This is greatly inspired by the success of wavelet-based GANs.

Background

The traditional diffusion process requires thousand timesteps to gradually diffuse each input x0x_{0}, following a data distribution p(x0)p(x_{0}), into pure Gaussian noise. The posterior probability of a diffused image xtx_{t} at timestep tt has a closed form:

where αt=1−βt\alpha_{t}=1-\beta_{t}, αtˉ=∏s=1tαs\bar{\alpha_{t}}=\prod_{s=1}^{t}\alpha_{s}, and βt∈(0,1)\beta_{t}\in(0,1) is defined to be small through a variance schedule which can be learnable or fixed at timestep tt in the forward process.

Since the diffusion process adds relatively small noise per step, the reverse process q(xt−1∣xt)q(x_{t-1}|x_{t}) can be approximated by Gaussian process q(xt−1∣xt,x0)q(x_{t-1}|x_{t},x_{0}). Therefore, the trained denoising process pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) can be parameterized according to q(xt−1∣xt,x0)q(x_{t-1}|x_{t},x_{0}). The common parameterized form of pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) is:

where μθ(xt,t)\mu_{\theta}(x_{t},t) and σt2\sigma^{2}_{t} are the mean and variance of the parametric denoising model, respectively. The objective is to minimize the distance between a true denoising distribution q(xt−1∣xt)q(x_{t-1}|x_{t}) and the parameterized one pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) through Kullback-Leibler (KL) divergence.

Unlike the traditional diffusion methods, DiffusionGAN enables large step sizes for faster sampling through generative adversarial networks. It introduces a discriminator DϕD_{\phi} and optimizes both the generator and the discriminator in an adversarial training manner:

where fake samples are sampled from a conditional generator pθ(xt−1∣xt)p_{\theta}\left(\mathbf{x}_{t-1}\mid\mathbf{x}_{t}\right). As large step sizes cause q(xt−1∣xt)q(x_{t-1}|x_{t}) to be no longer a Gaussian distribution, DiffusionGAN aims to implicitly model this complex multimodal distribution with a generator Gθ(xt,z,t)G_{\theta}(x_{t},z,t) given a DD-dimensional latent variable z∼N(0,I)z\sim\mathcal{N}(0,\mathbf{I}). Specifically, DiffusionGAN first generates unperturbed sample x0′x_{0}^{\prime} through the generator Gθ(xt,z,t)G_{\theta}(x_{t},z,t) and acquires the corresponding perturbed sample xt−1′x_{t-1}^{\prime} using q(xt−1∣xt,x0)q(x_{t-1}|x_{t},x_{0}). Meanwhile, the discriminator performs judgment on real pairs Dϕ(xt−1,xt,t)D_{\phi}(x_{t-1},x_{t},t) and fake pairs Dϕ(xt−1′,xt,t)D_{\phi}(x_{t-1}^{\prime},x_{t},t).

For convenience, we will abbreviate DiffusionGAN as DDGAN in later sections.

2 Wavelet Transform

Wavelet Transform is a classical technique widely used in image compression to separate the low-frequency approximation and the high-frequency details from the original image. While low subbands are similar to down-sampled versions of the original image, high subbands express the local statistics of vertical, horizontal, and diagonal edges. Notably, the Haar wavelet is widely adopted in real-world applications due to its simplicity. It involves two types of operations: discrete wavelet transform (DWT) and discrete inverse wavelet transform (IWT).

In this paper, we use this kind of transform to decompose input images and feature maps to emphasize high-frequency components and reduce the spatial dimensions to four folds for more efficient sampling.

Method

This section describes our proposed Wavelet Diffusion framework. First, we present the core wavelet-based diffusion scheme for more efficient sampling (Sec. 4.1). We then depict the design of a new wavelet-embedded generator for better frequency-aware image generation (Sec. 4.2).

First, we describe how to incorporate wavelet transform in the diffusion process. We decompose the input image into four wavelet subbands and concatenate them as a single target for the denoising process (illustrated in Fig. 2). Such a model does not perform on the original image space but on the wavelet spectrum. As a result, our model can leverage high-frequency information to increase the details of generated images further. Meanwhile, the spatial area of wavelet subbands is four times smaller than the original image, so the computational complexity of the sampling process is significantly reduced.

Let denote y0y_{0} be a clean sample and yty_{t} be a corrupted sample at timestep tt which is sampled from q(yt∣y0)q(y_{t}|y_{0}). In terms of the denoising process, a generator receives a tuple of variable yty_{t}, a latent z∼N(0,I)z\sim\mathcal{N}(0,\mathbf{I}), and a timestep tt to generate an approximation of original signals y0y_{0}: y0′=G(yt,z,t)y_{0}^{\prime}=G(y_{t},z,t). The predicted noisy sample yt−1′y^{\prime}_{t-1} is then drawn from tractable posterior distribution q(yt−1∣yt,y0′)q(y_{t-1}|y_{t},y^{\prime}_{0}). The role of the discriminator is to distinguish the real pairs (yt−1,yt)(y_{t-1},y_{t}) and the fake pairs (yt−1′,yt)(y^{\prime}_{t-1},y_{t}).

Adversarial objective Following, we optimize the generator and the discriminator through the adversarial loss:

Reconstruction term In addition to the adversarial objective in Eq. 4, we add a reconstruction term to not only impede the loss of frequency information but also preserve the consistency of wavelet subbands. It is formulated as an L1 loss between a generated image and its ground-truth:

The overall objective of the generator is a linear combination of adversarial loss and reconstruction loss:

where λ\lambda is a weighting hyper-parameter (default value is 1).

After a few sampling steps as defined, we acquire the estimated denoised subbands y0′y_{0}^{\prime}. The final image can be recovered via wavelet inverse transformation x0′=IWT(y0′)x_{0}^{\prime}=\text{IWT}(y_{0}^{\prime}). We depict the sampling process in Algorithm 1.

2 Wavelet-embedded networks

Next, we further incorporate wavelet information into feature space through the generator to strengthen the awareness of high-frequency components. This is beneficial to the sharpness and quality of final images.

Fig. 3 illustrates the structure of our proposed wavelet-embedded generator. It follows the UNet structure of with MM down-sampling and MM up-sampling blocks plus skip connections between blocks of the same resolution, with MM predefined. However, instead of using the normal downsampling and upsampling operators, we replace them with frequency-aware blocks. At the lowest resolution, we employ frequency-bottleneck blocks for better attention on low and high-frequency components. Finally, to incorporate original signals YY to different feature pyramids of the encoder, we introduce frequency residual connections using wavelet downsample layers. Let denote YY be the input image and FiF_{i} is the ii-th intermediate feature map of YY. We will discuss below the newly introduced components:

Frequency-aware downsampling and upsampling blocks. Traditional approaches relied on a blurring kernel for the downsampling and upsampling process to mitigate the aliasing artifact. We instead utilize inherent properties of the wavelet transform for better upsampling and downsampling (depicted in Fig. 4). This, in fact, strengthens the awareness of high-frequency information on these operations. Particularly, the downsampling block receives a tuple of input features FiF_{i}, a latent zz, and time embedding tt, which are then processed through a sequence of layers to return downsampled features and high-frequency subbands. These returned subbands are served as an additional input to upsample features based on frequency cues in the upsampling block.

Frequency bottleneck block locates at the middle stage, which includes two frequency bottleneck blocks and one attention block in-between. Each frequency bottleneck block first divides feature map FiF_{i} into the low-frequency subband Fi,llF_{i,ll} and the concatenation of high-frequency subbands Fi,HF_{i,H}. Fi,llF_{i,ll} is then passed as input to resnet block(s) for deeper processing. The processed low-frequency feature map and the original high-frequency subbands Fi,HF_{i,H} are transformed back to the original space via IWT. With such a bottleneck, the model can focus on learning intermediate feature representations of low-frequency subbands while preserving the high-frequency details.

Frequency residual connection The original design of the network in incorporates original signals YY to different feature pyramids of the encoder via a strided-convolution downsampling layer. We instead use a wavelet downsample layer to map residual shortcuts of input YY to the corresponding feature dimensions, which are then added to each feature pyramid. Specifically, the residual shortcuts of YY are decomposed into four subbands which are then concatenated and fed to a convolution layer for feature projection. This shortcut aims to enrich the perception of the frequency source of feature embeddings.

Experiments

In this section, we first provide details about the experimental setup and then present the empirical experiments on different datasets, including CIFAR-10, STL-10, CelebA-HQ, and LSUN. Finally, we ablate the important components of our proposed framework.

Datasets We experiment on CIFAR-10 32×3232\times 32, STL-10 64×6464\times 64 and two other higher-resolution datasets, including CelebA-HQ 256×256256\times 256 and LSUN-Church 256×256256\times 256. We also examine our model training on high-resolution images with CelebA-HQ (512 & 1024).

Evaluation metrics We measure image fidelity by Frechet inception distance (FID) and measure sample diversity by Recall metric . Following , FID and Recall are computed over 50k generated samples. To further demonstrate our faster sampling, we measure the average inference time over 300 trials for a batch size of 100. Besides, the inference time of high-resolution images like CelebA-HQ 512×\times512 is computed from batches of 25 samples.

Implementation details Our implementation is mainly based on DDGAN . We adopt the same training configurations as DDGAN for all experiments. Training epochs are set as 500 for CelebA-HQ 256×\times256, 500 for LSUN 256×\times256, and 1800 for CIFAR-10 32×\times32. For CelebA-HQ (512 & 1024), we train our model and DDGAN for 400 epochs. We train our models on 1 - 8 NVIDIA A100 GPUs for each corresponding dataset. Notably, our models require fewer GPU resources and computations than DDGAN thanks to the efficiency of our wavelet diffusion framework as demonstrated in Tab. 1. At the same dataset, our model has a comparable number of parameters but requires less computing FLOPs and memory usage compared with DDGAN . Following DDGAN , we use 2-sampling steps for CelebA-HQ (256, 512 & 1024) and 4-sampling steps for CIFAR-10 (32), STL-10 (64), and LSUN-Church (256) in both training and testing.

Our speed gain does not only come from the proposed framework but also from corresponding network configurations. As the wavelet input of our model is 4×4\times smaller than the original dimensions, we need to provide suitable network configurations. On CIFAR-10, we use only 3 layers for both the generator and the discriminator instead of 4 as in DDGAN. On CelebA-HQ (256) and LSUN-Church, we use 5 instead of 6 layers. On other datasets, we use the same configurations as in DDGAN, with 6 layers for CelebA-HQ (512 & 1024) and 4 layers for STL-10.

2 Experimental results

CIFAR-10 As shown in Tab. 2, we have greatly improved inference time by requiring only 0.08(s)0.08(s) with our Wavelet based diffusion process, which is 2.5×2.5\times faster than DDGAN. This gives us a real-time performance along with StyleGAN2 methods while exceeding other diffusion models by a wide margin in terms of sampling speed.

STL-10 In Tab. 3, we not only achieve a better FID score at 12.9312.93 but also gain faster sampling time at 0.38(s)0.38(s). We further analyze the convergence speed of our approach and DDGAN in Fig. 8a. Our approach offers a faster convergence than the baseline, especially in the early epochs. Generated samples of DDGAN can not recover objects’ overall shape and structure during the first 400 epochs. Besides, we provide sample images generated by our model in Fig. 8b.

CelebA-HQ At resolution 256×256256\times 256, we outperform some notable diffusion and GAN baselines with FID at 5.945.94 and Recall at 0.370.37 while achieving more than 2×2\times faster than DDGAN, as reported in Tab. 4. On high-resolution CelebA-HQ (512) in Tab. 5, our model is remarkably better than the DDGAN counterpart for both image quality (6.406.40 vs. 8.438.43) and sampling time (0.590.59 vs. 1.491.49). We also obtain a higher recall at 0.350.35 on high-resolution CelebA-HQ (512). Our further benchmarking on CelebA-HQ 1024 yielded a competitive FID score of 5.985.98, comparable to many high-res GAN baselines (StyleGAN - 5.065.06, PGGAN - 7.37.3). Notably, its inference time (0.59s for 25 samples) remains comparable to CelebA-HQ 512, thanks to our unchanged network configurations with only two modifications: removal of attention layers and use of patch size 2. For qualitative results, we provide the generated samples for the CelebA-HQ (256) and CelebA-HQ (512) datasets in Fig. 6 and Fig. 7, respectively.

LSUN-Church In Tab. 6, we obtain a superior image quality at 5.06 in comparison with other diffusion models while achieving comparable results with GAN counterparts. Notably, our model offers 2×2\times faster inference time compared with DDGAN while exceeding StyleGAN2 in terms of sample diversity at 0.400.40. For qualitative results, randomly generated samples are shown at Fig. 9.

3 Ablation studies

Reconstruction term In this section, we verify the contribution of reconstruction loss Eq. 5 to the model performance on CelebA-HQ 256×256256\times 256. The FID score experiences a remarkable reduction of around 0.60.6 points to 5.945.94 when employing the term, confirming its usefulness in improving model’s quality.

Wavelet-embedded networks We validate the contribution of each individual component of our proposed wavelet-based generator on CelebA-HQ 256×256256\times 256 in Tab. 7, where the full model includes residual connections, upsampling and downsampling blocks, and bottleneck blocks. As can be seen, each component has a positive effect on the model’s performance. By applying all three proposed components, our method achieves the best performance at 5.94, especially the bottleneck block is the least important component in the design of generator. Still, the performance gain comes along with a small cost in running speed.

4 Running time when generating a single image

We further demonstrate the superior speed of our models on single images, as expected in real-life applications. In Tab. 8, we present their time and key parameters. Our wavelet diffusion models can produce images up to 1024×10241024\times 1024 in a mere 0.1s, which is the first time for a diffusion model to achieve such almost real-time performance.

5 Why ours converges faster and more stably?

We believe that our model benefits from the frequency decomposition of wavelet transformation. Instead of learning from an entanglement of coarse and detailed information, our method separates them for efficient training and at multi scales in the feature space. First, it studies the low-frequency subbands easier due to lower spatial dimensions. Second, it quickly learns the sparse and repetitive high-frequency components, focusing on distinctive details.

Conclusions

This paper introduced a novel wavelet-based diffusion scheme that demonstrates superior performance on both image fidelity and sampling speed. By incorporating wavelet transformations to both image and feature space, our method can achieve the state-of-the-art running speed for a diffusion model, closing the gap with StyleGAN models while obtaining a comparable image generation quality to StyleGAN2 and other diffusion models. Besides, our method offers a faster convergence than the baseline DDGAN , confirming the efficiency of our proposed framework. With these initial results, we hope our approach can facilitate future studies on real-time and high-fidelity diffusion models.

References

Appendix A Sensitivity analysis of training batch size

We recognize that training batch size is a critical aspect affecting the final performance. Large batch size often results in worse performance, with the compensation being training time. Here, we carefully analyze the effect of training batch size on the model performance measured by Frechet inception distance (FID) (depicted in Fig. 10).

As expected, the model trained with batch size 64 consistently performs better than the model trained with batch size 128, as illustrated via the training curves on LSUN-Church in Fig. 10(a). The gap is initially large during the first 100 epochs and then shrunk in the following epochs. However, the performance gap is still significant. The best FID of the model with batch size 64 is 5.06, which is 0.6 points lower than the best FID of the model trained on batch size 128. More importantly, our model outperforms DDGAN (5.06 vs. 5.25) when using the same batch size of 64.

We further verify the effect of batch size to our trained models on CelebA-HQ (256) with batch sizes 128 and 64. As shown in Fig. 10(b), the model trained with batch size 64 achieves a minimum FID of 5.93, which is 0.280.28 points lower than the FID 6.21 of the model trained with batch size 128. It again confirms that batch size is an important factor to be considered when evaluating and comparing model performance.

Note that we used a larger batch size than DDGAN (32 vs. 16) on CelebA-HQ 512×512512\times 512 with 8 GPUsNVIDIA A100-40GB GPUs are used. Our model can fit a training batch size of 4 instead of 2 as DDGAN per GPU.. Due to time and resource limits, we cannot retrain our model with batch size 16, but we expect our result will be further improved in this fair experiment configuration, further increasing the gap in performance between our algorithm and DDGAN.

Appendix B Experimental details

We utilize the implementation of wavelet transformations, including Discrete wavelet transform (DWT) and Discrete inverse wavelet transform (IWT), from . We perform these transformations on both input images and feature maps for further processing in the proposed Wavelet-based Diffusion framework.

B.2 Network configurations

Generator. Our generator has a UNet alike architecture which is mainly based on NCSN++ . As can be seen in Tab. 9, we show the detailed configurations of the generator for each corresponding dataset. We adjust the number of layers in the generator according to the input resolution of wavelet coefficients. The number of channels of time embedding is 4×4\times larger than the base channels.

Discriminator. The number of layers in the discriminator is the same as the one of the generator. For more details of the discriminator structure, please refer to .

B.3 Training hyper-parameters

For reproductivity, we further provide a full table of tuned hyper-parameters in Tab. 10. Basically, our hyper-parameters are the same as the baseline except for the number of epochs and the allocated GPUs on specific datasets. Meanwhile, there are two new datasets, including STL-10 64×6464\times 64 and CelebA-HQ 512×512512\times 512, that share similar configurations of CIFAR-10 and CelebA-HQ 256×256256\times 256, respectively. Besides, the setting of CelebA-HQ 1024×10241024\times 1024 is almost similar to CelebA-HQ 512×512512\times 512.

For training time, CIFAR10 and STL10 models require 1.6 and 3.6 days on a single GPU, respectively. On CelebA-HQ 256 and LSUN-Church, they take 1.1 and 6.8 days on 2 and 4 GPUs, respectively. On high-resolution CelebA-HQ 512, it takes 4.3 days on 8 GPUs. Besides, the training time is mainly influenced by the number of denoising steps, the size of network architectures, and image resolutions (presented in Tab. 9) apart from the number of training epochs. This is also the same for the inference time.

Appendix C More qualitative results

We further provide additional qualitative results on CIFAR-10 in Fig. 12, STL-10 in Fig. 13(a), CelebA-HQ 256 in Fig. 14, CelebA-HQ 512 in Fig. 15, LSUN-Church in Fig. 16, and CelebA-HQ 1024 in Fig. 17.

A comparison of qualitative samples between ours and DDGAN on STL-10 is also presented in Fig. 13. Our model clearly achieves better sample quality with a plausible appearance of generated objects, while the counterpart fails to represent object-specific shapes in output samples. We also add a qualitative comparison on the CelebA-HQ 512 dataset (Fig. 11), which further illustrates the advantages of our proposal in producing clearer details, such as eyebrows and wrinkles.

Appendix D More discussion

Increment novelty. While wavelet transformation has been used in many previous works on different tasks, ours is the first work employing it in diffusion models and in a comprehensive manner. Wavelet transformation is carefully incorporated on both the pixel and feature levels. Particularly, for the feature level, we proposed three wavelet-based network components, and each component is designed to utilize low and high-frequency subbands for improved output quality. Thanks to these proposals, we achieve state-of-the-art running speed with a near real-time performance for a diffusion model, allowing this advanced technique to be applicable to real-time applications. Our method also provides a faster and more stable model training, as discussed in Sec 5.5. of the main paper. Hence, we believe our paper is essential and not just an incremental work.

Reasons why the high-frequency subbands are unprocessed and directly transmitted to the decoder in the frequency bottleneck block. When designing this block, we aimed to strengthen our model’s focus on learning low-level features while preserving the details; hence we directly transmitted the high-frequency components to the IWT module. We have tested processing both low and high subbands on STL-10 and got almost the same FID (12.9612.96 vs. 12.9312.93), suggesting that this design is not critical.

Progressive upsampling. We actually tried this direction first, but it produced quite poor results. On CelebA-HQ 256, the FID score of the 2-level upsampling network is 13.1113.11, much higher than ours (5.945.94). We suspect a discrepancy in conditional distribution between low and high-frequency p(xhi∣xlo)p(x_{hi}|x_{lo}), causing mismatching between generated subbands, and deteriorating the output quality.