Frequency Domain Image Translation: More Photo-realistic, Better Identity-preserving
Mu Cai, Hong Zhang, Huijuan Huang, Qichuan Geng, Yixuan Li, Gao Huang
Introduction
Image-to-image translation has attracted great research attention in computer vision, which is tasked to synthesize new images based on the source and reference images (see Figure 1). This task has been revolutionized since the introduction of GAN-based methods . In particular, a plethora of literature attempts to decompose the image representation into a content space and a style space . To translate a source image, its content representation is combined with a different style representation from the reference domain.
Despite exciting progress, existing solutions suffer from two notable challenges. First, there is no explicit mechanism that allows preserving the identity, and as a result, the synthesized image can over-adapt to the reference domain and lose the original identity characteristics. This can be observed in Figure 1, where Swapping Autoencoder generates images with identity and structure closer to the reference rather than the source image. For example, in the second row, the tree is absent from the source image yet occurs in the translation result. Second, the generation process may lose important fine-grained details, leading to suboptimal visual quality. This can be prohibitive for generating photo-realistic high-resolution images. The challenges above raise the following important question: how can we enable photo-realistic image translation while better preserving the identity?
Motivated by this, we propose a novel framework–Frequency Domain Image Translation (FDIT)–exploiting frequency information for enhancing the image generation process. Our key idea is to decompose the image into low- and high-frequency components, and regulate the frequency consistency during image translation. Our framework is inspired by and grounded in signal processing . Intuitively, the low-frequency component captures information such as color and illumination; whereas the high-frequency component corresponds to sharp edges and important details of objects. For example, Figure 2 shows the resulting images via adopting the Gaussian blur to decompose the original image into low- vs. high-frequency counterparts (top vs. bottom). The building identity is distinguishable based on the high-frequency components.
Formally, FDIT introduces novel frequency-based training objectives, which facilitates the preservation of frequency information during training. The frequency information can be reflected in the visual space as identity characteristics and important fine details. Formally, we impose restrictions in both pixel space as well as the Fourier spectral space. In the pixel space, we transform each image into its high-frequency and low-frequency components by applying the Gaussian kernel (i.e., low-frequency filter). A loss term regulates the high-frequency components to be similar between the source image and the generated image. Furthermore, FDIT directly regulates the consistency in the frequency domain by applying Fast Fourier Transformation (FFT) to each image. This additionally ensures that the original and translated images share a similar high-frequency spectrum.
Extensive experiments demonstrate that FDIT is highly effective, establishing state-of-the-art performance on image translation tasks. Below we summarize our key results and contributions:
We propose a novel frequency-based image translation framework, FDIT, which substantially improves the identity-preserving generation, while enhancing the image hybrids realism. FDIT outperforms competitive baselines by a large margin, across all datasets considered. Compared to the vanilla Swapping Autoencoder (SwapAE) , FDIT decreases the FID score by 5.6%.
We conduct extensive ablations and user study to evaluate the (1) identity-preserving capability and (2) image quality, where FDIT constantly surpasses previous methods. For example, user study shows an average preference of 75.40% and 64.39% for FDIT over Swap AE in the above two aspects. We also conduct the ablation study to understand the efficacy of different loss terms and frequency supervision modules.
We broadly evaluate our approach across five large-scale datasets (including two newly collected ones). Quantitative and qualitative evaluations on image translation and GAN-inversion tasks demonstrate the superiority of our method Code and dataset are available at: https://github.com/mu-cai/frequency-domain-image-translation.
Background: Image-to-image Translation
Given an image , the encoder maps it to a latent representation . Previous approaches rely on the assumption that the latent code can be composed into two components , where and correspond to the content and style information respectively. A reconstruction loss minimizes the norm between the original input and .
To perform image translation, the generator takes the content code from the source image, together with the style code from the reference image. The translated image is given by . However, existing methods can be limited by its feature disentanglement ability, where may not capture the identity of source image. As a result, such identity-related characteristics can be undesirably lost in translation (see Figure 5), which motivates our work.
Frequency Domain Image Translation
Our novel frequency-based image translation framework is illustrated in Figure 3. In what follows, we first provide an overview and then describe the training objective. Our training objective facilitates the preservation of frequency information during the image translation process. Specifically, we impose restrictions in both pixel space (Section 3.1) as well as the Fourier spectral space (Section 3.2).
We transform each input into two images and , which correspond to the low-frequency and high-frequency images respectively. Note that both and are in the same spatial dimension as . Specifically, we employ the Gaussian kernel, which filters the high frequency feature and keeps the low frequency information:
where denotes the spatial location within the image, and denotes the variance of the Gaussian function. Following , the variance is increased proportionally with the Gaussian kernel size . Using convolution of the Gaussian kernel on input , we obtain the low frequency (blurred) image :
where denotes the index of an 2D Gaussian kernel, i.e., .
To obtain , we first convert color images into grayscale, and then subtract the low frequency information:
where the rgb2gray function converts the color image to the grayscale. This removes the color and illumination information that is unrelated to the identity and structure. The resulting high frequency image contains the sharp edges, i.e. sketch of the original image.
We now employ the following reconstruction loss term, which enforces the similarity between the input and generator’s output, for both low-frequency and high-frequency components:
In addition to reconstruction loss, we also employ the translation matching loss:
where and are the content code of the source image and the style code of the reference image, respectively. Intuitively, the translated images should adhere to the identity of the original image. We achieve this by regulating the high frequency components, and enforce the generated image to have the same high frequency images as the original source image.
2 Fourier Frequency Space Loss
For the ease of post processing, we then transform from the complex number domain to the real number domain. Additionally, we take the logarithm to stabilize the training:
We then regulate the reconstruction loss in the frequency spectrum:
In a similar spirit as Equation 5, we devise a translation matching loss in the Fourier frequency domain:
where . is the frequency mask, for which we provided detailed explanation below. The loss constrains the high frequency components of the generated images for better identity preserving.
As illustrated in Figure 3, the low-frequency mask is a circle with radius , whereas the high-frequency mask is the complement region. The frequency masks and can be estimated empirically from the distribution of on the entire training dataset. We choose the radius to be 21 for images with resolution 256256. The energy within the low-frequency mask accounts for 97.8% of the total energy in the spectrum.
3 Overall Loss
Considering all the aforementioned losses, the overall loss is formalized as:
Gaussian kernel and FFT are complementary for preserving the frequency information. On one hand, the Gaussian kernel extracts the frequency information via the convolution, therefore representing the frequency features in a local manner. On the other hand, Fast Fourier Transformation utilizes the information from all pixels to obtain the FFT value for each spatial frequency, characterizing the frequency distribution globally. Gaussian kernel and FFT are therefore complementary in preserving the frequency information. We show ablation study on this in Section 4.2, where both are effective in enhancing the identity-preserving capability for image translation tasks.
When transforming the images in Figure 2 into the spectrum space, the effects of the Gaussian kernel size could be clearly reflected in Figure 4. To be specific, a large kernel would cause severe distortion on the low-frequency band while a small kernel would not preserve much of the high-frequency information. In this work, we choose the kernel size for images with resolution 256256, which could appropriately separate the high/low-frequency information, demonstrated in both image space and spectral space distribution. Our experiments also show that FDIT is not sensitive to the selection of as long as it falls into a mild range.
Experiments
In this section, we evaluate our proposed method on two state-of-the-art image translation architectures, i.e., Swapping Autoencoder , StarGAN v2 , and one GAN inversion model, i.e., Image2StyleGAN . Extensive experimental results show that FDIT not only better preserves the identity, but also enhances image quality.
We evaluate FDIT on the following five datasets: (1) LSUN Church , (2) CelebA-HQ , (3) LSUN Bedroom , (4) Flickr Mountains (100k self-collected images), (5) Flickr Waterfalls (100k self-collected images). (6) Flickr Faces HQ (FFHQ) dataset . All the images are trained and tested at resolution except FFHQ, which is trained at , and finetuned at resolution. For evaluation, we use a validation set that is separate from the training data.
1 Autoencoder
Autoencoder is widely used as the backbone of the deep image translation task . We use state-of-the-art Swapping Autoencoder (SwapAE) , which is built on the backbone of StyleGAN2 . Swap AE also uses the technique in PatchGAN to further improve the texture transferring performance. We incorporate our proposed FDIT training objectives into the vanilla SwapAE.
We contrast the image translation performance using FDIT vs. vanilla SwapAE in Figure 1 and Figure 5. The vanilla SwapAE is unable to preserve the important identity of the source images, and over-adapts to the reference image. For example, the face identity is completely switched after translation, as seen in rows 4 of Figure 5. SwapAE also fails to preserve the outline and the local sharp edges in the source image. As shown in Figure 1, the outlines of the mountains are severely distorted. Besides, the overall image composition has a large shift from the original source image. In contrast, using our method FDIT, the identity and structure of the swapped hybrid images are highly preserved. As shown in Figure 1 and Figure 5, the overall sketches and local fine details are well preserved while the coloring, illumination, and even the weather are well transferred from the reference image (top rows of Figure 1).
Lastly, we compare FDIT with the state-of-the-art image stylization method STROTSS and WCT2 . Image stylization is a strong baseline as it emphasizes on the strict adherence to the source image. However, as shown in Figure 5, WCT2 leads to poor transferability in image generation tasks. Despite strong identity-preservation, STROTSS and WCT2 are less flexible, and generate images that highly resemble the source image. In contrast, FDIT can both preserve the identity of the source image as well as maintain a high transfer capability. This further demonstrates the superiority of FDIT in image translation.
We show in Table 1 that FDIT can substantially improve the image quality while preserving the image content. We adopt the Fréchet Inception Distance (FID) as the measure of image quality. Small values indicate better image quality. Details about Im2StyleGAN and StyleGAN2 are shown in the supplementary material. FDIT achieves the lowest FID across all datasets. On average, FDIT could reduce the FID score by 5.6 compared to the current state-of-the-art method.
1.2 Image Attributes Editing
We show that FDIT enables image attribute editing task, which creates a series of smoothly changing images between two sets of distinct images . Vector arithmetic is one commonly used way to achieve this . For example, we can sample images from each of the two target domains, and then compute the average difference of the vectors between these two sets of images:
where denote the latent code from two domains.
We perform interpolation on the style code while keeping the content code unchanged. The generated images can be formalized as , where is the interpolation parameter. We show results on CelebA-HQ dataset in Supplementary material. FDIT performs image editing towards the target domain while strictly adhering to the content of the source image. Compared to the vanilla Swapping Autoencoder and StarGAN v2, our results demonstrate the better disentanglement ability of unique image attributes and identity characteristics. We also verify the disentangled semantic latent vectors using Principal Component Analysis (PCA). The implementation details and the identity-preserving results are shown in the supplementary materials.
2 Ablation Study
Pixel and Fourier space losses are complementary. To better understand our method, we isolate the effect of pixel space loss and Fourier spectral space loss. The results on the LSUN Church dataset are summarized in Table 2. The vanilla SwapAE is equivalent to having neither loss terms, which yields the FID score of 52.34. Using pixel space frequency loss reduces the FID score to 49.47. Our method is most effective when combining both pixel-space and Fourier-space loss terms, achieving the FID score of 48.21. Our ablation signifies the importance of using frequency-based training objectives.
3 GAN Inversion
FDIT improves reconstruction quality in GAN inversion. We evaluate the efficacy of FDIT on the GAN inversion task, which maps the real images into the noise latent vectors. In particular, Image2StyleGAN serves as a strong baseline, which performs reconstruction between the real image and the generated images via iterative optimization over the latent vector.
We adopt the same architecture, however impose our frequency-based reconstruction loss. The inversion results are shown in Figure 6. On high-resolution () images, the quality of the inverted images is improved across all scenes. FDIT better preserves the overall structure, fine details, and color distribution. We further measure the performance quantitatively, summarizing the results in Table 3. Under different metrics (MSE, MAE, PSNR, SSIM), our method FDIT outperforms Image2StyleGAN.
4 StarGAN v2
StarGAN v2 is another state-of-the-art image translation model which can generate image hybrids guided by either reference images or latent noises. Similar to the autoencoder-based network, we can optimize the StarGAN v2 framework with our frequency-based losses. In order to validate FDIT in a stricter condition, we construct a CelebA-HQ-Smile dataset based on the smiling attribute from CelebA-HQ dataset. The style refers to whether that person smiles, and the content refers to the identity.
Several salient observations can be drawn from Figure 7. First, FDIT can highly preserve the gender identity; whereas the vanilla StarGAN v2 model would change the resulting gender according to the reference image (e.g. first and second row). Secondly, the image quality of FDIT is better, where FID is improved from 17.32 to 16.86. Thirdly, our model can change the smiling attribute while maintaining other facial features strictly. For example, as shown in the third row, StarGAN v2 undesirably changes the hairstyle from straight (source) to curly (reference), whereas FDIT maintains the same hairstyle.
5 User Study
We conduct a user study to qualitatively measure the generated images. Specifically, we employ the two-alternative forced-choice setting, which was commonly used to train Learned Perceptual Image Patch Similarity (LPIPS) and to evaluate style transfer methods. We provide users with the source image, reference image, images generated by FDIT, and the baseline SOTA models. Each user is forced to choose which of the two image hybrids 1) better preserves the identity characteristics, and 2) has better image quality. We collected a total of 2,058 user preferences across 5 diverse datasets. Results are summarized in Table 4. On average, 75.40% of preferences are given to FDIT for identity preserving; and 64.39% of answers indicate FDIT produces more photo-realistic images.
Furthermore, comparing to StarGAN v2, 57.14% user preferences are given to FDIT for better content preservation; 53.34% user preferences indicate that FDIT produces better image quality compared to Image2StyleGAN. Therefore, the user study also verifies that FDIT produces better identity-preserving and photo-realistic images.
Related work
GAN has revolutionized revolutionized many computer vision tasks, such as super resolution , colorization , and image synthesis . Early work directly used the Gaussian noises as inputs to the generator. However, such an approach has unsatisfactory performance in generating photo-realistic images. Recent works significantly improved the image reality by injecting the noises hierarchically in the generator. These works adopt the adaptive instance normalization (AdaIN) module for image stylization.
Image-to-image translation synthesizes images by following the style of a reference image while keeping the content of the source image. One way is to use the GAN inversion, which maps the input from the pixel space into the latent noises space via the optimization method . However, these methods are known to be computationally slow due to their iterative optimization process, which makes deployment in mobile devices difficult . Furthermore, the quality of the reconstructed images can be suboptimal. Another approach is to utilize the conditional GAN (or autoencoder) to convert the input images into latent vectors , making the image translation process much faster than GAN inversion. However, exiting state-of-the-art image translation models such as StarGAN v2 and Swapping Autoencoder can lose important structural characteristics of the source image. In this paper, we show that frequency-based information can effectively preserve the identity of the source image and enhance photo-realism.
Frequency domain analysis is widely used in traditional image processing . The key idea of frequency analysis is to map the pixels from the Euclidean space to a frequency space, based on the changing speed in the spatial domain. Several works tried to bridge the connection between deep learning and frequency analysis . Chen et al. and Xu et al. showed that by incorporating frequency transformation, the neural network could be more efficient and effective. Wang et al. found that the high-frequency components are useful in explaining the generalization of neural networks. Recently, Durall et al. observed that the images generated by GANs are heavily distorted in high-frequency parts, and they introduced a spectral regularization term to the loss function to alleviate this problem. Czolbe et al. proposed a frequency-based reconstruction loss for VAE using discrete Fourier Transformation (DFT). However, this approach does not incorporate pixel space frequency information, and relies on a separate dataset to get its free parameters. In fact, no prior work has explored using frequency-domain analysis for the image-to-image translation task. In this work, we explicitly devise a novel frequency domain image translation framework and demonstrate its superiority in performance.
Neural style transfer aims at transferring the low-level styles while strictly maintaining the content in the source image . Typically, the texture is represented by the global image statistics while the content is controlled by the perception metric . However, existing methods could only handle the local color transformation, making it hard to transform the overall style and semantics. More specifically, they struggle in the cross-domain image translations, for example, gender transformation . In other words, despite strong identity-preservation ability, such methods are less flexible for the cross-domain translation and can generate images that highly resemble the source domain. In contrast, FDIT can both preserve the identity of the source images while maintaining a high domain transfer capability.
Conclusion
In this paper, we propose Frequency Domain Image Translation (FDIT), a novel image translation framework that preserves the frequency information in both pixel space and Fourier spectral space. Unlike the existing image translation models, FDIT directly uses high-frequency components to capture object structure akin to the identity. Experimental results on five large-scale datasets and multiple tasks show that FDIT effectively preserves the identity of the source image while producing photo-realistic image hybrids. Extensive user study and ablations further validate the effectiveness of our approach both qualitatively and quantitatively. We hope future research will increase the attention towards frequency-based approaches for image translation tasks.
Acknowledgment
Mu Cai and Yixuan Li are supported by funding from the Wisconsin Alumni Research Foundation (WARF). Gao Huang is supported in part by the National Key R&D Program of China under Grant 2020AAA0105200, the National Natural Science Foundation of China under Grants 62022048 and 61906106, the Institute for Guo Qiang of Tsinghua University and Beijing Academy of Artificial Intelligence.
References
Appendix A Image Attributes Editing Results
We demonstrate the identity preserving capability and photo realism of FDIT under the image attribute editing task via continuous interpolation and unsupervised semantic vector discovery.
We show that FDIT can generate a series of smoothly changing images between two sets of distinct images. We perform interpolation on the style code while keeping the content code unchanged. Figure 8 shows season transformation results using the Flicker Mountains dataset. Our identity-preserving image hybrids demonstrate that FDIT could achieve high-quality image editing performance towards the target domain while strictly adhering to the identity of the source image.
A.2 Unsupervised Semantic Vector Discovery for Image Editing
Another way to conduct image editing is to discover the underlying semantics via an unsupervised way. Here we adopt the Principal Component Analysis (PCA) to achieve this goal, which could find the orthonormal components in the latent space. Similar to the continuous interpolation approach in our paper, when manipulating the style code using PCA, a good image translation model would keep the content of the images as untouched as possible.
As shown in Fig. 9, FDIT is once again demonstrated to be an identity-preserving model. Specifically, the identities are well maintained, while the only facial attributes such as illumination and hair color are changed.
We additionally show results of image editing in the full latent space in Figure 10, which displays more variation.
Appendix B Frequency Domain Image Translation Results
We show the image generation results of the autoencoder based FDIT framework on LSUN Church , CelebA-HQ , Flickr Waterfalls, and LSUN Bedroom in Figure 11. FDIT framework achieves better performance in preserving the shape, which can be observed in the outline of the churches, the layout of the bedrooms, and the scene of the waterfalls.
Appendix C Constructing the Flicker Dataset
We collect the large-scale Flicker Mountains dataset and Flicker Waterfalls dataset from flickr.com. Each dataset contains 100,000 training images.
Appendix D Training Details
Our Frequency Domain Image Translation (FDIT) framework is composed of the pixel space and Fourier frequency space losses, which can be conveniently implemented for existing image translation models. For fair comparison, we keep all training and evaluation settings the same as the baselines (Swapping Autoencoder https://github.com/rosinality/swapping-autoencoder-pytorch , StarGAN v2 https://github.com/clovaai/stargan-v2 , and Image2StyleGAN https://github.com/pacifinapacific/StyleGAN_LatentEditor ). All experiments are conducted on the Tesla V100 GPU.
The encoder-decoder backbone is built on StyleGAN2 . We train the model on the 32GB Tesla V100 GPU, where the batch size is 16 for images of 256256 resolution, and 4 for images of resolution. During training, a batch of images are fed into the model, where reconstructed images and image hybrids would be produced. We adopt Adam optimizer where . The learning rate is set to be 0.002. The reconstructed quality is supervised by loss. The discriminator is optimized using the adversarial loss . A patch discriminator is utilized to enhance the texture transferring ability w.r.t. reference images.
We use the official implementation in StarGAN v2, where the backbone is built with ResBlocks . The batch size is set to be 8. Adam optimizer is adopted where . The learning rate for the encoder, generator, and discriminator is set to be . In the evaluation stage, we utilize the exponential moving averages over encoder and generator.
We adopt the Adam optimizer with the learning rate of 0.01, , and in the experiments. We use 5000 gradient descent steps to obtain the GAN-inversion images.
Appendix E Details of Image2StyleGAN and StyleGAN2 results in Table 1.
Both Im2StyleGAN and StyleGAN2 invert the image from the training domain, then use the mixed latent representations to create image hybrids. Image2StyleGAN adopts the iterative optimization on the ’-space’ to project images using the StyleGAN-v1 backbone; while StyleGAN2 utilizes an LPIPS-based projector under the StyleGAN-v2 backbone.
Appendix F The qualitative results for Section 4.2
The qualitative results are shown in Figure 12, where FDIT shows better identity preservation than using only pixel or Fourier loss. For example, using only Fourier loss preserves the identity but loses some style consistency in the pixel space.