GLEAN: Generative Latent Bank for Large-Factor Image Super-Resolution

Kelvin C. K. Chan, Xintao Wang, Xiangyu Xu, Jinwei Gu, Chen Change Loy

Introduction

In this study, we explore a new way to employ GAN for image super-resolution. We are interested in the regime of high magnification factors (8×8{\times} to 64×{64\times}), which typical SR methods fail to handle since most details and textures are lost during downsampling. Since the problem is severely underspecified, informative priors become inevitable in this setting, especially in restoring the textural details. Studying large-factor image SR is meaningful as it can potentially improve the state of the arts in SR, and more generally conditional generative models for images.

The notion of GAN has been extensively used in SR with the aim to enrich texture details in an upscaled image. There are two popular approaches to deploy GANs for this task. The more common paradigm trains a generator to handle the upscaling task, where adversarial training is performed by using a discriminator to differentiate real images from the upscaled images produced by the generator. Another possible way to exploit GAN for the task is by GAN inversion . In this setting, one will need to ‘invert’ the generation process of a pre-trained GAN by mapping a corrupted image back to the latent space. A restored image can then be reconstructed from the optimal vector in the latent space.

The second paradigm resolves the aforementioned problem by making better use of the latent space of GAN through optimization. However, as the low-dimensional latent codes and the constraints in the image space are insufficient to guide the restoration process, these methods often generate images with low fidelity. As shown in Fig. 1, despite being realistic, the output of a representative method, PULSE , fails to recover the structures of the ground-truth faithfully. In addition, as the optimization is usually conducted in an iterative manner for each image at runtime, these approaches are often time-consuming.

In our approach, we leverage pre-trained GANs such as StyleGAN to provide rich and diverse priors for the task. Unlike most GAN inversion methods, which also use pre-trained GANs, our method does not involve image-specific optimization at runtime. Once trained, the model only needs a single forward pass to upscale an image, which is more practical for applications that demand fast response. The idea is partially inspired by the classic notion of dictionary . But unlike conventional approaches that construct a finite and imagery-derived dictionary, we exploit GAN as a more effective way for storing priors.

Conditioning and retrieving from a GAN-based dictionary is a new and non-trivial question we need to address in this work. We show that pre-trained GANs can be employed as a latent bank in a succinct encoder-bank-decoder architecture. This novel architecture allows us to lift the burden of learning both fidelity and texture generation simultaneously in a typical encoder-decoder network since the latent bank already captures rich texture priors. In addition, we show that it is pivotal to condition the bank by passing both the latent vectors and multi-resolution convolutional features from the encoder to achieve high-fidelity results. Symmetrically, multi-resolution cues need to be passed from the bank to the decoder. We show the effectiveness of the proposed method in handling images with challenging poses and structures apart from the large magnification factor. We also demonstrate how the method can be generalized to different categories, \eg, human faces, cats, buildings, by switching different pre-trained GAN latent banks.

Related Work

Recent interests have shifted to large-factor SR beyond the typical upscaling factors (2×2{\times} or 4×4{\times}) . Dahl et al. propose a fully probabilistic pixel recursive network for upsampling extremely coarse images with resolution 8×88{\times}8. RFB-ESRGAN builds upon ESRGAN and adopts multi-scale receptive fields blocks for 16×16{\times} SR. VarSR achieves 8×8{\times} SR by matching the latent distributions of LR and HR images to recover the missing details. Zhang et al. perform 16×16{\times} reference-based SR on paintings with a non-local matching module and a wavelet texture loss. To handle even larger magnification factors, one would need to rely on stronger priors. SR methods specialized on large magnification factors are typically dedicated to the human face category as one could exploit the strong structural prior of faces. Facial priors including facial attributes , facial landmarks , and identity have been studied. Our work goes beyond previous works and pushes the limit to 64×64{\times} and generalizes to more categories. Such a large magnification factor is challenging due to its highly ill-posed nature.

GAN Inversion. Given a degraded image xx, GAN inversion-based methods in general produce a natural image best approximating xx by optimizing z∗=argmin⁡z∈ZL(G(z),x)z^{*}=\operatorname*{argmin}\nolimits_{z\in\mathcal{Z}}\mathcal{L}\left(G(z),x\right), where Z\mathcal{Z} is the latent space and L(⋅,⋅)\mathcal{L}(\cdot,\cdot) denotes the task-specific objective function. For instance, PULSE iteratively optimizes the latent code of StyleGAN with a pixel-wise constraint between the input and output. mGANprior optimizes multiple latent codes to increase the expressiveness of the model. DGP further finetunes the generator together with the latent code to reduce the gap between the distributions of the training and testing images. A common issue with GAN inversion is that important spatial information may not be faithfully kept due the low-dimensionality of the latent code. Thus, these methods often generate undesirable results that do not resemble the ground-truth. Different from GAN inversion, GLEAN conditions the pre-trained generator with both the latent codes and multi-resolution convolutional features, providing additional spatial guidance for restoration. In addition, GLEAN does not require iterative optimization during inference.

Methodology

A GAN model that is trained on large-scale natural images captures rich texture and shape priors. Previous studies have shown that such priors can be harvested through GAN inversion to benefit various image restoration tasks. Nonetheless, it remains underexplored how to exploit the priors without the expensive optimization during inversion.

In this study, we devise GLEAN within a novel encoder-bank-decoder architecture, which allows one to exploit the generative priors by needing just a single forward pass. An overview of the architecture is depicted in Fig. 2. Given a severely downsampled LR image, GLEAN applies an encoder to extract latent vectors and multi-resolution convolutional features, which capture important high-level cues as well as spatial structure of the LR image. Such cues are used to condition the latent bank, which further produces another set of multi-resolution features for the decoder. Finally, the decoder generates the final output by integrating the features from both the encoder and the latent bank. In this work, we adopt StyleGAN as the generative latent bank due to its exceptional performance. The idea of latent bank can be extended to other generators such as BigGAN .

To generate the latent vectors, we first use an RRDBNet (denoted as E0E_{0}) to extract features f0f_{0} from the input LR image. Then, we gradually reduce the resolution of the features by:

where EiE_{i}, i∈{1,⋯ ,N}{i\in\{1,\cdots,N\}}, denotes a stack of a stride-2 convolution and a stride-1 convolution. Finally, a convolution and a fully-connected layer are used to generate the latent vectors:

where CC is a matrix whose columns represent the latent vectors for the StyleGAN.

The latent vectors in CC capture a compressed representation of the images, providing the generative latent bank with high-level information. To further capture the local structures of the LR image and to provide additional guidance for structure restoration, we also feed multi-resolution convolutional features {fi}\{f_{i}\} into the latent bank.

2 Generative Latent Bank

Given the convolutional features {fi}\{f_{i}\} and the latent vectors CC, we leverage a pre-trained generator as a latent bank to provide priors for texture and detail generation. As StyleGAN is originally designed for image generation tasks, it cannot be directly integrated into the proposed encoder-bank-decoder framework. In this work, we adapt StyleGAN to our SR network by making three modifications:

Instead of taking one single latent vector as the input, each block of the generator takes a different latent vector to improve expressiveness. More specifically, we have C=(c0,⋯ ,ck−1)C{=}(\mathbf{c}_{0},\cdots,\mathbf{c}_{k-1}) for kk blocks, where each ci\mathbf{c}_{i} corresponds to one latent vector. We find that this modification leads to outputs with fewer artifacts. This modification is also seen in previous works .

To allow conditioning on the additional features from the encoder, we use an additional convolution in each style block for feature fusion:

where SiS_{i} denotes the augmented style block with an additional convolution, and gig_{i} corresponds to the output feature of the ii-th augmented style block.

Instead of directly generating outputs from the generator, we output the features {gi}\{g_{i}\} and pass them to the decoder to better fuse the features from the latent bank and encoder.

Advantages. The use of generative latent bank is reminiscent of the task of reference-based SR , where external HR reference image(s) are employed as an explicit imagery dictionary. While the external HR information leads to marked improvements, the performance is sensitive to the similarity between the inputs and references. This sensitivity may eventually lead to degraded results when the reference images/components are not well selected. Moreover, the size and diversity of those imagery dictionaries are limited by the selected components, impeding the generalization to diverse scenes in practice. In addition, computationally-intensive global matching or component detection/selection is often required to aggregate appropriate information from the references, hindering the applications to scenarios with tight computational constraints. Instead of constructing an imagery dictionary, GLEAN adopts a GAN-based dictionary conditioned on a pre-trained GAN. Our dictionary does not depend on any specific components or images. Instead, it captures the distribution of the images and has potentially unlimited size and diversity. Furthermore, GLEAN is computationally efficient without requiring global matching and reference images/components selection.

3 Decoder

GLEAN uses an additional decoder with progressive fusion to integrate the features from the encoder and latent bank to generate the output image. It takes the RRDBNet features as inputs and progressively fuse the features with the multi-resolution features from the latent bank:

where DiD_{i} and did_{i} denote a 3×33{\times}3 convolution and its output, respectively. Each convolution is followed by a pixel-shuffle layer except the final output layer. With the skip-connection between the encoder and decoder, the information captured by the encoder can be reinforced and hence the latent bank could focus more on the texture and detail generation.

4 Training

Similar to existing works , we adopt the standard l2l_{2} loss, perceptual loss , and adversarial loss for training. More details on the loss function can be found in the appendix. To exploit the generative prior, we keep the weights of the latent bank fixed throughout training. In our preliminary experiments, finetuning the latent bank with the encoder and decoder demonstrates no noticeable improvements. Moreover, it potentially harms the generalizability of the model as the latent bank may eventually bias to the training distribution. It is worth emphasizing that despite GLEAN is trained with similar objectives as in existing works (\eg ESRGAN), the main difference to these methods is that GLEAN leverages a pre-trained generator to directly incorporate the priors into the network, further improving the output quality. We show that the improvement is not due to additional parameters in the generator by comparing GLEAN with ESRGAN+, a larger ESRGAN that has similar FLOPs to GLEAN.

Experiments

We adopt pre-trained StyleGANGenForce: https://github.com/genforce/genforce or StyleGAN2BasicSR: https://github.com/xinntao/BasicSR (depending on the availability of pre-trained models) as our latent bank, and use the publicly available codes of existing methods for the comparison in this section. To maintain fairness, we train our model and baselines on the same datasets, including FFHQ and LSUN , so that the difference in restoration quality is mainly caused by the algorithms instead of the training distribution. Test set is strictly exclusive from the training. Detailed experimental settings are provided in the appendix.

Qualitative comparison. The qualitative comparison on 16×16{\times} SR is shown in Fig. 3. Guided by low-dimensional vectors and constraints in LR space, the outputs of GAN inversion methods are unable to maintain a good fidelity. In particular, PULSE and mGANprior fail to restore a face image with the same identity. In addition, artifacts are observed in their outputs. Through finetuning the generator during optimization, the result of DGP demonstrates significant improvements in both quality and fidelity. However, a slight difference between the identities of the output and ground-truth is still observed. For example, the eyes and lips show noticeable differences.

Methods trained with adversarial loss (SinGAN , ESRGAN+A larger version of ESRGAN with similar FLOPs to GLEAN. ) can preserve the local structures, but fail in synthesizing convincing textures and details. Specifically, SinGAN fails to capture the natural image style, producing a painting-like image. Although ESRGAN+ is capable of generating a realistic image, it struggles to synthesize fine details and introduces unnatural artifacts in detailed regions. It is worth emphasizing that although ESRGAN+ achieves competitive results on human faces, its performances on other categories such as cats and cars are less promising (see Fig. 1 and Fig. 4). With the latent bank providing natural image priors, GLEAN succeeds in both fidelity and naturalness. For example, when compared to ESRGAN+, GLEAN reconstructs eyes with better shape and details. We further extend our method to larger scale factors in Fig. 5. GLEAN successfully generates perceptually convincing images resembling the ground-truth for up to 64×64{\times} upscaling.

Robustness to poses and contents. Another appealing property of GLEAN is its robustness to the changes in poses and contents. As shown in Fig. 6, guided by the convolutional features, GLEAN is still able to construct realistic images when the images are non-aligned and contain non-human faces despite it is trained on aligned human faces. In contrast, the outputs of PULSE are biased to aligned human faces. Its outputs can only approximate the ground-truths in low resolution. Such robustness enables GLEAN to be applied to diverse categories and scenes such as cats, cars, bedrooms, and towers. Examples are shown in Fig. 4 and more results are provided in the appendix.

Quantitative comparison. To demonstrate the ability of GLEAN in producing outputs with high fidelity, we extract 100 images from CelebA-HQ and compute the cosine similarity to the ground-truth on the ArcFace embedding space . As shown in Table 1, GLEAN achieves higher similarity than the baseline methods, validating the superiority of GLEAN.

We additionally provide the quantitative comparison on different categories in Table 2. For each category, we select 100 images and compute their average PSNR and LPIPS . It is observed that mGANprior and PULSE perform significantly worse as they fail to restore the original objects. GLEAN outperforms these methods in most categories, suggesting its effectiveness in generating images with high quality and fidelity.

Ablation Studies

Importance of multi-resolution encoder features. We demonstrate how the convolutional features generated from the encoder assist in the restoration of fine details and local structures. We start with only the latent vectors and observe the transition when features are gradually introduced to the latent bank as conditions. To discard the effects brought by the decoder, we test with a variant of GLEAN where the generator directly produces the output images. The comparison is depicted in Fig. 7.

When all convolutional features are discarded, GLEAN resembles the typical GAN inversion methods that learn only the latent vectors. Similar to those methods, the network is able to synthesize realistic images given the latent vectors. However, guided only by low-dimensional vectors, in which spatial information is not well-preserved, the network restores only the global attributes such as hair color and poses, but fails to preserve finer details. When providing coarse (from 4×44{\times}4 to 16×1616{\times}16) convolutional features to the latent bank, more details are recovered and the outputs are better approximating the ground-truths. Further improvements in both quality and fidelity are observed when finer features are passed to the latent bank. The above observations corroborate our hypothesis that the convolutional features are pivotal in guiding the restoration of fine details and local structures, which cannot be reconstructed with only the latent vectors.

Effects of latent bank features. To understand the contributions of the latent bank, we investigate the effects brought by the latent bank features. We start by discarding all the latent bank features, and progressively pass the features to the decoder. The comparison is shown in Fig. 8. Lacking appropriate prior information, the network is responsible for both generating realistic details and maintaining fidelity to the ground-truths. Such a demanding objective eventually leads to outputs that contain flaws in both structure restoration and texture generation. With the latent bank, the burden of texture and details generation is reduced as the generator already captures rich image priors. Therefore, improvements in both structures and textures are observed when passing finer features to the decoder.

Importance of decoder. As shown in Fig. 9, without the decoder, despite being perceptually convincing overall, the output image contains unpleasant artifacts when zoomed in. The decoder allows the network to aggregate the information in a coarse-to-fine manner, leading to more natural details. In addition, the multi-scale skip-connections between the encoder and decoder reinforce the spatial information captured in the encoder features so that the latent bank could focus more on detail generation, further enhancing the output quality.

Comparisons with reference-based methods. We assess the efficacy of the new notion of GAN-based dictionary by comparing GLEAN with two representative methods adopting an imagery dictionary for SR – SRNTT and DFDNet . Examples are shown in Fig. 10.

For DFDNet, we evaluate the performance on LR images with unknown degradationsWe further downsample the LR images to 64×6464{\times}64 to match the input size of GLEAN.. Through pre-constructing a dictionary of facial components (\eg eyes, lips), DFDNet shows remarkable performance on face restoration. However, it cannot produce faithful results on parts absent in the dictionary, such as skin and hair. Therefore, significant incoherence is observed in the outputs. Despite GLEAN is trained on the bicubic kernel, it is still capable of producing appealing outputs. More importantly, GLEAN is not confined to improving the visual quality of specific components. Instead, the entire image is super-resolved, leading to coherent and pleasing results. The performance of GLEAN could be further improved by employing multiple degradations during training.

For SRNTT, we follow the same settings and downsample the ground-truth images using the bicubic kernel. With such low-resolution images (32×3232{\times}32), global matching becomes prohibitive, and hence SRNTT fails to transfer the textures from HR reference images. As a result, SRNTT tends to provide blurry textures. By capturing the distribution instead of specific imagery clues, GLEAN does not rely on any explicit textural transferal procedure. This enables the applicability to large-factor SR, where image matching is extremely difficult. More importantly, with no external images employed, GLEAN does not require any global matching to search for suitable textures/details. This allows GLEAN to be applied to images with larger resolutions, where global matching is computationally prohibitive.

Application – Image Retouching

In this section, we present one interesting application of GLEAN. In interactive image retouching, users can manually edit the images based on their preference. However, a perfect output requires tedious and precise retouching. As a result, artifacts are common in the outputs, especially those from amateur retouching. As a powerful super-resolver, GLEAN can be used as an image retouching tool to eliminate unpleasant artifacts.

As shown in Fig. 11, the blending operation in the interactive editing software produces a blurry and incoherent output. Thanks to the capability of GLEAN in producing high quality and fidelity images, GLEAN is able to eliminate the blurry region and generate a coherent output with natural textures. Furthermore, with only a single forward pass for generation, it can be easily incorporated into existing interactive editing software. More examples will be shown in the appendix.

Conclusion

We have presented a new way to exploit pre-trained GANs for the task of large-scale super-resolution, up to 64×64{\times} upscaling factor. We have shown that a pre-trained GAN can be used as a generative latent bank in an encoder-bank-decoder architecture. Reconstructing photorealistic HR images requires just a single forward pass, thanks to effective ways in conditioning and retrieving rich priors from the bank. The generality of the notion of GAN-based dictionary allows GLEAN to be potentially extended to not only diverse architectures but also various imaging tasks, such as image denoising, inpainting and colorization.

Acknowledgement. This research was conducted in collaboration with SenseTime and supported by the Singapore Government through the Industry Alignment Fund - Industry Collaboration Projects Grant. It is also partially supported by Singapore MOE AcRF Tier 1 (2018-T1-002-056) and NTU SUG.

References

Appendix A Training Details of GLEAN

We adopt pre-trained StyleGANGenForce: https://github.com/genforce/genforce or StyleGAN2BasicSR: https://github.com/xinntao/BasicSR as our generative latent bank. In this section, we assume the latent bank is pre-trained and present the training details of GLEAN (\ie the encoder-bank-decoder network). Note that the weights of the latent bank are fixed when training GLEAN to better employ the generative prior and to avoid biasing to the training distribution.

We train GLEAN on five categories including human faces, cats, cars, towers, and bedrooms. The training and test datasets used in our experiments are summarized in Table 3. Since StyleGAN produces images with fixed size, we resize the images in the datasets for our experiments.

Following previous works , the objective function for GLEAN consists of three terms. MSE loss is used to guide the fidelity of the output images:

where NN, y^\hat{y}, and yy denote the number of pixels, the output image, and the ground-truth image, respectively. We further incorporate perceptual loss and adversarial loss to improve the perceptual quality:

where f(⋅)f(\cdot) denotes the feature embedding space of the VGG16 network, and DD corresponds to the StyleGAN discriminator. The resulting objective function is a weighted mean of the three losses:

In all our experiments, we set αpercep=αgen=10−2\alpha_{percep}{=}\alpha_{gen}{=}10^{-2}. For the discriminator, we maximize

We adopt Cosine Annealing Scheme and Adam optimizer in training. The number of iterations is 300K and the initial learning rate is 10−410^{-4}. The batch size is 8 for human faces and 16 for other categories. We train our models using two Nvidia V100 GPUs.

Appendix B Qualitative Results

Randomly-Selected Examples. In Fig. 12, we show the results of randomly-selected examples from CelebA-HQ . By optimizing only the latent codes, PULSE produces outputs with low-fidelity. In contrast, guided by the encoder features and our generative latent bank, GLEAN achieves remarkable quality and fidelity, demonstrating the effectiveness of our designs.

Scale Factors and Categories. GLEAN is extensible to various scale factors (from 8×8{\times} to 64×64{\times}) and categories (\eg faces, cats, cars, bedrooms, towers). From Fig. 13 to Fig. 18, we see that GLEAN outperforms DGP and ESRGAN+ in both fidelity and quality. It is noteworthy that the performance of DGP and ESRGAN+ are less promising on categories other than human faces.

B.2 Image Retouching

In interactive image retouching, users can manually edit the images based on their preference. For instance, users can change the facial expression of an object and perform geometric transformations for enlarging eyes. However, a perfect output requires tedious and precise retouching. As a result, artifacts are common in the outputs from amateur retouching.

GLEAN allows the possibility of performing realistic refinement of imperfect retouching. More specifically, given a retouched image, we can first downsample the image to a smaller resolution, where the artifacts vanished. We can then upsample it back to the original resolution. With GLEAN as a powerful super-resolver, we can obtain an output with unnatural artifacts suppressed.

As shown in Fig. 19, GLEAN is able to correct the unnatural artifacts introduced by amateur retouching while being similar to the retouched images, realistic, and coherent with the unaltered regions. In addition, since GLEAN requires only a single forward pass, it can be used in interactive image editing software to allow a more flexible retouching.