Towards Robust Blind Face Restoration with Codebook Lookup Transformer
Shangchen Zhou, Kelvin C. K. Chan, Chongyi Li, Chen Change Loy
Introduction
Face images captured in the wild often suffer from various degradation, such as compression, blur, and noise. Restoring such images is highly ill-posed as the information loss induced by the degradation leads to infinite plausible high-quality (HQ) outputs given a low-quality (LQ) input. The ill-posedness is further elevated in blind restoration, where the specific degradation is unknown. Despite the progress made with the emergence of deep learning, learning a LQ-HQ mapping without additional guidance in the huge image space is still intractable, leading to the suboptimal restoration quality of earlier approaches. To improve the output quality, auxiliary information that 1) reduces the uncertainty of LQ-HQ mapping and 2) complements high-quality details is indispensable.
Various priors have been used to mitigate the ill-posedness of this problem, including geometric priors , reference priors , and generative priors . Although improved textures and details are observed, these approaches often suffer from high sensitivity to degradation or limited prior expressiveness. These priors provide insufficient guidance for face restoration, thus their networks essentially resort to the information of LQ input images that are usually highly corrupted. As a result, the LQ-HQ mapping uncertainty still exists, and the output quality is deteriorated by the degradation of the input images. Most recently, based on generative prior, some methods project the degraded faces into a continuous infinite space via iterative latent optimization or direct latent encoding . Despite great realness of outputs, it is difficult to find the accurate latent vectors in case of severe degradation, resulting in low-fidelity results (Fig. 1(d)). To enhance the fidelity, skip connections between encoder and decoder are usually required in this kind of methods , as illustrated in Fig. 1(a) (top), however, such designs meanwhile introduce artifacts in the results when inputs are severely degraded, as shown in Fig. 1(e).
Different from the aforementioned approaches, this work casts blind face restoration as a code prediction task in a small finite proxy space of the learned discrete codebook prior, which shows superior robustness to degradation as well as rich expressiveness. The codebook is learned by self-reconstruction of HQ faces using a vector-quantized autoencoder, which along with decoder stores the rich HQ details for face restoration. In contrast to continuous generative priors , the combinations of codebook items form a discrete prior space with only finite cardinality. Through mapping the LQ images to a much smaller proxy space (e.g., 1024 codes), the uncertainty of the LQ-HQ mapping is significantly attenuated, promoting robustness against the diverse degradation, as compared in Figs. 1(d-g). Besides, the codebook space possess greater expressiveness, which perceptually approximates the image space, as shown in Fig. 1(h). This nature allows the network to reduce the reliance on inputs and even be free of skip connections.
Though the discrete representation based on a codebook has been deployed for image generation , the accurate code composition for image restoration remains a non-trivial challenge. The existing works look up codebook via nearest-neighbor (NN) feature matching, which is less feasible for image restoration since the intrinsic textures of LQ inputs are usually corrupted. The information loss and diverse degradation in LQ images inevitably distort the feature distribution, prohibiting accurate feature matching. As depicted in Fig. 1(b) (right), even after fine-tuning the encoder on LQ images, the LQ features cannot cluster well to the exact code but spread into other nearby code clusters, thus the nearest-neighbor matching is unreliable in such cases.
Tailored for restoration, we propose a Transformer-based code prediction network, named CodeFormer, to exploit global compositions and long-range dependencies of LQ faces for better code prediction. Specifically, taking the LQ features as input, the Transformer module predicts the code token sequence which is treated as the discrete representation of the face images in the codebook space. Thanks to the global modeling for remedying the local information loss in LQ images, the proposed CodeFormer shows robustness to heavy degradation and keeps overall coherence. Comparing the results presented in Figs. 1(f-g), the proposed CodeFormer is able to recover more details than the nearest-neighbor matching, such as the glasses, improving both quality and fidelity of restoration.
Moreover, we propose a controllable feature transformation module with an adjustable coefficient to control the information flow from the LQ encoder to decoder. Such design allows a flexible trade-off between restoration quality and fidelity so that the continuous image transitions between them can be achieved. This module enhances the adaptiveness of CodeFormer under different degradations, e.g., in case of heavy degradation, one could manually reduce the information flow of LQ features carrying degradation to produce high-quality results.
Equipped with the above components, the proposed CodeFormer demonstrates superior performance in existing datasets and also our newly introduced WIDER-Test dataset that consists of 970 severely degraded faces collected from the WIDER-Face dataset . In addition to face restoration, our method also demonstrates its effectiveness on other challenging tasks such as face inpainting, where long-range clues from other regions are required. Systematic studies and experiments are conducted to demonstrate the merits of our method over previous works.
Related Work
Blind Face Restoration. Since face is highly structured, geometric priors of faces are exploited for blind face restoration. Some methods introduce facial landmarks , face parsing maps , facial component heatmaps , or 3D shapes in their designs. However, such prior information cannot be accurately acquired from degraded faces. Moreover, geometric priors are unable to provide rich details for high-quality face restoration.
Reference-based approaches have been proposed to circumvent the above limitations. These methods generally require the references to possess same identity as the input degraded face. For example, Li et al. propose a guided face restoration network that consists of a warping subnetwork and a reconstruction subnetwork, and a high-quality guided image of the same identity as input is used for better restoring the facial details. However, such references are not always available. DFDNet pre-constructs dictionaries composed of high-quality facial component features. However, the component-specific dictionary features are still insufficient for high-quality face restoration, especially for the regions out of the dictionary scope (e.g., skin, hair). To alleviate this issue, recent VQGAN-based methods explores a learned HQ dictionary, which contains more generic and rich details face restoration.
Recently, the generative facial priors from pre-trained generators, e.g., StyleGAN2 , have been widely explored for blind face restoration. It is utilized via different strategies of iterative latent optimization for effective GAN inversion or direct latent encoding of degraded faces . However, preserving high fidelity of the restored faces is challenging when one projects the degraded faces into the continuous infinite latent space. To relieve this issue, GLEAN , GPEN , and GFPGAN embed the generative prior into encoder-decoder network structures, with additional structural information from input images as guidance. Despite the improvement of fidelity, these methods highly rely on inputs through the skip connections, which could introduce artifacts to results when inputs are severely corrupted.
Dictionary Learning. Sparse representation with learned dictionaries has demonstrated its superiority in image restoration tasks, such as super-resolution and denoising . However, these methods usually require an iterative optimization to learn the dictionaries and sparse coding, suffering from high computational cost. Despite the inefficiency, their high-level insight into exploring a HQ dictionary has inspired reference-based restoration networks, e.g., LUT and self-reference , as well as synthesis methods . Jo and Kim construct a look-up table (LUT) by transferring the network output values to a LUT, so that only a simple value retrieval is needed during inference. However, storing HQ textures in the image domain usually requires a huge LUT, limiting its practicality. VQVAE is first to introduce a highly compressed codebook learned by a vector-quantized autoencoder model. VQGAN further adopts the adversarial loss and perceptual loss to enhance perceptual quality at a high compression rate, significantly reducing the codebook size without sacrificing its expressiveness. Unlike the large hand-crafted dictionary , the learnable codebook automatically learns optimal elements for HQ image reconstruction, providing superior efficiency and expressiveness as well as circumventing the laborious dictionary design. Inspired by the codebook learning, this paper investigates a discrete proxy space for blind face restoration. Different from recent VQGAN-based approaches , we exploit the discrete codebook prior by predicting code sequences via global modeling, and we secure prior effectiveness by fixing the encoder. Such designs allow our method to take full advantage of the codebook so that it does not depend on the feature fusion with LQ cues, significantly enhancing the robustness of face restoration.
Methodology
The main focus of this work is to exploit a discrete representation space that reduces the uncertainty of restoration mapping and complements high-quality details for the degraded inputs. Since local textures and details are lost and corrupted in low-quality inputs, we employ a Transformer module to model the global composition of natural faces, which remedies the local information loss, enabling high-quality restoration. The overall framework is illustrated in Fig. 2.
We first incorporate the idea of vector quantization and pre-train a quantized autoencoder through self-reconstruction to obtain a discrete codebook and the corresponding decoder (Sec. 3.1). The prior from the codebook combination and decoder is then used for face restoration. Based on this codebook prior, we then employ a Transformer for accurate prediction of code combination from the low-quality inputs (Sec. 3.2). In addition, a controllable feature transformation module is introduced to exploit a flexible trade-off between restoration quality and fidelity (Sec. 3.3). The training of our method is divided into three stages accordingly.
To reduce uncertainty of the LQ-HQ mapping and complement high-quality details for restoration, we first pre-train the quantized autoencoder to learn a context-rich codebook, which improves network expressiveness as well as robustness against degradation.
The decoder then reconstructs the high-quality face image given . The code token sequence forms a new latent discrete representation that specifies the respective code index of the learned codebook, i.e., when .
Training Objectives. To train the quantized autoencoder with a codebook, we adopt three image-level reconstruction losses: L1 loss , perceptual loss , and adversarial loss :
where denotes the feature extractor of VGG19 . Since, image-level losses are underconstrained when updating the codebook items, we also adopt the intermediate code-level loss to reduce the distance between codebook and input feature embeddings :
where stands for the stop-gradient operator and is a weight trade-off for the update rates of the encoder and codebook. Since the quantization operation in Eq. (1) is non-differentiable, we adopt straight-through gradient estimator to copy the gradients from the decoder to the encoder. The complete objective of codebook prior learning is:
where is set to 0.8 in our experiments.
Codebook Settings. Our encoder and decoder consist of 12 residual blocks and 5 resize layers for downsampling and upsampling, respectively. Hence we obtain a large compression ratio of , which leads to a great robustness against degradation and a tractable computational cost for our global modeling in Stage II. Although more codebook items could ease reconstruction, the redundant elements could cause ambiguity in subsequent code predictions. Hence, we set the item number of codebook to 1024, which is sufficient for accurate face reconstruction. Besides, the code dimension is set to 256.
2 Codebook Lookup Transformer Learning (Stage II)
Due to corruptions of textures in LQ faces, the nearest-neighbour (NN) matching in Eq. (1) usually fails in finding accurate codes for face restoration. As depicted in Fig. 1(b), LQ features with diverse degradation could deviate from the correct code and be grouped into nearby clusters, resulting in undesirable restoration results, as shown in Fig. 1(f). To mitigate the problem, we employ a Transformer to model global interrelations for better code prediction.
Training Objectives. We train Transformer module as well as finetune the encoder for restoration, while the codebook and decoder are kept fixed. Instead of employing reconstruction loss and adversarial loss in the image-level, only code-level losses are required in this stage: 1) cross-entropy loss for code token prediction supervision, and 2) L2 loss to force the LQ feature to approach the quantized feature from codebook, which eases the difficulty of token prediction learning:
where the ground truth of latent code is obtained from the pre-trained autoencoder in Stage I (Sec. 3.1), thus the quantized feature is then retrieved from codebook according to the . The final objective of Transformer learning is:
where is set to 0.5 in our method. Note that our network after this stage has already equipped with superior robustness and effectiveness in face restoration, especially for severely degraded faces.
3 Controllable Feature Transformation (Stage III)
Despite our Stage II has obtained a great face restoration model, we also investigate a flexible tradeoff between quality and fidelity of face restoration. Thus, we propose the controllable feature transformation (CFT) module to control information flow from LQ encoder to decoder . Specifically, as shown in Fig. 2, the LQ features are used to slightly tune the decoder features through spatial feature transformation with the affine parameters of and . An adjustable coefficient is then used to control the relative importance of the inputs:
where denotes a stack of convolutions that predicts and from the concatenated features of . We adopt the CFT modules at multiple scales between encoder and decoder. Such a design allows our network to remain high fidelity for mild degradation and high quality for heavy degradation. Specifically, one could use a small to reduce the reliance on input LQ images with heavy degradation, thus producing high-quality outputs. The larger will introduce more information from LQ images to enhance the fidelity in case of mild degradation.
Training Objectives. To train the controllable modules and finetune the encoder in this stage, we keep the code-level losses of in Stage II, and also add image-level losses of , , and , which are the same as that in Stage I except that is replaced by restoration output . The complete loss for this stage is the sum of above losses weighted with their original weight factors. We set the to 1 during training of this stage, which then allows network to achieve continuous transitions of results by adjusting in $w=0.5$ by default to make a good balance between the quality and fidelity of outputs.
Experiments
Training Dataset. We train our models on the FFHQ dataset , which contains 70,000 high-quality (HQ) images, and all images are resized to for training. To form training pairs, we synthesize LQ images from the HQ counterparts with the following degradation model :
where the HQ image is first convolved with a Gaussian kernel , followed by a downsampling of scale . After that, additive Gaussian noise is added to the images, and then JPEG compression with quality factor is applied. Finally, the LQ image is resized back to . We randomly sample , , , and from , , respectively.
Testing Datasets. We evaluate our method on a synthetic dataset CelebA-Test and three real-world datasets: LFW-Test, WebPhoto-Test, and our proposed WIDER-Test. CelebA-Test contains 3,000 images selected from the CelebA-HQ dataset , where LQ images are synthesized under the same degradation range as our training settings. The three real-world datasets respectively contain three different degrees of degradation, i.e., mild (LFW-Test), medium (WebPhoto-Test), and heavy (WIDER-Test). LFW-Test consists of the first image of each person in LFW dataset , containing 1,711 images. WebPhoto-Test consists of 407 low-quality faces collected from the Internet. Our WIDER-Test comprises 970 severely degraded face images from the WIDER Face dataset , providing a more challenging dataset to evaluate the generalizability and robustness of blind face restoration methods.
2 Experimental Settings and Metrics
Settings. We represent a face image of as a code sequence. For all stages of training, we use the Adam optimizer with a batch size of 16. We set the learning rate to for stages I and II, and adopt a smaller learning rate of for stage III. The three stages are trained with 1.5M, 200K, and 20K iterations, respectively. Our method is implemented with the PyTorch framework and trained using four NVIDIA Tesla V100 GPUs.
Metrics. For the evaluation on CelebA-Test with ground truth, we adopt PSNR, SSIM, and LPIPS as metrics. We also evaluate the identity preservation using the cosine similarity of features from ArcFace network , denoted as IDS. For the evaluation on real-world datasets without ground truth, we employ the widely-used non-reference perceptual metrics: FID and MUSIQ (KonIQ) .
3 Comparisons with State-of-the-Art Methods
We compare the proposed CodeFormer with state-of-the-art methods, including PULSE , DFDNet , PSFRGAN , GLEAN , GFP-GAN , and GPEN . We conduct extensive comparisons on both synthetic and real-world datasets.
Evaluation on Synthetic Dataset. We first show the quantitative comparison on the CelebA-Test in Table 1. In terms of the image quality metrics LPIPS, FID, and MUSIQ, our CodeFormer achieves the best scores than existing methods. Besides, it also faithfully preserves the identity, reflected by the highest IDS score and PSNR. Additionally, we present the qualitative comparison in Fig. 3. The compared methods fail to produce pleasant restoration results, e.g., DFDNet , PSFRGAN , GFP-GAN , and GPEN introduce obvious artifacts and GLEAN produces over-smoothed results that lack facial details. Moreover, all compared methods are unable to preserve the identity. Thanks to the expressive codebook prior and global modeling, CodeFormer not only produces high-quality faces but also preserves the identity well, even when inputs are highly degraded.
Evaluation on Real-world Datasets. As presented in Table 2, our CodeFormer achieves comparable perceptual quality of FID score with the compared methods on the real-world testing datasets with mild and medium degradation, and the best score on the testing dataset with heavy degradation. Although PULSE also obtains good perceptual MUSIQ score, it cannot preserve the identity of input images, as the identity score of IDS and visual results respectively suggested in Table 1 and Fig. 4. From the visual comparisons in Fig. 4, it is observed that our method shows exceptional robustness to the real heavy degradation and produces most visually pleasing results. Notably, CodeFormer successfully preserves the identity and produces natural results with rich details.
4 Ablation Studies
Effectiveness of Codebook Space. We first investigate the effectiveness of the codebook space. As shown in Exp. (a) of Table 3, removing the codebook (i.e., directly feeding the encoder features to the decoder) results in worse LPIPS and IDS scores. The results suggest that the discrete space of codebook is the key to ensure the robustness and effectiveness of our model.
Superiority of Transformer-based Code Prediction. To verify the superiority of our Transformer-based code prediction for codebook lookup, we compare it with two different solutions, i.e., nearest-neighbour (NN) matching, i.e., Exp. (b), and a CNN-based code prediction module, i.e., Exp. (c),
that adopts a Linear layer for prediction following encoder . As shown in Table 3, the comparison of Exps. (b) and (c) indicates that adopting code prediction for codebook lookup is more effective than NN feature matching. However, the local nature of convolution operation of CNNs restricts its modeling capability for long code sequence prediction. In comparison to the pure CNN-based method, i.e., Exp. (c), our Transformer-based solution produces better-fidelity results in terms of LPIPS and IDS scores, as well as higher accuracy of code prediction under all degradation degrees, as shown in Fig. 6. Besides, the superiority of CodeFormer is also demonstrated in visual comparisons shown in Fig. 5 and Fig. 9.
Flexibility of Controllable Feature Transformation Module. Considering the diverse degradation in real-world LQ face images, we provide a controllable feature transformation module (CFT) to allow a flexible trade-off between quality and fidelity. As shown in Fig. 7, a smaller tends to produce a high-quality result while a larger improves the fidelity. While such a flexibility is rarely explored in previous work, here we show that it is an appealing strategy to improves the adaptiveness of our method for different scenarios. As shown in Table 3, Exp. (f), i.e., setting the coefficient to increases the reconstruction and identity scores but decreases the visual quality. In this work, we trade between the quality and fidelity, and set the coefficient to by default.
5 Running time
We compare the running time of state-of-the-art methods and the proposed CodeFormer. All existing methods are evaluated on face images using their publicly available code. As shown in Table 4.5, the proposed CodeFormer has a similar running time as PSFRGAN and GPEN that can infer one image within 0.1s. Meanwhile, our method achieves the best performance in terms of LPIPS on the Celeb-Test dataset.
6 Extensions
Face Color Enhancement. We finetune our model on face color enhancement using the same color augmentations (random color jitter and grayscale conversion) as GFP-GAN (v1) . We compare our method with GFP-GAN (v1) on the real-world old photos (from CelebChild-Test dataset ) with color loss. The proposed CodeFormer generates high-quality face images with more natural color and faithful details.
Face Inpainting. The proposed Codeformer can be easily extended to face inpainting, and it shows great performance even in large mask ratios. To build training pairs, we use a publicly available script to randomly draw irregular polyline masks for generating masked faces. We compare our method with two state-of-the-art face inpainting methods CTSDG and GPEN , as well as Nearest-Neighbor matching for codebook lookup. As shown in Fig. 9, CTSDG and GPEN struggle in cases with large masks. Using Nearest-Neighbor matching within our framework roughly reconstructs the face structures, but it also fails in restoring complete visual parts such as the glasses and the eyes. In contrast, our CodeFormer generates high-quality natural faces without strokes and artifacts.
7 Limitation
Our method is built on a pre-trained autoencoder with a codebook. Thus, the capability and expressiveness of the autoencoder could affect the performance of our method. 1) Though the identity inconsistency issue is significantly relieved by the Transformer’s global modeling, inconsistency still exists in some rare visual parts such as accessories, where the current codebook space cannot seamlessly represent the image space. Using multiple scales in the codebook space to explore more fine-grained visual quantization may be a solution. 2) Although CodeFormer exhibits great robustness in most cases, when it comes to side faces, CodeFormer offers limited superiority to other methods and also cannot produce good results, as failure cases shown in Fig. 10. This is expected because there are only few side faces in the FFHQ training dataset, thus, the codebook is unable to learn sufficient codes for this case, leading to less effectiveness in both reconstruction and restoration.
Conclusion
This paper aims to address the fundamental challenges in blind face restoration. With a learned small discrete but expressive codebook space, we turn face restoration to code token prediction, significantly reducing the uncertainty of restoration mapping and easing the learning of restoration network. To remedy the local loss, we explore global composition and dependency from degraded faces via an expressive Transformer module for better code prediction. Benefiting from these designs, our method shows great expressiveness and strong robustness against heavy degradation. To enhance the adaptiveness of our method for different degradation, we also propose a controllable feature transformation module that allows a flexible trade-off between fidelity and quality. Experimental results suggest the superiority and effectiveness of our method.
Acknowledgement
This study is supported under the RIE2020 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s). It is also partially supported by the NTU NAP grant.
References
A More Discussions on CodeFormer
The proposed CodeFormer is affected by the reconstruction capability of learned codebook prior in Stage I. The better reconstruction capability of codebook, the greater the expressiveness of our method. We investigate the effect of the different numbers of learned codebook for face image reconstruction. Table A.1 shows the reconstruction results of LPIPS and PSNR on FFHQ dataset when different numbers of codebooks are learned. The reconstruction performance is better as more codebook items are activated and learned. Thus, we adopt a larger codebook with 1024 items.
A.2 Transformer Structure
The main novelty of our work is casting face restoration as a code token prediction task. Therefore, improving the accuracy of code prediction is significant in our method. The proposed CodeFormer employs a Transformer module to model the global interrelations of low-quality faces for better code prediction. Since the Transformer capability could be affected by its structure, we investigated different Transformer structures in pursuit of the best choice. Table A.2 shows the accuracy of code prediction and LPIPS scores for the comparison of different structures.
Single-scale and Multi-scale Inputs. We design two different Transformer structures that respectively take the single-scale () and multi-scale () features of the encoder as input. For multi-scale inputs, we first use a PixelUnshuffle layer to rearrange the features to the resolution of , and then utilize a Convolutional layer with a kernel of size to reduce feature channel to the same dimension. Finally, we add up all the reshaped multi-scale features and feed them to the Transformer . As shown in Table A.2, multi-scale input features do not boost the performance of code prediction, but slightly reduce the accuracy. An explanation is that degradation still exists in the larger-scale features of the shallow layers, which affects the accuracy of code prediction.
Number of Transformer Layer. We conduct an ablation study on Transformers with different numbers of layers. Table A.2 shows that more Transformer layers can boost code prediction accuracy and enhance the restoration performance. But the number of Transformer layer beyond nine only leads to slight performance gains of code prediction, and the LPIPS score is not boosted. The ablation study confirms our choice of using nine Transformer layers in CodeFormer, which gives a better trade-off among the computational complexity and performance.
B More Results on Blind Face Restoration
In this section, we provide more visual comparisons with state-of-the-art methods, including DFDNet , PSFRGAN , GLEAN , GFP-GAN , and GPEN . As shown in Fig. 11, the proposed CodeFormer not only produces high-quality faces but also preserves the identifies well, even when the input faces are highly degraded. Moreover, compared with other methods, the proposed CodeFormer is able to recover richer details and produce more natural faces.
C More Results on Extensions
The proposed Codeformer can be easily extended to other tasks, including face inpainting, face color enhancement, old photo enhancement, and AI-generated image correction. In this section, we compare our CodeFormer with state-of-the-art methods on face color enhancement (Sec. C.1) and face inpainting (Sec. C.2). Besides, we present the enhanced results of old historic photos (Sec. C.3) and fixed results of AI-arts (Sec. C.4).
We finetune our model on face color enhancement using the same color augmentations (random color jitter and grayscale conversion) as GFP-GAN . We evaluate and compare our method with GFP-GAN (version 1)Version 1 of GFP-GAN is jointly trained on blind face restoration and color enhancement. on the real-world old photos (from CelebChild-Test dataset ) with color loss. The proposed CodeFormer generates high-quality face images with more natural color and faithful details.
C.2 Face Inpainting
For face inpainting, we retrain the Codeformer on the synthetic paired data that generate masked faces by randomly drawing irregular polyline masks . We compare our method with two state-of-the-art face inpainting methods of CTSDG and GPEN and give more results in Fig. 13. The proposed CodeFormer generates more natural face images without strokes and artifacts.
C.3 Old Photo Enhancement
We evaluate our method on an old historic photo of the 5-th Solvay conference that was taken in 1927. For inference, we first crop out and align all faces, then enhance faces using the proposed CodeFormer, and finally paste back the enhanced faces to the original photo. Fig. 14 shows both overall results and cropped face results.
C.4 AI-Generated Image Correction
The proposed CodeFormer can also be used to fix faces in AI-generated artwork. Fig. 15 shows two examples that generated by Stable Diffusion and fixed results by CodeFormer.