Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild
Fanghua Yu, Jinjin Gu, Zheyuan Li, Jinfan Hu, Xiangtao Kong, Xintao Wang, Jingwen He, Yu Qiao, Chao Dong
Introduction
With the development of image restoration (IR), expectations for the perceptual effects and intelligence of IR results have significantly increased. IR methods based on generative priors leverage powerful pre-trained generative models to introduce high-quality generation and prior knowledge into IR, bringing significant progress in these aspects. Continuously enhancing the capabilities of the generative prior is key to achieving more intelligent IR results, with model scaling being a crucial and effective approach. There are many tasks that have obtained astonishing improvements from scaling, such as SAM and large language models . This further motivates our effort to build large-scale, intelligent IR models capable of producing ultra-high-quality images. However, due to engineering constraints such as computing resources, model architecture, training data, and the cooperation of generative models and IR, scaling up IR models is challenging.
In this work, we introduce SUPIR (Scaling-UP IR), the largest-ever IR method, aimed at exploring greater potential in visual effects and intelligence. Specifically, SUPIR employs StableDiffusion-XL (SDXL) as a powerful generative prior, which contains 2.6 billion parameters. To effectively apply this model, we design and train a adaptor with more than 600 million parameters. Moreover, we have collected over 20 million high-quality, high-resolution images to fully realize the potential offered by model scaling. Each image is accompanied by detailed descriptive text, enabling the control of restoration through textual prompts. We also utilize a 13-billion-parameter multi-modal language model to provide image content prompts, greatly improving the accuracy and intelligence of our method. The proposed SUPIR model demonstrates exceptional performance in a variety of IR tasks, achieving the best visual quality, especially in complex and challenging real-world scenarios. Additionally, the model offers flexible control over the restoration process through textual prompts, vastly broadening the possibility of IR. Fig. 1 illustrates the effects by our model, showcasing its superior performance.
Our work goes far beyond simply scaling. While pursuing an increase in model scale, we face a series of complex challenges. First, when applying SDXL for IR, existing Adaptor designs either too simple to meet the complex requirements of IR or are too large to train together with SDXL . To solve this problem, we trim the ControlNet and designed a new connector called ZeroSFT to work with the pre-trained SDXL, aiming to efficiently implement the IR task while reducing computing costs. In order to enhance the model’s ability to accurately interpret the content of low-quality images, we fine-tune the image encoder to improve its robustness to variations in image degradation. These measures make scaling the model feasible and effective, and greatly improve its stability. Second, we amass a collection of 20 million high-quality, high-resolution images with descriptive text annotations, providing a solid foundation for the model’s training. We adopt a counter-intuitive strategy by incorporating poor quality, negative samples into training. In this way, we can use negative quality prompts to further improve visual effects. Our results show that this strategy significantly improves image quality compared to using only high-quality positive samples. Finally, powerful generative prior is a double-edged sword. Uncontrolled generation may reduce restoration fidelity, making IR no longer faithful to the input image. To mitigate this low-fidelity issue, we propose a novel restoration-guided sampling method. All these strategies, coupled with efficient engineering implementation, are key to enabling the scaling up of SUPIR, pushing the boundaries of advanced IR. This comprehensive approach, encompassing everything from model architecture to data collection, positions SUPIR at the forefront of image restoration technology, setting a new benchmark for future advancements.
Related Work
The goal of IR is to convert degraded images into high-quality degradation-free versions . In the early stage, researchers independently explored different types of image degradation, such as super-resolution (SR) , denoising , and deblurring . However, these methods are often based on specific degradation assumptions and therefore lack generalization ability to other degradations . Over time, the need for blind restoration methods that are not based on specific degradation assumptions has grown . In this trend, some methods approximate synthesize real-world degradation by more complex degradation models, and are well-known for handling multiple degradation with a single model. Recent research, such as DiffBIR , unifies different restoration problems into a single model. In this paper, we adopt a similar setting to DiffBIR and use a single model to achieve effective processing of various severe degradations.
Generative Prior.
Generative priors are adept at capturing the inherent structures of the image, enabling the generation of images that follow natural image distribution. The emergence of GANs has underscored the significance of generative priors in IR. Various approaches employ these priors, including GAN inversion , GAN encoders , or using GAN as the core module for IR . Beyond GANs, other generative models can also serve as priors . Our work primarily focuses on generative priors derived from diffusion models , which excel in controllable generation and model scaling . Diffusion models have also been effectively used as generative priors in IR . However, these diffusion-based IR methods’ performance is constrained by the scale of the used generative models, posing challenges in further enhancing their effectiveness.
Model Scaling
is an important means to further improve the capabilities of deep-learning models. The most typical examples include the scaling of language models , text-to-image generation models , and image segmentation models . The scale and complexity of these models have increased dramatically, with billions or even hundreds of billions of parameters, but these parameters also lead to extraordinary performance improvements, demonstrating the potential of model scaling . However, scaling up is a systematic problem, involving model design, data collection, computing resources, and other limitations. Many other tasks have not yet been able to enjoy the substantial performance improvements brought by scaling up. IR is one of them.
Method
An overview of the proposed SUPIR method is shown in Fig. 2. We introduce our method from three aspects: Sec. 3.1 introduces our network designs and training method; Sec. 3.2 introduces the collection of training data and the introduction of text modality; and Sec. 3.3 introduces the diffusion sampling method for image restoration.
There are not many choices for the large-scale generative models. The only ones to consider are Imagen , IF , and SDXL . Our selection settled on SDXL for the following reasons. Imagen and IF prioritize text-to-image generation and rely on a hierarchical approach. They first generate small-resolution images and then hierarchically upsample them. SDXL directly generates a high-resolution image without hierarchical design, which is more aligned with our objectives, as it utilizes its parameters effectively for image quality improvement rather than text interpretation. Additionally, SDXL employs a Base-Refine strategy. In the Base model, diverse but lower-quality images are generated. Subsequently, the Refine model enhances the perceptual quality of these images. Compared to the Base model, the Refine model uses training images with significantly higher quality but less diverse. Considering our strategy to train with an extensive dataset of high-quality images, the two-phase design of SDXL becomes superfluous for our needs. We opt for the Base model, which has a greater number of parameters, making it an ideal backbone for our generative prior.
Degradation-Robust Encoder.
Large-Scale Adaptor Design.
Considering the SDXL model as our chosen prior, we need an adaptor that can steer it to restore images according to the provided LQ inputs. The Adaptor is required to identify the content in the LQ image and to finely control the generation at the pixel level. LoRA , T2I adaptor , and ControlNet are existing diffusion model adaptation methods, but none of them meet our requirements: LoRA limits generation but struggles with LQ image control; T2I lacks capacity for effective LQ image content identification; and ControlNet’s direct copy is challenging for the SDXL model scale. To address this issue, we design a new adaptor with two key features, as shown in Fig. 3(a). First, we keep the high-level design of ControlNet but employ network trimming to directly trim some blocks within the trainable copy, achieving an engineering-feasible implementation. Each block within the encoder module of SDXL is mainly composed of several Vision Transformer (ViT) blocks. We identified two key factors contributing to the effectiveness of ControlNet: large network capacity and efficient initialization of the trainable copy. Notably, even partial trimming of blocks in the trainable copy retains these crucial characteristics in the adaptor. Therefore, we simply trim half of the ViT blocks from each encoder block, as shown in Fig. 3(b). Second, we redesign the connector that links the adaptor to SDXL. While SDXL’s generative capacity delivers excellent visual effects, it also renders pixel-level precise control challenging. ControlNet employs zero convolution for generation guidance, but relying solely on residuals is insufficient for the control required by IR. To amplify the influence of LQ guidance, we introduced a ZeroSFT module, as depicted in Fig. 3(c). Building based on zero convolution, ZeroSFT encompasses an additional spatial feature transfer (SFT) operation and group normalization .
2 Scaling Up Training Data
The scaling of the model requires a corresponding scaling of the training data . But there is no large-scale high-quality image dataset available for IR yet. Although DIV2K and LSDIR offer high image quality, they are limited in quantity. Larger datasets like ImageNet (IN) , LAION-5B, and SA-1B contain more images, but their image quality does not meet our high standards. To this end, we collect a new large-scale dataset of high-resolution images, which includes 20 million 10241024 high-quality, texture-rich, and content-clear images. A comparison on scales of the collected dataset and the existing dataset is shown in Fig. 3. We also included an additional 70K unaligned high-resolution facial images from FFHQ-raw dataset to improve the model’s face restoration performance. In Fig. 5(a), we show the relative size of our data compared to other well-known datasets.
Multi-Modality Language Guidance.
Negative-Quality Samples and Prompt.
3 Restoration-Guided Sampling
Experiments
2 Comparison with Existing Methods
Our method can handle a wide range of degradations, and we compare it with the state-of-the-art methods with the same capabilities, including BSRGAN , Real-ESRGAN , StableSR , DiffBIR and PASD . Some of them are constrained to generating images of 512512 size. In our comparison, we crop the test image to meet this requirement and downsample our results to facilitate fair comparisons. We conduct comparisons on both synthetic data and real-world data.
To synthesize LQ images for testing, we follow previous works and demonstrate our effects on several representative degradations, including both single degradations and complex mixture degradations. Specific details can be found in Tab. 1. We selected the following metrics for quantitative comparison: full-reference metrics PSNR, SSIM, LPIPS , and the non-reference metrics ManIQA , ClipIQA , MUSIQ . It can be seen that our method achieves the best results on all non-reference metrics, which reflects the excellent image quality of our results. At the same time, we also note the disadvantages of our method in full-reference metrics. We present a simple experiment that highlights the limitations of these full-reference metrics, see Fig. 7. It can be seen that our results have better visual effects, but they do not have an advantage in these metrics. This phenomenon has also been noted in many studies as well . We argue that with the improving quality of IR, there is a need to reconsider the reference values of existing metrics and suggest more effective ways to evaluate advanced IR methods. We also show some qualitative comparison results in Fig. 6. Even under severe degradation, our method consistently produces highly reasonable and high-quality images that faithfully represent the content of the LQ images.
Restoration in the Wild.
We also test our method on real-world LQ images. We collect a total of 60 real-world LQ images from RealSR , DRealSR , Real47 , and online sources, featuring diverse content including animals, plants, faces, buildings, and landscapes. We show the qualitative results in Fig. 10, and the quantitative results are shown in LABEL:tab:real. These results indicate that the images produced by our method have the best perceptual quality. We also conduct a user study comparing our method on real-world LQ images, with 20 participants involved. For each set of comparison images, we instructed participants to choose the restoration result that was of the highest quality among these test methods. The results are shown in Fig. 8, revealing that our approach significantly outperformed state-of-the-art methods in perceptual quality.
3 Controlling Restoration with Textual Prompts
After training on a large dataset of image-text pairs and leveraging the feature of the diffusion model, our method can selectively restore images based on human prompts. Fig. 1(b) illustrates some examples. In the first case, the bike restoration is challenging without prompts, but upon receiving the prompt, the model reconstructs it accurately. In the second case, the material texture of the hat can be adjusted through prompts. In the third case, even high-level semantic prompts allow manipulation over face attributes. In addition to prompting the image content, we can also prompt the model to generate higher-quality images through negative-quality prompts. Fig. 11(a) shows two examples. It can be seen that the negative prompts are very effective in improving the overall quality of the output image. We also observed that prompts in our method are not always effective. When the provided prompts do not align with the LQ image, the prompts become ineffective, see Fig. 11(b). We consider this reasonable for an IR method to stay faithful to the provided LQ image. This reflects a significant distinction from text-to-image generation models and underscores the robustness of our approach.
4 Ablation Study
We compare the proposed ZeroSFT connector with zero convolution . Quantitative results are shown in LABEL:tab:connectors. Compared to ZeroSFT, zero convolution yields comparable performance on non-reference metrics and much lower full-reference performance. In Fig. 9, we find that the drop in non-reference metrics is caused by generating low-fidelity content. Therefore, for IR tasks, ZeroSFT ensures fidelity without losing the perceptual effect.
Training data scaling.
We trained our large-scale model on two smaller datasets for IR, DIV2K and LSDIR . The qualitative results are shown in Fig. 12, which clearly demonstrate the importance and necessity of training on large-scale high-quality data.
Negative-quality samples and prompt.
LABEL:tab:prompts shows some quantitative results under different settings. Here, we use positive words describing image quality as “positive prompt”, and use negative quality words and the CFG methods described in Sec. 3.2 as negative prompt. It can be seen that adding positive prompts or negative prompts alone can improve the perceptual quality of the image. Using both of them simultaneously yields the best perceptual results. If negative samples are not included for training, these two prompts will not be able to improve the perceptual quality. Fig. 4 and Fig. 11(a) demonstrate the improvement in image quality brought by using negative prompts.
Restoration-guided sampling method.
The proposed restoration-guided sampling method is mainly controlled by the hyper-parameter . The larger is, the fewer corrections are made to the generation at each step. The smaller is, the more generated content will be forced to be closer to the LQ image. Please refer to Fig. 13 for a qualitative comparison. When , the image is blurry because its output is limited by the LQ image and cannot generate texture and details. When , there is not much guidance during generation. The model generates a lot of texture that is not present in the LQ image, especially in flat area. Fig. 8(a) illustrates the quantitative results of restoration as a function of the variable . As shown in Fig. 8(a), decreasing from 6 to 4 does not result in a significant decline in visual quality, while fidelity performance improves. As restoration guidance continues to strengthen, although PSNR continues to improve, the images gradually become blurry with loss of details, as depicted in Fig. 13. Therefore, we choose as the default parameter, as it doesn’t significantly compromise image quality while effectively enhancing fidelity.
Conclusion
We propose SUPIR as a pioneering IR method, empowered by model scaling, dataset enrichment, and advanced design features, expanding the horizons of IR with enhanced perceptual quality and controlled textual prompts.
References
Appendix
Appendix A Discussions
As shown in Fig. 2 of the main text, a degradation-robust encoder is trained and deployed prior to feeding the low-quality input into the adaptor. We conduct experiments using synthetic data to demonstrate the effectiveness of the proposed degradation-robust encoder. In Fig. 14, we show the results of using the same decoder to decode the latent representations from different encoders. It can be seen that the original encoder has no ability to resist degradation and its decoded images still contain noise and blur. The proposed degradation-robust encoder can reduce the impact of degradation, which further prevents generative models from misunderstanding artifacts as image content .
A.2 LLaVA Annotation
Our diffusion model is capable of accepting textual prompts during the restoration process. The prompt strategy we employ consists of two components: one component is automatically annotated by LLaVA-v1.5-13B , and the other is a standardized default positive quality prompt. The fixed portion of the prompt strategy provides a positive description of quality, including words like “cinematic, High Contrast, highly detailed, unreal engine, taken using a Canon EOS R camera, hyper detailed photo-realistic maximum detail, 32k, Color Grading, ultra HD, extreme meticulous detailing, skin pore detailing, hyper sharpness, perfect without deformations, Unreal Engine 5, 4k render”. For the LLaVA component, we use the command “Describe this image and its style in a very detailed manner” to generate detailed image captions, as exemplified in Fig. 18. While occasional inaccuracies may arise, LLaVA-v1.5-13B generally captures the essence of the low-quality input with notable precision. Using the reconstructed version of the input proves effective in correcting these inaccuracies, allowing LLaVA to provide an accurate description of the majority of the image’s content. Additionally, SUPIR is effective in mitigating the impact of potential hallucination prompts, as detailed in .
A.3 Limitations of Negative Prompt
Figure 23 presents evidence that the use of negative quality prompts substantially improves the image quality of restored images. However, as observed in Fig. 15, the negative prompt may introduce artifacts when the restoration target lacks clear semantic definition. This issue likely stems from a misalignment between low-quality inputs and language concepts.
A.4 Negative Samples Generation
While negative prompts are highly effective in enhancing quality, the lack of negative-quality samples and prompts in the training data results in the fine-tuned SUPIR’s inability to comprehend these prompts effectively. To address this problem, in Sec. 3.2 of the main text, we introduce a method to distill negative concepts from the SDXL model. The process for generating negative samples is illustrated in Fig. 19. Direct sampling of negative samples through a text-to-image approach often results in meaningless images. To address this issue, we also utilize training samples from our dataset as source images. We create negative samples in an image-to-image manner as proposed in , with a strength setting of .
Appendix B More Visual Results
We provide more results in this section. Fig. 16 presents additional cases where full-reference metrics do not align with human evaluation. In Fig. 17, we show that using negative-quality prompt without including negative samples in training may cause artifacts. In Figs. 20, 21 and 22, we provide more visual caparisons with other methods. Plenty of examples prove the strong restoration ability of SUPIR and the most realistic of restored images. More examples of controllable image restoration with textual prompts can be found in Fig. 23.