ReFIR: Grounding Large Restoration Models with Retrieval Augmentation

Hang Guo, Tao Dai, Zhihao Ouyang, Taolin Zhang, Yaohua Zha, Bin Chen, Shu-tao Xia

Introduction

Restoring a high-quality image (HQ) from its low-quality counterpart (LQ) is a well-known ill-posed problem and has been studied over the years . Previous efforts attempt to handle this problem through employing various neural network architectures, including CNNs, GANs and Transformers. Recently, diffusion models have emerged as a promising alternative, delivering noteworthy results in real-world image restoration . In particular, some works have successfully leveraged the powerful generative prior of pre-trained text-to-image (T2I) diffusion models for scaling up, to obtain the Large Restoration Model (LRM) with billions of parameters, bringing significant progress in restoring photo-realistic images.

Although scaling up restoration models has achieved remarkable success, existing LRMs may not always produce results that are faithful to the original scene, particularly when faced with heavily degraded images that surpass the LRMs’ capabilities (see Fig. 1). This issue is similar to the hallucination problem observed in large language models (LLMs) , e.g. ChatGPT might generate nonsense responses when highly specialized questions exceed its knowledge boundary. Similarly, if one LRM has never seen a specific scene, it will struggle to restore corresponding images faithfully. By analogizing LLM to LRM, we define the phenomenon where LRMs generate textures inconsistent with the original scene when facing hard samples as the hallucination of LRMs.

To address the hallucination problem in LRMs, simply expanding the internal knowledge through additional training data and parameters might seem straightforward, but it can significantly increase computational and storage costs. Instead, this work considers another orthogonal strategy that enhances the external knowledge of LRMs without adding parameter counts. Drawing inspiration from the retrieval-augmented generation (RAG) used in LLMs , we aim to use the retrieved high-quality content-relevant images as external knowledge to alleviate the hallucination of LRMs. However, applying RAG to image restoration poses specific challenges. Specifically, in natural language, simply feeding the retrieved documents along with the original user query to LLMs can allow it to produce grounded responses. However, in the context of image restoration, allowing low-quality images to attend to retrieved images during their restoration process is non-trivial, which motivates us to develop novel techniques to enable LRMs to utilize external knowledge in restoration.

To this end, we delve deep into the working mechanisms of LRMs for insightful observations. Details of the experimental setup are described in Sec. 3. Our key findings indicate that the workflow of LRMs can be divided into two distinct stages: the Denoising Structure Reconstruction stage, during which the self-attention in the ControlNet reconstructs a clear overall structure from the noised representation. After that, in the Detail Texture Restoration stage, the self-attention in the UNet decoder fills scene-specific textures based on the denoised structure map. Based on these findings, a natural solution emerged: we can transfer high-quality, scene-specific textures from the retrieved images to the low-quality images during the detail texture restoration stage. In this way, the restored image is allowed a consistent texture with the retrieved image, thus mitigating the hallucination.

Inspired by the above observation, in this work, we propose the Retrieval-augmented Framework for Image Restoration, dubbed ReFIR, to offer a simple but effective way to expand the knowledge boundary of LRMs using the external knowledge from the retrieved images. Specifically, we first construct the retriever which employs the nearest neighbor lookup in the semantic embedding space to retrieve content-relevant reference images in the high-quality image database. After that, we develop the cross image injection which modifies the self-attention layer of original LRMs to enable the queries from the low-quality denoising chain to attend to the keys and values from the denoising chain of retrieved reference. To avoid the domain preference problem during injection, we propose separate attention to perform intra-chain and inter-chain attention, respectively. Given the spatial misalignment between the LQ and the retrieved HQ, we further adopt spatial adaptive gating to mask meaningless pixels during injection. At last, we employ the distribution alignment to narrow the domain gap between LQ and retrieved images. Thanks to the proposed ReFIR, the restoration of the LQ image can make full use of the external knowledge from the reference to generate high-fidelity images. Notably, the proposed pipeline is training-free and can be applied to multiple LRMs.

The contribution of this paper can be summarized as: (i) We introduce retrieval-augmented restoration, a novel concept to mitigate the hallucination problems in existing LRMs. (ii) We conduct an in-depth analysis of the working mechanisms of LRMs, based on which we propose a training-free framework to utilize the retrieved images. (iii) Extensive experiments validate that our proposed method effectively mitigates hallucination and is applicable to a broad spectrum of existing LRMs.

Related Works

Diffusion models have recently achieved significant advancements across various computer vision tasks . In the realm of image restoration, early explorations often involved training diffusion models from scratch to obtain the restoration tailored models . While these models are capable of producing high-fidelity results, they usually fall short of generating perceptually pleasing images. To leverage the powerful generative capabilities of large pre-trained text-to-image diffusion models like Stable Diffusion , recent attempts have focused on using the ControlNet with a LQ image as the condition to generate HQ images. Benefiting from the scaling law , these large restoration models with billions of parameters have shown impressive restoration results with photo-realistic textures and details. However, similar to the large language models, when the user query, i.e., the LQ image in this setting, exceeds the knowledge boundary of the large models, the models often fail to generate meaningful or correct responses, which is unacceptable for image restoration tasks that pursue high-fidelity.

2 Reference-based Image Super-resolution

Compared with single image super-resolution , Reference-based Image Super-Resolution (RefSR) can achieve enhanced performance by employing content-similar reference images as the additional input, and has attracted great research interests in the past few years . For instance, C2-Matching introduces a teacher-student correlation distillation and a dynamic DCN aggregation module for more precise alignment between low-quality and reference images. Following this, DATSR employs reciprocal learning and SwinTransformer to further boost performance. Additionally, MRefSR introduces a simple baseline to facilitate RefSR with multiple reference images. It is worth mentioning that despite both using additional images as references, our proposed retrieval augmented restoration pipeline differs from previous RefSR methods in several key aspects. Firstly, current RefSR models are typically small-scale due to limited training data, leading to performance degradation under challenging real-world conditions. Secondly, most RefSR methods can only use one single reference image and even fail to work in the absence of reference images. Thirdly, different from RefSR models that require training, our method can inject image-specific external knowledge into LRMs in a training-free manner. We give a detailed discussion about the difference in Appendix B.

3 Retrieval Augmented Generation

In the domain of natural language processing, Retrieval-Augmented Generation (RAG) leverages the strengths of pre-trained Large Language Models (LLMs) combined with knowledge retrieved from an external document database to enhance the quality of generated content . Typically, a RAG system initially retrieves documents relevant to the user’s query from the knowledge base and then integrates the retrieved document along with the original user query into the LLMs without any tuning to generate a response. Even when no relevant document is available, this system can still operate by using the internal knowledge embedded in the LLMs’ parameters. The integration of RAG allows LLMs to produce outputs that are not only contextually rich but also factually accurate, effectively mitigating the hallucination problem in knowledge-intensive tasks . In this work, we extend the concept of RAG to image processing and propose retrieval-augmented restoration to alleviate the hallucination issues in LRMs. By utilizing external textures embedded in the retrieved reference images, our tuning-free framework significantly facilitates faithful restoration results.

Probing Large Restoration Models

In order to manipulate the LRM so that it can utilize the retrieved reference images as external knowledge, we first delve into the underlying mechanism of existing LRMs to find useful insights. We choose the current popular LRM method SUPIR as a representative. Inspired by previous image editing efforts , which show that the self-attention layer of diffusion models contains important spatial correlation of an image, we thus follow this clue and employ the PCA to visualize the principal components of the latent from self-attention layers of SUPIR. We further utilize the Fourier analysis to allow for quantitative results. The results are shown in Fig. 2.

It can be seen that the ControlNet of the LRM can denoise the latent as the layers deepen, facilitating the reconstruction of a clear overall structure. However, this process is accompanied by a reduction in the high-frequency meaningful texture of the original image. This qualitative visualization can be also verified by the frequency characteristic plots, with high-frequency components decaying as layer number increases. On the other hand, the role of the UNet decoder is significantly different. Based on the previous clear structural map, the decoder restores the high-frequency details and textures with the help of skip connections, which is also shown through the strengthening high-frequency component in the decoder’s frequency curve.

Considering the above observations, we can divide the image restoration process of the LRM into two phases: the Denosing Structure Reconstruction phase in the ControlNet, and the Detail Texture Restoration phase in the UNet decoder. Inspired by these probing experiments, in this work, we employ the detail texture restoration nature in the self-attention layer of the decoder to inject the high-fidelity textures of retrieved images into the restoration process of the low-quality image.

Methodology

This work considers using retrieved reference images as an explicit part of the model. In contrast to the existing restoration pipeline, our ReFIR is parameterized by not only the internal knowledge from the network weights but also the external knowledge retrieved from suitable data representations. Fig. 3 gives an overview of our ReFIR. In the following part, we will first give the technical details of the retriever for reference image retrieval in Sec. 4.1, followed by the cross image injection to inject the external data knowledge into the restoration process of LRMs in Sec. 4.2.

In this work, we implement a conceptually simple solution of R\mathcal{R}, which uses the query image ILQI_{LQ} to retrieve its kk nearest neighbor in D\mathcal{D} using cosine similarity in the compact feature space derived from any feature extractors, such as VGG , ResNet or CLIP . Since the D\mathcal{D} is fixed, in practice, we can pre-extract and store the compact feature before training. Given a sufficiently large database D\mathcal{D}, this strategy ensures that the set of neighbors IR\mathbf{I_{R}} shares sufficient semantic consistency with ILQI_{LQ} and thus provides useful visual information for the restoration. Although this scheme seems simple, we show that it is efficient and effective, please see Sec. 5.3 for discussion.

2 Cross Image Injection for High-fidelity Image Restoration

Given the retrieved reference images IR=R(ILQ,D)\mathbf{I_{R}}=\mathcal{R}(I_{LQ},\mathcal{D}), we further propose the cross image injection to allow the original LRMs to use the external knowledge from IR\mathbf{I_{R}}. As shown in Fig. 4, we first construct two parallel denoising chains: the target restoration chain CT\mathcal{C}_{T} which is used to restore ILQI_{LQ}, and the source reference chain CS\mathcal{C}_{S} which unfolds IR\mathbf{I_{R}} into denoising time steps. After that, we introduce separate attention to separately perform attention within and between chains, followed by spatial adaptive gating to filter out irrelevant pixels. At last, we use the distribution alignment to mitigate the domain gap between chains. More details are given below.

Separate attention. To allow the CT\mathcal{C}_{T} to learn the knowledge from the CS\mathcal{C}_{S}, an effective interaction between the latents is crucial. Inspired by the observation in Sec. 3, we aim to transfer the knowledge embedded in the self-attention layer of CS\mathcal{C}_{S}’s decoder to the counterpart of CT\mathcal{C}_{T}. To this end, we modify the original self-attention in CT\mathcal{C}_{T} to our proposed separate attention. The core idea of our separate attention is to add “inter-chain cross-attention” to the original “intra-chain self-attention” so that CT\mathcal{C}_{T} can attend high-quality texture knowledge from CS\mathcal{C}_{S} while preserving its original features. As shown in Fig. 4(a), formally, denote QTQ_{T}, KTK_{T}, VTV_{T} as the query, key and value from the CT\mathcal{C}_{T}, and KSK_{S}, VSV_{S} as the key and value from the CS\mathcal{C}_{S}, the intra-chain self-attention preserves the original attention of CT\mathcal{C}_{T} to obtain the output OintraO_{intra}, and the inter-chain cross-attention uses the QTQ_{T} to query the KSK_{S} and VSV_{S} to facilitate CT\mathcal{C}_{T} utilizing the knowledge from CS\mathcal{C}_{S} to get the result OinterO_{inter}. In short, the proposed separate attention can be formalized as follows:

It is worth mentioning that directly using QTQ_{T} to query the concatenate results of KTK_{T} and KSK_{S} can only yield sub-optimal results due to the domain preference issue, i.e., QTQ_{T} will prefer latent from the same domain CT\mathcal{C}_{T} even though CS\mathcal{C}_{S} is more helpful for reconstruction. By using the proposed separate attention, the QTQ_{T} is separated to attend KTK_{T} and KSK_{S}, thus effectively mitigating this problem. We give more discussion in Sec. 5.3.

Spatial adaptive gating. We then consider fusing the separate attention results OintraO_{intra} and OinterO_{inter}. The main challenge is the spatial misalignment between ILQI_{LQ} and IR\mathbf{I_{R}}. For instance, the same objects may appear in different locations in ILQI_{LQ} and IR\mathbf{I_{R}}, or some objects in ILQI_{LQ} may not present in IR\mathbf{I_{R}} and vice versa. As a result, some pixels in QTQ_{T} may not find the corresponding reference in KSK_{S}, resulting in some pixels in OinterO_{inter} meaningless.

where ss is a user-defined scalar to control the degree to which the restored image attends the retrieved images, ⊗\otimes denotes the Hardamard product, and 1\mathbf{1} is an all one tensor with the same shape as M\mathcal{M}.

Distribution alignment. Using the OfuseO_{fuse} to replace the original intra-chain self-attention results OintraO_{intra} seems to be a promising way to integrate useful external knowledge from CS\mathcal{C}_{S}. However, it should be noticed that there is a domain gap between CT\mathcal{C}_{T} and CS\mathcal{C}_{S} due to the image quality and content differences, and thus a direct insertion of OintraO_{intra} into CT\mathcal{C}_{T} will result in a distribution shift of the original denoising chain in CT\mathcal{C}_{T}.

To this end, we propose the distribution alignment as a complementary to calibrate the distribution shift. Specifically, considering the latent in the diffusion chain is a Gaussian, we propose to use the Adaptive Instance Normalization (AdaIN) to align the mean and variance of OfuseO_{fuse} to the original statistics of OintraO_{intra}:

where AdaIN(u,v)\mathtt{AdaIN}(u,v) denotes replacing the mean and variance of uu with the corresponding part of vv. Finally, we replace the original self-attention result in CT\mathcal{C}_{T} with the well-aligned Ofuse′O^{\prime}_{fuse} to finish the cross image injection process.

Experiments

Datasets and metrics. In this work, we include experiments with two difficulty levels for performance evaluation. The first setup considers restoration with manually provided ideal reference images, which share a high content similarity with the LQ image, to evaluate the ability to utilize the reference knowledge. The datasets for this setting employ the widely used RefSR dataset including CUFED5 and WR-SR , in which the reference images are already provided. Since these datasets only contain HQ images, we thus use the second-order degradation model from Real-ESRGAN with ×4\times 4 down-sampling scale to generate the real-world degraded images. The second setup turns to more challenging practice where the reference images have to be retrieved using the retriever, and we use the RealPhoto60 which contains 60 real-world degraded images without ground truth for evaluation. And we use DIV2K as the high-quality image database for retrieval and employ the image encoder of VGG16 as the feature extractor. As for the evaluation metrics, we use both the fidelity metrics containing PSNR and SSIM, as well as the perceptual metrics including LPIPS , NIQE , FID , MUISQ , and CLIPIQA , to assess the performance of the different methods.

Implementation details. For a fair comparison, we use one reference image if not specified. Experiments with multiple reference images are given in Appendix A. Following the common practice of existing LRMs , the ILQI_{LQ} is up-sampled to the desired size using Bicubic before going through the LRMs. We use reflective padding to ensure the input size of CT\mathcal{C}_{T} and CS\mathcal{C}_{S} are the same. We use fixed random seeds for results reproducibility in all experiments. The hyperparameters of different baselines follow their original settings. We apply the proposed retrieval augmented restoration framework to two popular LRMs, namely SeeSR and SUPIR , and denoted the models augmented with our ReFIR as “SeeSR+ReFIR” and “SUPIR+ReFIR”, respectively.

2 Comparison to State-of-the-Arts

Restoration with ideal reference. We first compare on the RefSR dataset with real-world degradation. The compared methods includes state-of-the-art RefSR methods , GAN-based methods , and recent Diffusion-based methods . Tab. 1 gives the results. It can be seen that our method brings significant gains in all metrics on both fidelity (PSNR, SSIM) and perceptual quality (LPIPS, NIQE, FID) for the LRMs. Taking SUPIR as an example, our method brings a FID improvement of even 19.57 on the CUFED5 dataset. Moreover, similar performance gains can also be observed in SeeSR. For instance, equipping our ReFIR to SeeSR can lead to 0.38dB PSNR improvement, demonstrating the generalization of our ReFIR. It is noteworthy that the above superiority is obtained without any training or fine-tuning. Moreover, we also give visual comparisons in Fig. 5, and it can be seen that our method can generate details that are faithful to the original scene with the help of external knowledge from retrieved reference images.

Restoration in the wild. The above experiments on RefSR datasets focus on utilizing the already provided reference images from the dataset, which applies when the user has relevant HQ images. In this section, we turn to more challenging scenarios in which the reference image has to be obtained by retrieval. Since the ground truth of RealPhoto datasets is unavailable, we use non-reference image quality assessment metrics, i.e. NIQE, MUSIQ, and CLIPIQA for evaluation. As shown in Tab. 2, our approach continues to produce significant gains over its non-ReFIR counterparts. For instance, our SeeSR+ReFIR surpasses the original SeeSR by 0.2866 NIQE and 1.59 MUSIQ. Since the retrieved image can not serve as an ideal reference, the above favorable results demonstrate the robustness of our ReFIR in the face of real-world retrieved images. We also give quantitative results in Fig. 6. Even under severe real-world degradation, our method maintains good perceptual quality.

Complexity analysis. Tab. 3 gives the comparison of the computational complexity, including the number of parameters, GPU cost, and the inference latency. We also give the restoration performance for a more comprehensive comparison. As for the parameters, our ReFIR can facilitate both fidelity and realistic image restoration using the same #param as the original base LRMs. For the GPU memory, since our ReFIR uses two images as input, i.e., one LQ image, and one reference image, the GPU cost will become larger than the original one. For instance, it rises 1.38 times the increase of SUPIR+ReFIR than the original SUPIR model. Moreover, the inference time also increases due to more inputs as well as the additional interaction between two chains. In the future, we will delve deep into the effective utilization of retrieved images while maintaining efficiency.

3 Ablation Studies

Effectiveness of the reference retriever. In order to obtain content-relevant retrieved images, we present a simple but inference-efficient retriever R\mathcal{R} that uses the high-level semantic vectors from the pre-trained deep models for similarity matching in the high-quality image dataset D\mathcal{D}. Despite the simple design, we here demonstrate its effectiveness in Fig. 7. Since semantically consistent images usually contain similar textures, e.g., the texture in the first elephant image can help in the restoration of the LQ elephant image, and thus the proposed retriever can yield satisfactory retrieval results. Although texture-based retrieval may be a better choice for image restoration, it usually necessitates additional training of new retrieval models. For simplicity, we adopt semantic-based retrieval and leave the exploration of more advanced reference retrievers for future work.

Ablation on cross image injection. In the proposed cross image injection, we use separate attention (SA), spatial adaptive gating (SG), and distribution alignment (DA) for effective external knowledge injection. Here, we ablate to validate the effectiveness of different components. We use SUPIR+ReFIR as a representative on the CUFED5 dataset and use the scalar weighted sum when SG is removed. The results are shown in Tab. 4. One can see that using fixed scalar weights instead of spatial adaptive gating results in a 0.18 NIQE drop. This is because not all pixels of the reference image are useful, and thus fine-grain gated mask is needed. Moreover, removing the distribution alignment also impairs performance, e.g., 4.36 FID drop, since the distribution of raw fusion results OfuseO_{fuse} does not match CT\mathcal{C}_{T}, and directly inject OfuseO_{fuse} to the denoising chain of CT\mathcal{C}_{T} can cause sub-optimal results.

The motivation behind the proposed separate attention is to address the domain preference problem, i.e., the attention in CT\mathcal{C}_{T} will prefer to use latent from the same chain even though the latent from CS\mathcal{C}_{S} is more helpful for reconstruction. To verify the existence of the domain preference, we use the ground truth IHRI_{HR} as the input of CS\mathcal{C}_{S} and compute the normalized attention scores between QTQ_{T} and KTK_{T}, QTQ_{T} and KSK_{S}. It can be seen in Fig. 8 that even using the spatially strictly aligned IHQI_{HQ} as the reference, QTQ_{T} still has significantly high attention for the latent from the same chain, indicating that the domain preference problem interferes with the CT\mathcal{C}_{T}’s utilization of external knowledge in CS\mathcal{C}_{S}. By contrast, the proposed separate attention can effectively mitigate this problem by forcing the QTQ_{T} to separately attend KTK_{T} and KSK_{S}.

Other choices on injection position. In Sec. 3, we find the diffusion decoder is responsible for restoring textures. Based on this observation we propose to apply cross-image injection on the UNet decoder. Here, we ablate to analyze the impact of different cross-image injection positions. The results are shown in Tab. 5. It can be seen that performing cross-image injection only on the encoder will cause 19.57 FID drops. This is because the encoder focuses on the structure reconstruction, thus transferring the structure of CS\mathcal{C}_{S} will destroy the layout of the CT\mathcal{C}_{T}. Moreover, performing injection only in the decoder achieves the best results since it can transfer the high-quality textures from the CS\mathcal{C}_{S}. Due to the page limit, more ablation experiments can be seen in Appendix C.

4 Discussions

What is the impact of the control scale? The scale ss in Eq. 2 can control the extent to which the LRMs use external knowledge from the retrieved reference image for restoration. Here, we conduct an ablation study to explore the effect of ss. The results are shown in Fig. 9. It can be seen that when ss takes smaller values, the model mainly uses the internal knowledge embedded in its own parameters, which can make the model hallucinate when the degradation is severe. For example, the model produces incorrect textures when s=0s=0. As ss increases, the model starts to use external knowledge from the retrieved reference image, from which the model’s hallucination problem can be alleviated. We also provide quantitative ablation experiments on ss in Appendix C.

How much do the reference images affect performance? In the proposed framework, the retrieved images IR\mathbf{I_{R}} is crucial in alleviating hallucinations. Here, we try to answer the role of IR\mathbf{I_{R}} during restoration process, by manually controlling different types of retrieved images. As shown in Tab. 6, we find that using the exact ground truth IHQI_{HQ} as the IR\mathbf{I_{R}} can further improve the performance, which can be seen as an ideal up-bound. Interestingly, using ILQI_{LQ} itself as its own retrieved image instead brings a slight improvement compared with no retrieval, which we attribute to the regularization effect from the distribution alignment strategy. Finally, randomly selecting a high-quality reference image even resulted in a huge performance degradation, suggesting that the content correlation is more important than the image quality for a favorable retrieved reference image.

How does the proposed ReFIR work? Extensive experiments have shown the state-of-the-art performance of our ReFIR. However, it seems not straightforward to understand how the retrieved reference images influence the image restoration process of the original LRMs. Here, we give an intuitive explanation. As shown in Fig. 10, for the latent at the tt-th time step on the latent manifold, there are two forces in different directions pulling it to produce the latent at the next t−1t-1-th time step. One force is from the internal knowledge of frozen weights in LRMs, and the other is the external knowledge from the retrieved reference image through the proposed cross image injection mechanism. These two forces ultimately determine the latent of the next time step. Therefore, a restored image from our ReFIR can utilize both the internal knowledge in the original LRMs as well as the external knowledge in the retrieved image, thus alleviating the hallucination of the LRMs.

Conclusion

This paper presents ReFIR, a training-free and generic framework that can alleviate the hallucination of LRMs to facilitate high-fidelity and photo-realistic restoration results through retrieval augmentation. We introduce the nearest neighbor lookup as a simple retriever to obtain relevant high-quality images and further propose the cross-image injection which employs separate attention to transfer knowledge while avoiding the domain preference problem, the spatial adaptive gating to address the spatial misalignment, and the distribution alignment to mitigate the domain gap during injection. Through expanding the knowledge boundary using the additional external knowledge from retrieved images, our ReFIR exhibits significant improvements on both fidelity and perceptual quality, as demonstrated through extensive qualitative and quantitative evaluations. Moreover, with its training-free and generic nature, our ReFIR can be easily applied to multiple LRMs.

Acknowledgements

This work is supported in part by the National Natural Science Foundation of China, under Grant (62302309,62171248), Shenzhen Science and Technology Program (JCYJ20220818101014030, JCYJ20220818101012025), and the PCNL KEY project (PCL2023AS6-1).

References

Appendix

Appendix A Adaptive Multi-reference Injection

In the main paper, we mainly focus on the case of one single retrieved image. However, in practice, there may be multiple available reference images at hand, and using multiple reference images for resemblance could intuitively gain better performance. To this end, we extend the original cross-image injection to allow to incorporation of multiple reference images for reconstruction. Our key idea is to modify the scale factor ss in Eq. 2 from a scalar into a vector: s={s1,s2,⋯ ,sk}\mathbf{s}=\{s_{1},s_{2},\cdots,s_{k}\}, where ∑sn=1\sum s_{n}=1. Each sns_{n} can be obtained by computing the cosine similarity between ILQI_{LQ} and the corresponding nn-th retrieved image in IR\mathbf{I_{R}} followed by Softmax normalization. Then we can modify the original single-reference cross-image injection of Eq. 2 to the following multi-reference version:

where Mn\mathcal{M}_{n} denotes the gated mask of the nn-th reference image.

Experiments with multiple reference images.

For experiments with multiple reference images, we use SUPIR+ReFIR as a representative. Since the CUFED5 dataset contains multiple reference images, we directly use the provided images as the retrieved reference for reproducibility. Tab. 8 gives the results. It can be seen that using multiple reference images produces better results than one single reference image, e.g. the 2.08 improvement in FID. However, it is worth noting that the marginal gain from adding reference images is diminishing, accompanied by a notable increase in computational cost. Therefore, in practice, we use one single reference image to balance the model performance and inference efficiency.

Appendix B More Discussions

Our ReFIR uses retrieved images as the reference for high-fidelity restoration. Despite both RefSR methods and ours appears the reference image, we would like to clarify the difference between our ReFIR and previous RefSR methods. Firstly, current RefSR models are typically small-scale (#param <50M) and use simple Bicubic degradation, while our ReFIR focuses on the recent diffusion-based large-scale restoration model (#param >1B) for more challenging real-world SR. Secondly, most RefSR methods can only use one reference image and even fail to work in the absence of reference images, by contrast, our ReFIR can flexibly use 0∼k0\sim k images. Thirdly, different from RefSR models that require training, our method can be applied in various LRMs in a training-free manner.

Performance under extreme conditions.

Computational overhead from retrieval and attention modification. Since we employ additional Ref images as input and modify the attention layers, we adiscuss the impact of these trchnuques on the inferenve efficiency. First, in order to reduce the computational overhead of the retrieval process, we pre-calculated the feature vectors of all images in the retrieval database before inference. Furthermore, the cosine similarity between the LR image vectors and all retrieval vectors is computed in parallel. These strategy results in an almost negligible (less than 3% inference time) cost of computational overhead. Second, the modification of self-attention layers only happens in the last 20 timestep in the decoder layers, i.e., only 12% attention layers are modified while the left is kept intact. These analysis is also supported by practice, in which we find these two process only take up <5% inference time, with most computational cost coming from the original LRM. Future LRM acceleration (e.g. pruning, quantization, one-step diffusion) will benefit our ReFIR, and we will explore more efficient implementation in the future.

Why use the self-attention as the external knowledge?

In the proposed cross-image injection, we use the features of the self-attention layer of the CS\mathcal{C}_{S}’s decoder as external knowledge to guide CT\mathcal{C}_{T} to produce textures faithful to the original scene. Here, we give the reason behind this. Firstly, previous image-to-image efforts , e.g., image editing, has demonstrated through extensive experiments that the self-attention layer of the diffusion model contains important spatial correlations in images, which inspired us to follow this clue to utilize this prior. Secondly, leveraging the attention mechanism allows CT\mathcal{C}_{T} to query features in CS\mathcal{C}_{S} without any training, whereas using features from other parts of CS\mathcal{C}_{S} may require introducing additional training.

What about the quality of cross-image attention?

In the proposed cross-image injection, the inter-chain attention is used to perform attention between QTQ_{T} and KSK_{S}. Considering the domain gap between CT\mathcal{C}_{T} and CS\mathcal{C}_{S} due to the input quality difference, one may ask whether the results of the inter-chain attention are meaningful. Here, We visualize the attention map to validate the effectiveness of the inter-chain attention (see Fig. 13). It can be seen that for a given query pixel query in CT\mathcal{C}_{T}, the inter-chain attention can effectively overcome the spatial misalignment, and find relevant pixel features in CS\mathcal{C}_{S} for reference.

Appendix C More Ablation Results

We also provide quantitative ablation results on the control scale ss in Fig. 11. It can be seen that when ss is too small, the LRM will mainly use the knowledge contained within its parameters to restore high-quality images, which can lead to performance degradation due to the hallucination problem. On the other hand, when ss is too large, the LRM will overuse the content in the retrieved reference image, thus producing patterns that are not present in the original LQ image. In practice, we adopt a moderate s=0.5s=0.5 to trade off the hallucination and the overuse of the reference image.

Other choices for cross image injection.

The proposed cross image injection mitigates the domain preference problem by using separate attention to promote latent in CT\mathcal{C}_{T} to attend CS\mathcal{C}_{S}. Here, we conduct ablation to study the impact of different design choices of cross image injection. As shown in Tab. 9, directly replacing the original self-attention results from OintraO_{intra} in CT\mathcal{C}_{T} with corresponding latent in CS\mathcal{C}_{S} causes severe performance degradation, due to the significant loss of original knowledge in CT\mathcal{C}_{T}. In addition, using QTQ_{T} to query the concatenation results of KTK_{T} and KSK_{S} also causes a performance drop, which further confirms that the domain preference problem, i.e., QTQ_{T} prefers to use latent from the same chain CT\mathcal{C}_{T}, even though CS\mathcal{C}_{S} is more helpful for reconstruction.

Appendix D Statistical Significance on Performance

In Tab. 1 of the main paper, we give the performance gains of incorporating the proposed ReFIR into the existing LRMs. Considering the randomness of the generative models, we give the performance fluctuations of ReFIR under multiple trials with exactly the same experimental setting and random seed. The results are given in Tab. 10. It can be seen that the randomness of the diffusion-based generative model is very small when using a fixed seed, reducing the disturbance from noise errors for evaluation. In addition, we further use hypothesis testing to verify the significance of performance gains, and the test results reject the original hypothesis H0 at 95% confidence level on all metrics and datasets, indicating that the performance gains from the proposed method are statistically significant.

Appendix E Extension to Specific Restoration Scenarios

An important application of our method is in scenarios with high fidelity demands, such as scene text images with a specific stylistic structure, or face images with identity preservation, and here we preliminarily explore the application of the proposed ReFIR to real-world face image restoration. The results are given in Fig. 15. It can be seen that by using a high quality image of a specific person’s identity as a reference, the resulting restoration results can better preserve the person’s attributes. However, it should be noted that this experiment is just a preliminary attempt, and we will leave the further improvement of our ReFIR for specific downstream restoration tasks for future work.

Appendix F Limitation and Future Works

Although the proposed ReFIR can effectively mitigate the hallucination of LRMs by introducing external knowledge from retrieved reference images, the proposed framework can be further improved in the following aspects. First, since the computational complexity of the current LRMs is costly, the computational complexity will be further increased when using the proposed method, which may hinder the use of resource-constrained mobile devices. With the advent of accelerated diffusion-based image restoration methods in the future, we believe that the proposed method can further improve its efficiency. In addition, this paper proposes a simple retriever based on semantic vector matching for presentation. With the development of image retrieval techniques , designing specialized retrievers, e.g., using textures as key matching cues, will further improve the performance. Finally, for some slightly degraded images, which can already be handled well by only using the internal knowledge of the LRMs, designing hyper-networks to adaptively decide whether to use retrieval augmentation or not is also promising. We leave the above considerations for future work.

Appendix G Broader Impact

The development of our ReFIR offers significant positive societal impacts, including advancements in medical imaging, historical preservation, and media restoration by enhancing the fidelity and realism of image restoration. However, it also poses potential negative societal impacts, such as the misuse of improved restoration capabilities for generating disinformation, deepfakes, and surveillance, raising ethical concerns about privacy, security, and fairness. To mitigate these risks, implementing safeguards like gated releases of models, monitoring mechanisms, and transparency in model training and deployment is crucial. Continuous ethical evaluation and adherence to strict guidelines are essential to prevent potential harms.

Appendix H Additional Visual Results

In this section, we provide more visual results, which are organized as follows:

In Fig. 12, we give more samples of the PCA visualization on the top three principal components of the self-attention layer latent.

In Fig. 13, we give a visualization of the attention map from the cross-image injection, to help better understand the feasibility of cross-image attention.

Fig. 14 gives more quantitative comparison results against the state-of-the-art method on real-world degradation without ground truth.

Fig. 15 gives the visualization results of the extension experiments of applying the proposed ReFIR to blind face image restoration.