RFormer: Transformer-based Generative Adversarial Network for Real Fundus Image Restoration on A New Clinical Benchmark

Zhuo Deng, Yuanhao Cai, Lu Chen, Zheng Gong, Qiqi Bao, Xue Yao, Dong Fang, Shaochong Zhang, Lan Ma

introduction

Due to the safety and cost-effectiveness in acquiring, fundus images are widely used by ophthalmologists for early eye disease detection and diagnosis, including glaucoma , diabetic retinopathy , cataract , and age-related macular degeneration . However, different equipments and ophthalmologists pose large variations to the quality of fundus images. A screening study of 5,575 patients found that about 12%12\% of fundus images are of inadequate quality to be readable by ophthalmologists . We analyze the factors causing the degradation in real fundus image capturing. Firstly, patients, especially infant patients, do not cooperate with the capturing process of fundus images. Specifically, most patients are reluctant to undergo pupil dilation, which causes poorly lit and blurred fundus images. Besides, infant patients usually can not resist the eye-closing reflex caused by a bright light during flash photography. Secondly, in practice, spatial pixel misalignment, color, and brightness mismatch are inevitable due to the changes in light conditions and misoperations of inexpert ophthalmologists. Thirdly, high-quality (HQ) fundus images can be collected in hospitals of developed areas using high-precision fundus cameras. However, these equipments are expensive and unaffordable for hospitals in some remote areas of under-developed or developing countries. As a result, low-precision and portable fundus cameras are used to capture low-quality (LQ) fundus images. These LQ fundus images easily mislead the clinical diagnosis and lead to unsatisfactory results of downstream tasks like blood vessels segmentation. Various biomarkers of the retina (e.g.e.g., hemorrhage, microaneurysm, exudate, optic nerve and optic cup) are essential in different diseases. Therefore, it is necessary to ensure the prominence and visibility of each marker for precise clinical diagnosis. Thus, when LQ fundus images are captured in clinical diagnosis, ophthalmologists often repeat dozens of shots until HQ fundus images are obtained. Nonetheless, this repeated capturing process harms patients, degrades hospital efficiency, prevents reliable diagnosis of ophthalmologists, and impacts automated image analysis systems.

We observe the clinical fundus images and find that the main degradation types of LQ images include out-of-focus blur, motion blur, artifact, over-exposure, and over-darkness. Compared with other types of degradation, blur, especially out-of-focus blur, poses the most severe threat to image analysis and clinical diagnosis. An example is shown in Fig. 2, where the uneven illumination, haze, and out-of-focus blur degradation are presented in (a), (b), and (c), respectively. It can be observed that the performance of blood vessel segmentation only collapses on out-of-focus blurred fundus images.

Traditional fundus image restoration methods are mainly based on handcrafted priors. However, these model-based methods achieve unsatisfactory performance and generality due to the poor representing capacity. Recently, deep Convolutional Neural Networks (CNNs) have been widely used in natural image restoration and enhancement , e.g., super resolution , deraining , deblurring , enlighten , etc. Inspired by the success of natural image restoration, CNNs have also been applied to fundus image restoration . Although impressive results have been achieved, CNN-based methods show limitations in capturing long-range dependencies. In recent years, the natural language processing (NLP) model, Transformer has been introduced into computer vision and outperformed CNN-based methods in many tasks. The Multi-head Self-Attention (MSA) in Transformer excels at modeling non-local similarity and long-range dependencies. This advantage of Transformer may provide a possibility to address the limitations of CNN-based methods.

Existing deep learning methods rely on a large amount of LQ and HQ fundus image pairs. Unfortunately, real clinical benchmark has not been explored for fundus image reconstruction. There remains a data-hungry problem. As shown in Fig. 1 (b), to get more image pairs, artificially designed degradation models such as Gaussian filter are used to synthesize degraded fundus images from their high-quality counterparts. However, as depicted in Fig. 1 (c), artificial degradation is fundamentally different from clinical degradation. As shown in Fig. 1 (d), the two CNN-based methods trained with synthesized data fail in real fundus image restoration.

In this paper, we investigate the real fundus image restoration problem, which has not been studied in the literature. Our work is the first attempt. To begin with, we establish a clinical benchmark, Real Fundus (RF), including 120 LQ and HQ real clinical fundus image pairs to alleviate the data-hungry issue. Based on this dataset, we propose a novel method, namely Transformer-based Generative Adversarial Network (RFormer), for real fundus image restoration. Specifically, the generator and discriminator are built up by the basic unit, Window-based Self-Attention Blocks (WSABs). The self-attention mechanism equipped with each basic block excels at capturing the non-local self-similarity and long-range dependencies, which are the main limitations of existing CNN-based methods. In particular, the generator adopts a U-shape structure to aggregate multi-resolution contextual information. Unlike previous CNN-based Generative Adversarial Networks (GANs), we adopt a Transformer-based discriminator to extract non-local image prior information and thus improve the ability of discriminator to distinguish restored fundus images from the ground-truth HQ fundus images. Our Transformer-based adversarial training scheme encourages the generator to create more plausible-looking natural and visually-pleasant images with more detailed contents and structural textures.

Our contributions can be summarized as follows:

We establish a new clinical benchmark, RF, to evaluate algorithms in real fundus image restoration. To the best of our knowledge, this is the first real fundus image dataset.

We propose a novel Transformer-based method, RFormer, for real fundus image restoration. To the best of our knowledge, it is the first attempt to explore the potential of Transformer for this task in the literature.

Comprehensive quantitative and qualitative experiments demonstrate that our RFormer significantly outperforms SOTA algorithms. Extensive experiments of downstream tasks further validate the effectiveness of our method.

related work

Traditional fundus image restoration and enhancement methods are mainly based on hand-crafted priors. For example, Setiawan et al.\textit{et al}. apply contrast limited adaptive histogram equalization (CLAHE) to fundus image enhancement. Some methods decompose the reflection and illumination, achieving image enhancement and correction by estimating the solution in an alternate minimization scheme. However, these model-based methods achieve unsatisfactory performance and generality due to the poor representing capacity. With the development of deep learning, fundus image restoration has witnessed a significant progress. CNNs apply a powerful learning model to restore LQ fundus images. For instance, Zhao et al.\textit{et al}. propose an end-to-end deep CNN to remove the lesions on the fundus images of cataract patients. However, the cataract lesions are not caused by clinical fundus imaging. Sourya et al.\textit{et al}. , Shen et al.\textit{et al}. , and Raj et al.\textit{et al}. customize different synthetic degradation models to better simulate the degradation types in actual clinical practice. However, real fundus image degradation is more sophisticated than synthesized degradation. It is hard to simulate real degradation by artificial degradation models completely. Thus, models trained on synthesized data easily fail in real fundus image restoration. In addition, the CNN-based methods show limitations in capturing non-local self-similarity and long-rang dependencies, which are critical for fundus image reconstruction.

2 Generative Adversarial Network

Generative Adversarial Network (GAN) is firstly introduced in and has been proven successful in image synthesis , and translation . Subsequently, GAN is applied to image restoration and enhancement, e.g., super resolution , deraining , deblurring , enlighten , dehazing , image inpainting , style transfer , image editing , medical image enhancement , and mobile photo enhancement . Although GAN is widely applied in low-level vision tasks, few works are dedicated to improving the underlying framework of GAN, such as replacing the traditional CNN framework with Transformer. Jiang et al.\textit{et al}. propose the first Transformer-based GAN, TransGAN, for image generation. Nonetheless, to the best of our knowledge, the Transformer-based GAN has not been involved in fundus image restoration.

3 Vision Transformer

Transformer is proposed by for machine translation. Recently, Transformer has achieved great success in high-level vision, such as image classification , semantic segmentation , human pose estimation , object detection , etc.\textit{etc}. Due to the advantage of capturing long-range dependencies and excellent performance in many high-level vision tasks, Transformer has also been introduced into low-level vision . SwinIR uses Swin Transformer blocks to build up a residual network and achieve SOTA results in natural image restoration. Chen et al.\textit{et al}. propose a large model IPT pre-trained on large-scale datasets with a multitask learning scheme. MST presents a spectral-wise Transformer for HSI reconstruction. Although Transformer has achieved impressive results in many tasks, its potential in fundus image restoration remains under-explored.

Methodology

The architecture of RFormer is shown in Fig. 3, where (a) and (b) depict the generator and discriminator. Fig. 3 (c) illustrates the proposed Window-based Self-Attention Blocks (WSABs), which consists of a Feed-Forward Network (FFN) (detailed in Fig. 3 (d)), a Window-based Multi-head Self-Attention (W-MSA), and two layer normalization.

2 Window-based Self-Attention Block

The emergence of Transformer provides an alternative to address the limitations of CNN-based methods in modeling non-local self-similarity and long-range dependencies. However, as analyzed in Swin Transformer , the computational cost of the standard global Transformer is quadratic to the spatial size of the input feature (HWHW). This burden is nontrivial and sometimes unaffordable. To tackle this problem, we adopt the Window-based Multi-head Self-Attention (W-MSA) as the self-attention mechanism and integrate it with the basic Transformer unit. The computational complexity of W-MSA is linear to the spatial size, which is much cheaper than that of standard global MSA. Inspired by Swin Transformer , we add window shift operations(WSO) in our proposed Window-based Multi-head Self-Attention Block (WSAB) to introduce cross-window connections. The components of our proposed WSAB are shown in Fig. 3 (c). WSAB consists of a W-MSA, an FFN, and two layer normalization. The details of FFN are shown in Fig. 3 (d). Then WSAB can be formulated as

where Fin\mathbf{F}_{in} represents the input feature maps of a WSAB. LN(⋅)\text{LN}(\cdot) represents the layer normalization. F′\mathbf{F}^{\prime} and Fout\mathbf{F}_{out} denote the output feature of W-MSA and FFN respectively.

2.2 Feed-Forward Network

As depicted in Fig. 3 (d), the Feed-Forward Network (FFN) consists of a 1×11\times 1 convconv layer with a GELU activation, a depth-wise 3×33\times 3 convconv layer with a GELU activation, and another 1×11\times 1 convconv layer.

3 Loss Functions

During the training procedure, we exploit the weighted sum of four loss functions as the overall training objective. They are described and analyzed in the following part.

The first loss function is the Charbonnier loss between the restored and ground-truth HQ images:

where IR\mathbf{I}_{R} denotes the restored fundus image, IHQ\mathbf{I}_{HQ} represents the ground-truth HQ fundus image, and ε\varepsilon denotes a constant which is empirically set to 10−310^{-3} for all the experiments.

3.2 Fundus Quality Perception Loss

Unlike natural images, fundus images have specific acquisition process and anatomical structures. This indicates fundus images have highly similar styles. Therefore, we exploit high-level feature constraints to improve the perceptual quality and encourage the network to capture the fundus anatomical structures and styles. To this end, we propose Fundus Quality Perception Loss (FQPLoss). More specifically, we adopt VGG-19 as the perception network and pre-train it on the fundus image quality evaluation dataset, Eye-Q with the fundus image quality classification task. The Eye-Q dataset has 28,792 fundus images with three-level quality grading. The perception network trained on the Eye-Q dataset is capable of extracting the difference of high-level features between different qualities of fundus images. Subsequently, our FQPLoss can be formulated as

where HH and WW denote the height and width of the fundus image. ϕ(⋅)\phi(\cdot) denotes the feature extraction function of the pre-trained perception network. Our FQPLoss is customized to assess a solution with respect to perceptually relevant characteristics. By minimizing the FQPLoss Lfqp\mathcal{L}_{fqp}, the model is encouraged to capture more high-level discriminative features and generate more visually-pleasant results.

3.3 Adversarial Loss

where D(⋅)D(\cdot) denotes the mapping function of our proposed Transformer-based discriminator. LadvG\mathcal{L}_{adv}^{G} trains the generator to fool the discriminator by generating more realistic restored fundus images. In contrast, LadvD\mathcal{L}_{adv}^{D} encourages the discriminator to distinguish the restored images from real images.

3.4 Edge Loss

To enhance the high-frequency edge details, we exploit the edge loss function that focuses on the gradient information of images and enhances edge textures. To be specific, the edge loss function is formulated as

where Δ(⋅)\Delta(\cdot) represents the Laplacian operator.

3.5 The Overall Loss Function

Finally, the overall training objective is the weighted sum of the above four loss functions:

where λ1,λ2,λ3\lambda_{1},\lambda_{2},\lambda_{3} are three hyper-parameters controlling the importance balance of different loss functions. Our proposed RFormer is end-to-end trained by minimizing L\mathcal{L}. The weights of the perception network are fixed. Each mini-batch training procedure can be divided into two steps: (i) Fix the discriminator and train the generator. (ii) Fix the generator and train the discriminator. This adversarial training scheme encourages the reconstructed fundus images to be more photo-realistic and closer to the real clinical HQ fundus image manifold.

Real Fundus

This section introduces our clinical benchmark, Real Fundus (RF). It consists of 120 LQ and HQ clinical fundus image pairs with the spatial size of 2560×25602560\times 2560. The training and testing subsets are split in proportional to 3:1. Since blur significantly impacts clinical diagnosis and automated image analyzing systems, it is set to the primary degradation type of LQ fundus images. Besides, there are other degradation types such as artifacts and uneven illumination which are inevitably introduced in the fundus image capturing process.

The collection process of our RF obtains the exemption determination from Shenzhen Eye Hospital and contains three steps: capturing, selecting, and calibrating fundus images.

Instead of exploiting artificial degradation models (e.g., Gaussian Filter.) to synthesize LQ fundus images as shown in Fig. 1 (b), we directly use the degraded fundus images from the fail cases in practical capturing. As depicted in Fig. 1 (a), the fundus images are captured by ophthalmologists using a ZEISS VISUCAM200 fundus camera, which is a mainstream product of fundus camera. The price of ZEISS VISUCAM200 fundus camera is about 350,000 RMB. We select clinical fundus images from patients of different ages and different fundus states (e.g., leopard fundus, hemorrhage, microaneurysms, and drusen ) to expand the scope of our RF.

1.2 Selecting

When LQ fundus images are captured in practice, the operator will repeat capturing until HQ fundus images are obtained. Subsequently, we manually select LQ and HQ fundus image pairs of the same eye. To ensure the diversity of RF and avoid similar data, only one image pair is selected with one eye. Note that only HQ clear fundus images captured by experienced ophthalmologists can be used as the ground truths of degraded LQ images. Based on these strict criteria, we finally select 120 LQ and HQ fundus image pairs from the eye hospital database containing more than 30,00 eye instances. Each instance contains multiple fundus images.

1.3 Calibrating

After selecting fundus image pairs, we observe two issues in raw unprocessed fundus data. Firstly, the LQ and HQ fundus images are spatially misaligned (as illustrated in Fig. 1(a)). Secondly, there is a large black area around the eyeball. This black area is uninformative and may easily degrade the performance of the restoration model during the training procedure. Thus, to improve the quality of our RF, we calibrate the collected dataset using the software, Photoshop. Specifically, we first spatially align the image pairs and then cut off the black area around the eyeball.

2 Comparisons with Synthetic Dataset

We compare the LQ images from our RF and synthetic dataset in Fig. 1 (c). As can be seen from the zoom-in patches that the artificially synthesized degradation is fundamentally different from the real clinical degradation.

2.2 Domain Discrepancy

To validate the huge domain discrepancy between the synthetic and real clinical datasets, we adopt two CNN-based fundus image restoration methods, I-SECRET and Cofe-Net , to conduct ablation study. We train them with the synthetic data and then test them on our RF. As shown in Fig. 1 (d), the two models fail to reconstruct the real clinical LQ fundus images. They either yield over-smooth results sacrificing detailed contents, or introduce visually unpleasant artifacts. Since the synthetic data can not be applied to real fundus image restoration, it still remains a severe data-hungry issue. To meet with this research requirement, we establish a large scale clinical dataset, RF. To the best of our knowledge, this is the first work contributing a real clinical fundus image restoration benchmark.

Experiments

During the training procedure, fundus images are first cropped into the patches with the size of 128×\times128. Then the patches are fed into our proposed RFormer. The Adam optimizer (β1\beta_{1}=0.9, β2\beta_{2}=0.999) is adopted. The initial learning rate is set to 1×10−41\times 10^{-4}. The cosine annealing strategy is employed to steadily decrease the learning rate from the initial value to 1×10−61\times 10^{-6} during the training procedure. Our RFormer is implemented by PyTorch. It takes about 12h using an NVIDIA RTX 3090 GPU to train for 100 epochs. The mini-batch size is set to 4. Random flipping and rotation are used for data augmentation. In the testing phase, the input is the whole image with the size of 2560×\times2560 for fair comparison with other methods. We adopt peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) as the metrics to evaluate the fundus image reconstruction performance.

2 Comparisons with State-of-the-Art Methods

We provide quantitative comparisons between our RFormer with seven SOTA methods including two model-based methods (GLCAE and Bicubic+RL ), four CNN-based methods (RealSR , ESRGAN , I-SECRET , and Cofe-Net ), and one Transformer-based method (MST ). The quantitative comparisons on our RF are shown in Table 1, the proposed RFormer outperforms other competitors in terms of PSNR and SSIM. Specifically, RFormer achieves 0.33 and 1.59 dB improvement in PSNR when compared to RealSR and ESRGAN .

To verify the robustness of our RFormer, we conduct 5-fold and 10-fold cross-validation. The results are shown in Table 1. As can be seen that our Rformer still achieves robust results, e.g., 28.15 dB in 5-fold cross-validation and 28.38 dB in 10-fold cross-validation. The small gap in performance suggests that the overfitting is moderate while the effectiveness of RFormer is reliable and promising.

Fig. 5 depicts the qualitative comparisons on RF. It can be observed that Bicubic+RL , GLCAE , I-SECRET , and Cofe-Net yield over-smooth results and fail to restore the LQ blurry fundus images. Although ESRGAN and RealSR can reconstruct more high-frequency edge details, they do not maintain the authenticity and anatomical structure of the original fundus. Some undesired artifacts are introduced to the restored images, which may severely mislead the clinical diagnosis. In contrast, our RFormer is capable of restoring more fine-grained contents and structural details without introducing artifacts. Thus, the fundus anatomical structure can be well preserved.

3 Ablation Study

We adopt RFormer as the restoration model to conduct an ablation study to validate the effect of our FQPLoss. As listed in Table 2, when the FQPLoss is applied, the PSNR and SSIM are increased by 1.07 dB and 0.129, respectively. In addition, we provide visual comparisons in Fig. 7. As depicted in Fig. 7 (b), the model yields an over-smooth fundus image and fails in reconstructing the fine-grained vessel details without FQPLoss. As shown in Fig. 7 (c), when using our FQPLoss, the model restores more detailed anatomical structure contents and high-frequency textures.

3.2 Discriminator

We conduct ablation study to compare our proposed Transformer-based discriminator with Traditional CNN-based discriminators. Please note that the Transformer-based generator remains unchanged. The results are reported in Table 3. Compared with the discriminators in CNN-based PatchGAN , PixelGAN , etc,\textit{etc}, our Transformer-based discriminator yields the best performance. We provide qualitative comparisons in Fig. 6. It can be observed that our Transformer-based discriminator significantly outperforms CNN-based discriminators in terms of recovering detailed contents and preserving the anatomical structure consistency.

3.3 Patch Size

We experimentally analyze the effect of the patch size set in the Transformer-based discriminator. The results are shown in Table 4. Our RFormer achieves the best restoration result with the patch size of 40×\times40.

3.4 Window Size

We change the window size of W-MSA and conduct experiments to study its effect. The results are reported in Table 6. It can be observed that our RFormer yields the best result when the window size is set to 8×\times8.

3.5 Window Shift Operations

We conduct ablation study to analyze the effect of the window shift operations. The results are reported in Table 2. The results indicate that the window shift operations can build cross-window connections and improve the performance of RFormer.

4 Clinical Image Analysis and Applications

The ultimate goal of restoring and enhancing fundus images is to serve the real clinical tasks better and improve the accuracy of clinical diagnosis. To validate the effectiveness of our proposed RFormer, we use it as a pre-processing technique for downstream clinical image analysis tasks, including vessel segmentation and optic disc/cup detection. LadderNet and M-Net are employed as the segmentation baselines.

Since the restored fundus images should meet the requirements of ophthalmologists, we adopt 30 LQ fundus images for user study. We use different image restoration methods to enhance these LQ images. Subsequently, we display these results in random order and ask the experienced ophthalmologists to score the quality of the restored images based on their extensive clinical experience. The score ranges from 0 to 100, larger values are better. The suppression of artifacts and preservation of lesions are taken into account. Finally, we collect responses from five ophthalmologists. The score results for each method are shown in Table.5. Our RFormer receives the highest score for best restored results.

4.2 Vessel Segmentation

We test LadderNet pre-trained on DRIVE dataset for vessel segmentation on our collected RF. Please note that the vessel segmentation maps of real HQ fundus images serve as the references for comparison due to the lack of segmentation labels on our RF. The vessel segmentation results are shown in the third row of Fig.8. As can be seen that LadderNet fails in segmenting the blood vessels of clinical LQ fundus images. In contrast, LadderNet extracts obvious vessel structure of the fundus images restored by our RFormer, which is closest to the segmentation results of real HQ fundus images. These results clearly suggest the effectiveness of our proposed method.

4.3 Optic Disc/Cup Detection

We also evaluate the effect of our RFormer for the downstream disc/cup detection task. We test M-Net pre-trained on ORIGA dataset for optic disc/cup detection on our collected RF. Similar to the vessel segmentation task, the optic disc/cup detection results of real HQ fundus images function as the references due to the lack of segmentation labels. The qualitative comparisons of different fundus image restoration methods are depicted in Fig. 8. The fourth line is the optic disc/cup detection map and the fifth line depicts the zoom-in patches of the fourth line. It can be observed that M-Net fails to detect the disc/cup on clinical LQ fundus images. On the contrary, M-Net detects the optic cup and disc more accurately on the fundus images reconstructed by our RFormer. This evidence verifies that our method benefits the optic disc/cup detection task.

5 Fail Cases

Although RFormer achieves good performance, it may not work in some scenes. Fig. 9 shows some fail cases of RFormer on our RF dataset. In the (a) and (e) column, from top to bottom are LQ fundus image, restored fundus image, and HQ fundus image. (b), (c), and (d) are three zoom-in patches of (a). (f), (g), and (h) are three zoom-in patches of (e). It can be clearly observed from 9 (c), (f), and (g) that our RFormer fails to remove the bright spots. As can be seen from 9 (b), (d), and (h) that our RFormer fails in enhancing the low-lights regions. It is difficult for RFormer to learn the feature in areas with insufficient contrast and brightness. We will continue to improve our work according to these fail cases.

Conclusion

In this paper, we establish the first real clinical fundus image restoration benchmark, Real Fundus, which contains LQ and HQ fundus image pairs to alleviate the data-hungry issue. Our dataset can help better evaluate restoration algorithms in clinical scenes. Based on this dataset, we propose a novel Transformer-based method, RFormer, for clinical fundus image restoration. To the best of our knowledge, it is the first attempt to explore the potential of Transformer in this task. Comprehensive qualitative and quantitative results demonstrate that our RFormer significantly outperforms a series of SOTA methods. Extensive experiments verify that the proposed RFormer serving as a data pre-processing technique can boost the performance of different downstream tasks, such as vessel segmentation and optic disc/cup detection. We hope this work can serve as a baseline for real clinical fundus image restoration and benefit the community of medical imaging.

References