HQG-Net: Unpaired Medical Image Enhancement with High-Quality Guidance

Chunming He, Kai Li, Guoxia Xu, Jiangpeng Yan, Longxiang Tang, Yulun Zhang, Xiu Li, Yaowei Wang

I Introduction

D ue to the variability of light transmission and clinical imaging conditions, medical images often exhibit uneven illumination or blurry texture details (see Fig. 1 B and C). These low-quality (LQ) images can significantly impede automated disease screening, examination, and diagnosis. Medical image enhancement aims to

transform an LQ image to a high-quality (HQ) one that fulfills the modality-specific HQ definitions (in Fig. 1), such as enhanced illumination quality and texture details . Early approaches optimize energy-based objective functions based on certain imaging priors, such as Retinex equation and region line prior . However, these methods are limited in their ability to generalize to different types of degradation conditions, as one imaging prior may only apply to certain types of degradation conditions, restricting their applicability. Recently, deep learning-based methods provide new avenues to this problem and use neural networks to approximate various degradation models. A naive approach involves training a neural network, such as Pix2Pix , to translate an LQ image from the LQ domain to the HQ domain at the pixel level. However, this approach requires paired LQ-HQ images with pixel-to-pixel correspondence for training, which can be challenging to obtain for medical applications .

To address the enhancement task more practically, several methods have investigated the unpaired medical image enhancement (UMIE) problem, which requires only unpaired HQ and LQ images as training data. The most intuitive idea is to retain the main structure of Pix2Pix but use specific loss functions for inter-domain structure preservation, e.g., structure loss (Fig. 2 (a)). For instance, EnlightenGAN employs a structure-preserving loss with a global-local discriminator. However, the generator of the Pix2Pix-based method is mainly trained with an unsupervised loss between real LQ and generated HQ images, which neglects valuable real HQ information during the training phase, leading to the generation of artifacts in the enhanced medical images. Another approach is to exploit CycleGAN , which involves a two-sided translation between LQ and HQ domains for the intra-domain distribution similarity with a cycle-consistent constraint (Fig. 2 (b)). By employing both HQ and LQ images for network training, CycleGAN-based techniques outperform Pix2Pix-based methods in exploiting valuable information from unpaired real HQ images. However, since the real HQ images are not directly involved in the enhancement of LQ images, there exists a domain shift between the enhanced LQ (generated HQ) and the real HQ domains, leading to potential influences on structure fidelity and texture distortion. For structure preservation, StillGAN , a CycleGAN-based approach, proposes an SSIM-based structure loss. However, this loss function is applied to the entire image space and can hardly distinguish the complete structure information of LQ medical images with complex degradation.

Considering the intra-modality homogeneity of medical images (Fig. 1), it is desirable to enhance LQ images with the guidance of unpaired but homogeneous HQ images. Therefore, we propose HQG-Net, a GAN-based UMIE network that explicitly utilizes unpaired HQ images as guidance for LQ image enhancement (see Fig. 2 (c)). To ensure the stability of HQG-Net, we introduce the feature vector vz\mathbf{v}_{z} of an unpaired HQ sample z\mathbf{z} as guidance. This feature vector is a condensed extraction of HQ information and can moderate the slight degradation from the HQ sample. To preserve structure information from the LQ input, we propose a Variational Information Normalization (VIN) module (see Fig. 3 (c)) that guides the enhancement in a variational manner, where the concise and general variational information ensures the structure preservation of the original image during the domain translation. By introducing the variational guidance information, HQG-Net models the UMIE task under a joint distribution between the LQ and HQ domains, which is supposed to learn more comprehensive and unbiased information and thereby ensures visual fidelity .

Furthermore, we propose a content-aware loss to enhance texture details and improve visual fidelity for automated disease screening and diagnosis, incorporating both pixel-level and feature-level constraints. Specifically, the wavelet-based pixel-level constraint ensures inter-domain structure preservation in the high-frequency component, which contains the most abundant texture information, rather than in the whole image space, encouraging the network to focus on translating critical structures from the LQ source image. To address the limited information representation of the pixel-level cycle-consistent loss, we propose a multi-encoder-based feature-level regularization is proposed for intra-domain distribution similarity with deep feature consistency. Due to the domain gap between low-level and high-level vision tasks, a high-quality reconstruction result may not necessarily guarantee good performance on downstream tasks (see Fig. 9). To alleviate this, we propose a cooperative training strategy with a bi-level learning scheme that jointly learns the UMIE task and the downstream tasks such as medical image segmentation and medical image classification. The aim is to generate HQ enhanced results that are both visually appealing and favorable for the downstream tasks.

To comprehensively evaluate the proposed method, we collect two datasets (named Fundus and Colonoscopy datasets) in addition to the one currently used in this field . Our experiments on the three datasets comprehensively show that HQG-Net consistently outperforms the existing methods.

Our contributions are summarized in the following respects:

We propose HQG-Net, which incorporates HQ cues to guide LQ image enhancement. In this way, we model UMIE under the joint distribution between LQ and HQ domains that contains more complete and unbiased information, and thus ensures the stability of the framework and improves the visual fidelity of the enhanced result.

We introduce a content-aware loss for textural enhancement and visual fidelity with the wavelet-based pixel-level and the multi-encoder-based feature-level constraints.

We propose a plug-and-play cooperative training strategy with a bi-level optimization formulation to produce a better downstream performance as well as enhanced results with better visually appealing effects.

We conduct extensive experiments on three datasets, including two newly collected datasets, and the results show that HQG-Net achieves state-of-the-art (SOTA) performance in terms of enhancement quality, downstream task performance, and computational efficiency. The newly collected datasets will be made publicly available.

II Related Work

Approaches to medical image enhancement can be broadly categorized into model-based and learning-based methods. Model-based methods involve exploiting a specific manually designed imaging prior and then empirically correcting the LQ image correction. Early techniques focused on the global enhancement of the LQ images through stretching the parameters of dynamic ranges, such as histogram equalization (HE) , image contrast normalization (ICN) , and contrast limited adaptive histogram equalization (CLAHE) . Retinex-based methods were later proposed by decomposing the LQ image into reflectance and illumination parts and recovering them with the corresponding prior knowledge. Additionally, IDRLP discovered a region line prior to enhance the LQ images with a region fidelity constraint. However, those traditional methods are primarily based on manually designed imaging priors, which fail to handle the complex degradation conditions and thus suffer from the limited application scope.

Learning-based methods are more flexible and can be extended to the unpaired setting, where paired LQ-HQ images are not required for model training. Intuitively, the unpaired problem can be solved by the structure of Pix2Pix with task-specific loss functions in an unsupervised framework . ADN employed multiple encoders and decoders to separate the content information from the degraded component and train the network with specialized loss in an unsupervised manner. EnlightenGAN proposed a multi-scale discriminator with a structure-preserving loss. Nevertheless, those Pix2Pix-based networks neglect the valuable information from the unpaired HQ images, which weakens their resilience to complex degradation conditions. To exploit the unpaired LQ and HQ images, CycleGAN was proposed with the cycle consistency constraint to learn a mapping between the LQ and the HQ domains. StillGAN was a CycleGAN-based method with a structure preservation loss and an illumination constraint. Unfortunately, the two-sided cycle-consistency strategy lacks the direct involvement of HQ images in the enhancement of LQ images, which inevitably brings domain shift and thus results in structure distortion . To solve these problems, we propose the explicit guidance-based HQG-Net with the variational information normalization module and the content-aware loss for structure preservation and visual fidelity. Furthermore, we design a plug-and-play cooperative training strategy to simultaneously acquire visual-appealing enhancement results and satisfactory downstream applications.

II-B Reference-based Super-Resolution and Style Transfer

Our HQG-Net is also related to reference-based image super-resolution (Ref-SR) and Style Transfer (Ref-ST), which mainly focus on natural images. While it may seem straightforward to extend methods from these fields to medical images, we argue that such naive extension often fails to produce satisfactory results as the unique characteristics of medical images are not properly modeled. In reference-based image super-resolution (Ref-SR), the generation of texture details heavily relies on the reference image via point-wise or patch-wise correspondence. CrossNet established a relationship between the reference image and the low-resolution (LR) image by flow estimation for point-wise alignment. For patch matching, NTT matched the multi-level features extracted from a pretrained VGG between the reference image and LR image in a concatenated manner. Style transfer aims to generate a re-stylized image by rendering the source image with the style component from a given reference image. By embedding the encoded reference image into a latent space, StyleGAN enabled controlled modifications in style transfer. To generate diverse images across multiple domains, StarGAN V2 introduced a style mapping network with domain-specific style codes extracted by specific reference images.

It should be noted that we are the first to handle the UMIE problem with the explicit assistance of reference images. Besides this, HQG-Net differs from existing Ref-SR and Ref-ST methods in several aspects. First, unlike source images in ref-SR or ref-ST that are of high quality, LQ images in UMIE often suffer from complex degradation conditions , such as light transmission disturbance and environment-induced artifacts. Therefore, a stable HQ cue extractor and general guidance are required to handle the complex degradation. To achieve this, we pretrain the HQ cue extraction module GG with images from both LQ and HQ domains to achieve stable HQ cue extraction and guide the enhancement in a condensed variational manner with the purpose of general guidance. Second, different from ref-ST and ref-SR, which prioritize visually appealing results, UMIE also emphasizes structure preservation for clinical requirements . However, both the feature-level alignment (ref-SR) and the reference-only guidance manner (ref-ST) can inevitably suppress the structure preservation capacity. To accommodate this requirement, we propose the task-specific content-aware loss and make two additional changes to our HQG-Net, including guiding the enhancement with the joint variational vector vc\mathbf{v}_{c} (including vz\mathbf{v}_{z} and vy\mathbf{v}_{y}), and using the adaptive normalization operator AdaLINAdaLIN.

By employing the joint vector and the adaptive normalization operator, we expect the enhanced results to simultaneously ensure high-quality reconstruction performance and preserve the critical structure information with the joint guidance strategy and the learnable hybrid normalization operation.

III Methodology

Unpaired Medical Image Enhancement (UMIE) aims to learn an enhancement model with a low-quality (LQ) image dataset Y\mathcal{Y} and a high-quality (HQ) image dataset Z\mathcal{Z}. Note that Y\mathcal{Y} and Z\mathcal{Z} are unpaired, i.e., there is not an image from Z\mathcal{Z} that is the high-quality version of an image from Y\mathcal{Y}. We propose a novel UMIE technique, dubbed HQG-Net, by explicitly employing HQ information to guide the enhancement of the LQ image. For structure preservation and visual fidelity, a novel content-aware loss is designed with the wavelet-based pixel-level constraint and the multi-encoder-based feature-level regularization. Furthermore, we propose a cooperative training strategy with a bi-level optimization formulation to produce a better downstream performance as well as visually appealing enhanced results. For better comprehension, we summarize the primary notations utilized in this paper, along with their corresponding meanings, in Table I.

HQG-Net consists of four main parts, including a high-quality cue extraction module GG, a generator EgenE_{gen} of the enhancement module EE, a variational information normalization module VINV_{IN}, and a discriminator EdisE_{dis}, where VINV_{IN} convey the HQ cue from GG to EE for high-quality (HQ) guidance. The specific network architecture of HQG-Net is shown in Fig. 3.

High-quality cue extraction. An HQ image z\mathbf{z} is randomly selected and then feed into GG to extract high quality cues,

Guided low-quality image enhancement. Following , our enhancement network is based on ResUNet50 with the last layer abandoned for computational efficiency. Besides, we plus several extra Variational Information Normalization (VIN) Modules that inject HQ cues in multiple levels. Specifically, given an LQ image y∈Y\mathbf{y}\in\mathcal{Y}, the ResNet50 encoder EeE_{e} generates a set of feature maps fkf_{k} (k∈{1,2,3,4}k\in\left\{1,2,3,4\right\}) in different levels.

which encodes the gist information from both the LQ image y\mathbf{y} and the HQ image z\mathbf{z}, and we use vc\textbf{v}_{c} as one input for the VIN modules inserted into the decoder in a multi-scale manner.

The decoder EdE_{d} also has a ResNet50-like backbone with four layers, except a VIN module appended after each residual block in the decoder. Each layer of the decoder consists of two components, the residual block and the VIN, which produce the corresponding output feature maps rkr_{k} and hkh_{k}, k∈{1,2,3,4}k\in\left\{1,2,3,4\right\}. In the residual block, the spatial resolution of the feature map hk+1h_{k+1} is first up-sampled by a factor of 2. Then the up-sampled feature is concatenated with the feature fkf_{k} from the encoder and compressed by a 1×11\times 1 convolution to reduce channel dimensions. The compressed feature map is further processed by a 3×33\times 3 convolution and generates the output feature map of the residual block rkr_{k}:

where up_sampleup\_sample, conv1conv1, conv3conv3 denote up-sampling, 1×11\times 1, and 3×33\times 3 operations, respectively. Note that r4=conv3(f4)r_{4}=conv3\left(f_{4}\right).

Variational information normalization (VIN). As shown in Fig. 3 (c), the VIN module takes the feature map rkr_{k} and the guidance of vector vc\mathbf{v}_{c} as inputs, and outputs an HQ-aware feature map hkh_{k} that is combined with the decoded feature map rkr_{k} for further processing. VIN is a residual connection structure with AdaLIN, LReLU, and a 3×33\times 3 convolution layer. Therefore, the final feature map hkh_{k} is formulated as follows:

where AdaLINAdaLIN is the AdaLIN operator, which is an information normalization block by combining the adaptive instance normalization (AdaIN) and the adaptive layer normalization (AdaLN) with a learnable parameter. The definition of the proposed AdaLIN is presented as follows:

where ρ\rho is a learnable parameter and is initialized as 0.5. Although AdaIN and AdaLN focus on instance and layer separately, they share a common structure with the same mathematical formulation. Take AdaIN as an example:

where μ\mu and σ\sigma correspond to mean and variance value. Being normalized by the concise and general variational information from the joint vector vc\textbf{v}_{c} (including vz\textbf{v}_{z} and vy\textbf{v}_{y}), HQG-Net can simultaneously exploit the HQ cues extracted from the vector-based HQ guidance and preserve significant structural information from the LQ input. This ensures the generation of enhancement results with visual fidelity and structure preservation, thereby facilitating clinical decision-making.

Discriminator. Our discriminator EdisE_{dis} is constructed following PatchGAN with patch-wise strategy, which is particularly beneficial for the LQ medical images with complex degraded conditions , e.g., blur and retinal artifact , and thus ensures the enhancement performance.

III-B Content-Aware Loss

Previous enhancement losses focus on the pixel-level consistency on both paired settings, e.g., mean square error, and unpaired conditions, e.g., cycle-consistency loss , for the structure preservation and visual fidelity, which brings two problems for UMIE. First, it can be challenging for pixel-level constraints to distinguish the complete structural information between the concerning biological tissues and LQ factors, especially for complex degradations. Therefore, a preliminary structure extraction operation can benefit the pixel-level loss function. Besides, the pixel-level constraint suffers from limited information, whereas in-depth features have a more robust representation capacity. Hence, an additional feature-level constraint is desirable to supplement this deficiency.

Since structural information is predominantly encoded in high-frequency components while exhibiting the texture sparsity in the whole image space, it is preferable to encourage structure fidelity in separately extracted high-frequency components rather than the whole-frequency image space. The wavelet-based operator is well-suited to this requirement . To this end, we adopt the Haar wavelet to construct a structure preservation loss, termed LHaar\mathcal{L}_{Haar}, between the LQ input y\mathbf{y} and the enhanced result x^\widehat{\mathbf{x}}. LHaar\mathcal{L}_{Haar} is defined as follows:

where x^HF\widehat{\mathbf{x}}^{HF} and yHF\mathbf{y}^{HF} are the integrated high-frequency parts of x^\widehat{\mathbf{x}} and y\mathbf{y} filtered by the wavelet-based high pass filters, including the vertical, horizontal, and diagonal filters. LHaar\mathcal{L}_{Haar} enforce the network to learn how to utilize the extracted HQ cues without sacrificing structure preservation from the LQ input, thus reducing the risk of misleading medical decisions.

For the second problem, a feature consistency loss, dubbed LFC\mathcal{L}_{FC}, is proposed by constraining the feature-level consistency between the enhanced result x^\widehat{\mathbf{x}} and the comprehensive guidance vc\textbf{v}_{c} in a vector form, which is formulated as:

where vx^=concate(G(x^),FCe(Ee(x^)))\mathbf{v}_{\widehat{\mathbf{x}}}=concate\left(G(\widehat{\mathbf{x}}),FC_{e}\left(E_{e}(\mathbf{\widehat{\mathbf{x}}})\right)\right), Gr(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) represents the Gram matrix for a more abstract feature characterization. \|\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}\|_{F} denotes Frobenius norm. As presented in Eq. 8, the enhanced result x^\widehat{\mathbf{x}} is expected to share a similar representation with the generated encoding information of the LQ input and the variational encoding feature of the HQ guidance in a compact vector form. Such a constraint contributes to the extraction of general HQ cues and ensures that the extracted HQ cues can be fully leveraged for image enhancement, thereby jointly ensuring the retention of feature-level information and visual fidelity.

Combining the Haar wavelet-based loss and feature consistency loss, we propose a content-aware loss, LCA\mathcal{L}_{CA}, for structure preservation, textural enhancement, and visual fidelity:

where λ1\lambda_{1} is a trade-off parameter. Intuitively, the content-aware loss balances the structural preservation from LQ input y\mathbf{y}, and the feature-based visual fidelity from HQ guidance z\mathbf{z}.

Apart from the content-aware loss, a standard GAN loss is used for the adversarial training to simultaneously regularize the generation ability and discrimination capability:

where E_{gen}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) generates an enhanced result x^\widehat{\mathbf{x}} conditioned on the LQ image y\mathbf{y} and the guidance image z\mathbf{z}, while {E_{dis}}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) tries to distinguish the differences between LQ and HQ images.

Combining the content-aware loss and the standard GAN loss, we reach our enhancement-oriented learning objective as

where λ2\lambda_{2} is the trade-off parameter.

III-C Bi-level Optimization and Cooperative Training

Bi-level optimization. Due to discrepancies in domain knowledge and training strategy, there inevitably exists a significant domain gap between low-level vision problems and subsequent high-level tasks, i.e., a high-quality reconstruction result in the low-level domain may not necessarily result in a good performance in high-level tasks . Inspired by the success of bi-level optimization (BLO) in parameter optimization that jointly updates the model parameters and the hyper-parameters , we introduce BLO to UMIE, aiming to achieve satisfactory performance on both the UMIE problem and downstream tasks. Following the truism Stackelberg’s theory , the segmentation-oriented enhancement is defined as a BLO formulation, which is formulated as follows:

where Ld\mathcal{L}^{d} is the objective function for the segmentation task and \Theta\left(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}};\bm{\omega}_{d}\right) represents the segmentation network enabled by the trainable parameters ωd\bm{\omega}_{d}. f\left(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}\right) is the enhancement term jointly driven by the two feasible constraints g_{e}\left(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}\right) and g_{v}\left(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}\right).

As presented in Fig. 4, the proposed bi-level formulation can improve the performance of both enhancement and segmentation tasks. Paradoxically, the ill-posed constraint in Eq. 13 is not bound in the form of an equation, which inevitably brings challenges for further optimization. Consequently, it is desirable to replace the complicated constraints with an adaptive framework and unroll the bi-level formulation into two networks, i.e., the segmentation network Θ\Theta and the enhancement network Φ\Phi, which is defined as follows:

where \Phi\left(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}};\bm{\omega}_{e}\right) is the proposed HQG-Net with the corresponding parameters ωe\bm{\omega}_{e}. In practice, considering the discrepancy of medical data and medical applications, different deep frameworks are adopted as the backbone for the downstream network Θ\Theta under the specific task. In particular, CS-Net and PNS-Net are correspondingly applied as the segmentation network for CCM and colonoscopy datasets, and Deepgrading are used as the classification network for nerve fiber tortuosity classification task with CCM data.

Cooperative training strategy. In this section, we further propose a cooperative training strategy for the jointly optimal network parameters, i.e., ω=(ωd,ωe)\bm{\omega}=\left(\bm{\omega}_{d},\bm{\omega}_{e}\right), with the aforementioned bi-level framework. By incorporating the downstream regularizer Ld\mathcal{L}^{d} and the enhancement regularizer Le\mathcal{L}^{e}, Eq. 14 can be rewritten from the perspective of mutual optimization:

where λ3\lambda_{3} is a hyperparameter for balance. We further present the gradient propagation flow of the parameters of downstream task ωd\bm{\omega}_{d} and enhancement network ωe\bm{\omega}_{e}:

where L\mathcal{L} denotes the overall energy function Eq. 15, which is subject to Eq. 16. The loss gradient of ωe\bm{\omega}_{e} is jointly determined by the loss functions of both enhancement and the downstream task, yielding visually pleasant results with favorable downstream performance.

Note that under the cooperative training strategy, the enhancement task is not unilaterally catered to the downstream task. Conversely, the downstream task can also assist the enhancement task to obtain better enhanced results. For instance, in medical image segmentation, mis-segmentation often occurs at the boundaries that contain crucial structural information and texture details, which are also challenging to highlight for UMIE. In this regard, the enhancement network, with additional constraints on the loss function of the segmentation task, can focus more on the boundary information and thus generate visually pleasant results with preserved structure and enhanced texture. Hence, as shown in Fig. 4, compared to the enhancement network optimized solely with the enhancement loss, the network optimized by our cooperative training strategy with a BLO formulation can achieve a more optimal enhancement solution with the assistance of constraints from the downstream task.

IV Experiment

Datasets. We employ three datasets, i.e., the corneal confocal microscopy (CCM) dataset, the Fundus dataset, and the Colonoscopy dataset, to evaluate enhancement performance under complex degeneration conditions. CCM dataset is publicly available , while Fundus and Colonoscopy datasets are the private datasets collected and relabelled by our collaborative clinicians into HQ and LQ subsets from the iSee dataset and the CVC-EndoSceneStill dataset . Details of the three datasets are presented in Table II. Note that CCM and Colonoscopy contain paired segmentation labels for the HQ and LQ images, which enables us to quantitatively evaluate the enhancement quality by taking segmentation as the downstream task and retrain our framework with the proposed cooperative training strategy for a bi-level optimization.

Implementation details. Our HQG-Net is trained with random flipping for data augmentation. Adam optimizer is applied with momentum terms (0.5,0.99)(0.5,0.99) and the learning rate is set as 1×10−41\times 10^{-4}. The batch size is set as 4, and the trade-off parameters λ1\lambda_{1}, λ2\lambda_{2}, λ3\lambda_{3} are set as 10, 1, 5, respectively. In the test phase, one HQ image in the corresponding test set is randomly selected as guidance for the LQ input. All the experiments are implemented with PyTorch on two RTX3090TI GPUs.

Compared methods. Ten state-of-the-art (SOTA) methods are selected in comparison, including three traditional methods, i.e., CARFI (Retinex-based), ATA (Retinex-based), and IDRLP (region line prior), and seven learning-based techniques, i.e., StarGAN V2 (style transfer), SSEN (ref-SR), TTSR (ref-SR), C2C^{2}-Matching (ref-SR), MASA-SR (ref-SR), StillGAN (CycleGAN-based), and EnlightenGAN (Pix2Pix-based). Note that we incorporate the style transfer-oriented or ref-SR-based techniques in natural scenarios into our experiment to evaluate the effectiveness of the task-specific modifications of our HQG-Net mentioned in Section II. All the compared methods are trained and tested with their corresponding default settings.

CCM. All the learning-based methods are trained on the CCM dataset . Four metrics, i.e., signal-to-noise ratio (SNR) with a radius of 3, 5, 7, and 9 pixels, are used for quantitative evaluation, where a higher score indicates a better enhancement result. As shown in Table III, the proposed HQG-Net achieves the best performance in all the metrics with a significant improvement. Furthermore, we also provide the results of StarGAN V2+, MASA-SR+, StillGAN+, EnlightenGAN+, and HQG-Net+, where + means the network is trained with the cooperative training strategy, and the higher metric scores obtained by all the re-trained networks indicate the generalization and effectiveness of the cooperative strategy.

The visualization results are presented in Fig. 5, where we only present partially enhanced results for space limitation. As shown in Fig. 5, CAEFI fails to enhance the LQ image with good visual fidelity. StarGAN V2 utilizes the same guidance image as our method, while the generated images suffer from texture blur and local luminance distortion, which lies in that the style transfer technique has neither an information maintenance framework nor a degradation-specific loss function. MASA-SR also fails to preserve critical structures and can be influenced by complex degradation due to its ignorance of the characteristics of medical images. StillGAN fails to enhance texture information and highlight the contrast for lacking the direct involvement of real HQ images in the enhancement, as well as the coarse structure loss. Global illumination uniformity is a problem for EnlightenGAN for the excessive attention to local details. In contrast, the enhanced results from HQG-Net enjoy high contrast and enhanced texture details, which is mainly attributed to the explicit HQ information guided HQG-Net and the content-aware loss function. In addition, the results of HQG-Net+ better conform to the definitions of HQ images in Fig. 1. Specifically, the LQ CCM data enhanced by HQG-Net+ can better enhance the texture details than those enhanced by HQG-Net, which owes to the segmentation-oriented cooperative training strategy.

Colonoscopy. We train all the learning-based methods with the private Colonoscopy dataset and evaluate the enhancement performance with four metrics, i.e., average gradient (AG) , entropy (EN) , natural image quality evaluator (NIQE) , and blind/reference image spatial quality evaluator (BRISQUE) . A higher value in AG or EN means better results, which is reversed in NIQE and BRISQUE. Table III exhibits the qualitative results, and the proposed HQG-Net has the best performance on all the metrics. Besides, we also demonstrate the superiority of the cooperative training strategy in terms of its capacity to improve enhancement performance. Fig. 5 also demonstrates the superiority of HQG-Net in visual comparison, where our enhanced results have excellent visual fidelity for our variation-guided framework and the task-specific loss function. Moreover, in Fig. 5, profiting from the cooperative training strategy, the enhanced results in HQG-Net+ have more uniform illumination than those in HQG-Net.

Fundus. Fundus dataset is used for training the compared networks with 640 HQ images and 700 LQ images. High-quality score (SHQS_{HQ}) and perception-based image quality evaluator (PIQE) are used for evaluation, where a higher value in SHQS_{HQ} or a lower score in PIQE means a better outcome. The qualitative and quantitative analyses are presented in Fig. 6. In qualitative analysis, we achieve the best performance in SHQS_{HQ} and PIQE, where our SHQS_{HQ} is 0.64 and is almost twice as high as the second-best result. Furthermore, the qualitative analysis also demonstrates that the enhanced fundus image has uniform illumination, enhanced texture details, and excellent visual fidelity. Note that the cooperative training strategy is not applicable to this task for the lack of paired labels.

Computational efficiency. We compare the parameters of deep networks and the corresponding FLOPs (at the size of 384×384384\times 384) on the enhancement task in Table IV. As shown in Table IV, the proposed HQG-Net achieves the lowest FLOPs, which demonstrates the efficiency of our HQG-Net.

IV-B Ablation Study

Effect of the HQ cue extraction module. As shown in Table V(a), when directly removing the HQ cue extraction module or replacing the HQ vector with a learnable tensor, the enhancement performance greatly declines, which verifies the effectiveness of our HQ guidance mechanism.

Consistency of the guidance representation. Furthermore, to verify the consistency of HQG-Net on the HQ cue extraction, we consider regularizing the HQ vector vz\textbf{v}_{z} to exhibit minimal variance when the HQ guidance image is altered. To achieve this, we add two constraints, namely LInterVarL_{InterVar} and LIntraVarL_{IntraVar}, to minimize the variance in either different homogeneous HQ guidance images or the same HQ guidance image with different views, which are formulated as:

where vz1\textbf{v}_{\textbf{z}_{1}} and vz2\textbf{v}_{\textbf{z}_{2}} are the vectors of two randomly selected HQ guidance images z1\textbf{z}_{1} and z2\textbf{z}_{2}. vz1\textbf{v}_{\textbf{z}^{1}} and vz2\textbf{v}_{\textbf{z}^{2}} are the vectors of HQ images z1\textbf{z}^{1} and z2\textbf{z}^{2}, where z2\textbf{z}^{2} is a different view of z1\textbf{z}^{1} with the random rotation and flip. As presented in Table V(b), the limited increment from the extra constraints, i.e., LInterVarL_{InterVar} and LIntraVarL_{IntraVar}, indicates that HQG-Net has already learned such consistent representation and thereby extract sufficiently stable HQ cues. This mainly attributes to the explicit regularization of our vector-based HQ guidance mechanism and the implicit constraint of the proposed content-aware loss.

Robustness analysis with the HQ cue extraction module trained with different data. To verify the feasibility to pretrain the HQ cue extraction module GG with both HQ and LQ data, we conduct a robustness analysis for HQG-Net with GG pretrained with different data, i.e., “HQ data only” and “HQ & LQ data”, by simulating different kinds of noise, including Gaussian, Poisson, and Salt & Pepper noises, for the HQ guidance image. As reported in Fig. 7, GG trained with “HQ data only” can be inevitably undermined by the fake HQ guidance with perturbations, while GG trained with “HQ & LQ data” can better resist the noises and extract more stable HQ cues. Guided by such HQ cues, HQG-Net can better withstand the complex degradation from the LQ inputs and thereby generate HQ enhanced results.

Robustness to the HQ guidance image. We first attest to whether the proposed HQG-Net is robust to the HQ guidance image and mainly consider three conditions, i.e., all HQ data from the corresponding test set (200 HQ images), 10 HQ images, and ours (1 fixed image), to randomly select the HQ image to guide the enhancement in the test phase. As shown in Table V(c), all three conditions yield similar performance, which evidently verifies our robustness to the HQ guidance image and further indicates that HQG-Net can extract general HQ cues from the HQ guidance images.

Using other types of images for guidance. To verify the advantage of using the HQ images from the same dataset for guiding the enhancement process, we replace the HQ images with other types of images, namely, (1) HQ heterogeneous medical images from another dataset, (2) LQ medical data, (3) natural images. Table VI shows the results. We can see that these alternative choices produce inferior results to using the HQ images from the same dataset as guidance.

However, as shown in Table VI, when being guided by HQ heterogeneous medical data, HQG-Net still generates promising enhanced results that exceed most comparison methods, which lie in the similarities of HQ medical images (see Fig. 1). For example, both HQ colonoscopy and HQ fundus data hold uniform illumination. Besides, HQG-Net introduces vector-based HQ guidance to ensure the extraction of valid HQ cues and employs the variational information normalization module to guide the enhancement with the concise and general variational information from the HQ cues. Therefore, even when the relevance of the HQ guidance image to the LQ image is not very tight, HQG-Net can also extract some valuable HQ cues and apply them to improve the enhancement performance.

Additionally, as depicted in Table VI, even under the guidance of the least relevant and even seriously degraded data, the enhanced performance is gracefully degraded to our Pix2Pix-based version, which demonstrates that introducing the vector-based HQ guidance in a variational manner is a stable and harmless way to improve the enhancement performance.

To sum up, the superiority of the enhanced results under the highly relevant HQ guidance, i.e., the HQ homogeneous images, and the robustness of them under the least relevant and seriously degraded guidance, i.e., the LQ natural images, validate that our guidance-based framework is a meaningful exploration for the UMIE task and can draw the attention of the community to handle the UMIE task with the explicit assistance of reference images. Besides, considering that it is very easy for a hospital to obtain an HQ image that is highly correlated with the LQ image, e.g., LQ and HQ data acquired by the same instrument with the same set of parameters for the same organs, our HQG-Net has a high clinical utility.

Effect of the vector-based guidance. We further validate the efficacy of the vector-based guidance with the joint vector vc\textbf{v}_{c}. In Fig. 8, HQG-Net with vector-based guidance can better enhance LQ images and moderate the slight perturbations from the HQ guidance, which verifies that the vector-based guidance helps to extract valid HQ cues. Besides, as exhibited in Table LABEL:table:EffectAdaLIN, guided by the joint vector vz\textbf{v}_{z} that combines the HQ vector vz\textbf{v}_{z} and structure vector vc\textbf{v}_{c}, HQG-Net is capable of preserving critical structural information and achieving superior enhancement performance.

Effect of the variational information normalization module. We verify the effect of the variational information normalization module by ablating AdaLIN (Table LABEL:table:EffectAdaLIN) and comparing AdaLIN with other integration strategies at the feature level (Table LABEL:table:AdaLINVersus). In Table LABEL:table:EffectAdaLIN, by canceling the multiscale strategy and ablating AdaLIN, we validate the effect of the multiscale AdaLIN operator.

Besides, when substituting AdaLIN for other integration strategies, namely feature-level concatenate, add, or multiple, we directly use the combined feature maps rather than the joint vector vc\textbf{v}_{c} for feature matching. For fairness, we also provide the results of AdaLIN with the combined features. Table LABEL:table:AdaLINVersus validates the superiority of AdaLIN over other information integration strategies, which attributes to the concise and general variational information from AdaLIN.

Effect of the content-aware loss. We further ablate the content-aware loss LCAL_{CA} (Table LABEL:table:AblationCAloss) and evaluate the optimal combination in LHaarL_{Haar} (Table V(g)). In Table LABEL:table:AblationCAloss, we compare several combinations of loss functions, where LsL_{s} is the structure loss in StillGAN . Table LABEL:table:AblationCAloss attests to the superiority of our content-aware loss. Besides, as shown in Table V(g), we try the different combinations of high-frequency components in LHaarL_{Haar} and find that combining all the high-frequency sub-bands results in optimal enhancement performance.

IV-C Downstream Tasks

Medical image segmentation task. Thanks to the paired segmented label for the CCM and Colonoscopy datasets, we further test the segmentation results with the corresponding segmentors, i.e., CS-Net and PNS-Net with their pre-trained models. Five commonly used metrics are applied for assessment, i.e., area under the ROC curve (AUC), accuracy (ACC), sensitivity (SEN), Dice coefficient (Dice), and G-mean score (G-Mean), where a higher value indicates a better performance. As shown in Table VII, the proposed HQG-Net achieves nine best values and one second-best value in the ten scores, demonstrating that the image enhanced by our method can best facilitate the downstream segmentation task. As shown in Fig. 9, the segmented results of HQG-Net are more accurate and continuous, which further verifies the superiority of HQG-Net for the highlight of the geometric structure, whether the foreground is filamentous or blocky. Moreover, in Table VII, compared with networks optimized solely with the enhancement loss, the corresponding network optimized by our cooperative training strategy achieves better performance, indicating the potential of the cooperative training strategy to generate enhancement results that are segmentation-friendly.

Nerve fiber tortuosity classification task. Previous studies indicate that the tortuosity level grading in CCM is highly relevant to some diseases , e.g., diabetic neuropathy and hypertensive retinopathy. Hence, we conduct an experiment on nerve fiber tortuosity with a SOTA classifier, Deepgrading , to explore how the enhancement task contributes to disease-related classifications on CORN-3 dataset , which contains about 300 images that are divided into four levels according to the tortuosity level. Three metrics are used to estimate the performance, i.e., ACC, F1-score, and precision (PRE), where a higher value means better performance. In Fig. 10, the proposed HQG-Net achieves ten best and two second-best values among the twelve metrics, which shows that the proposed HQG-Net can improve the identification rate of nerve fiber tortuosity and further demonstrate the clinical value of HQG-Net. In addition, as shown in Table VIII, the cooperative training strategy improves the accuracy of nerve fiber tortuosity classification by 2.3%2.3\% (HQG-Net), 2.0%2.0\% (StillGAN), 2.1%2.1\% (EnlightenGAN), and increases the F1-score of that by 3.2%3.2\% (HQG-Net), 4.4%4.4\% (StillGAN), 3.7%3.7\% (EnlightenGAN). These increments provide strong evidence of the superiority of the cooperative training strategy.

IV-D Clinical Decision-Making

Medical image reclassification task. In the image reclassification task, we invite two collaborative clinicians, who have participated in the production of our private datasets, to reclassify the LQ images (60 in CCM and 70 in Colonoscopy) and the HQ images (200 in CCM and 200 in Colonoscopy) enhanced by HQG-Net and HQG-Net+, where we randomly select homogeneous HQ guidance in the enhancement. To avoid bias from the experts, we leave out that these images have already been enhanced. As shown in Table IX(a), both HQG-Net and HQG-Net+ achieve promising performance in this reclassification task, which verifies that HQG-Net and HQG-Net+ have been successful in improving the quality of most LQ images from a clinical perspective. While in Table IX(b), none of the enhanced HQ images are identified as the LQ ones, thereby confirming that both HQG-Net and HQG-Net+ do not weaken the quality of the enhanced images.

Disease diagnosis task. In the disease diagnosis task, we create a new dataset, DiaRet, comprised of 300 LQ fundus images, obtained from 150 healthy eyes and 150 eyes affected by diabetic retinopathy. These images were selected from the “usable” grade of the Eye-Quality with their labels, i.e., with or without diabetic retinopathy. We invite two ophthalmologists to diagnose diabetic retinopathy from the given images, i.e., the original data and the data enhanced by our HQG-Net. To minimize subjective factors, the ophthalmologists are invited to perform the diagnostic task based on the original data first and to review the enhanced images three days later. The diagnostic results from the two ophthalmologists are presented in Table X, where HQG-Net improves the diagnosis results by 11.8%11.8\% (ACC), 3.4%3.4\% (SEN), 9.2%9.2\% (F1-score), and 14.8%14.8\% (PRE). Such results firmly demonstrate that HQG-Net is a promising method for facilitating clinical decision-making.

V Conclusion

In this paper, we propose a GAN-based network for unpaired medical image enhancement with the explicit guidance of unpaired HQ images (HQG-Net), which is the first work to model the medical image enhancement task under the joint distribution between both the LQ and HQ domains. The joint distribution-based framework implicitly contains more comprehensive and unbiased information than those methods, which only rely on the LQ domain. Therefore, HQG-Net can generate enhanced images with better visual fidelity. For high image contrast and enhanced textural details, the content-aware loss function is proposed with both wavelet-based pixel-level and feature-level consistency constraints. We further propose a bi-level formulation to jointly optimize the UMIE framework with the downstream tasks. Compared with the state-of-the-art techniques, the proposed HQG-Net achieves the best performance in the three distinct modalities both in the medical enhancement problem and the subsequent applicability validation tasks. Fundus and Colonoscopy datasets with enhanced quality labels will be released to the public with the source code in the future for community research.

VI ACKNOWLEDGMENT

The authors would like to express their sincere appreciation to the anonymous reviewers for their insightful comments, which greatly improved the quality of this paper.

References