Real-world Image Dehazing with Coherence-based Pseudo Labeling and Cooperative Unfolding Network

Chengyu Fang, Chunming He, Fengyang Xiao, Yulun Zhang, Longxiang Tang, Yuelin Zhang, Kai Li, Xiu Li

Introduction

Real-world image dehazing (RID) is a challenging task that aims to restore images affected by complex haze in real-world scenarios. The goal is to generate visual-appealing results while enhancing the performance of downstream tasks . The atmospheric scattering model (ASM), providing a physical framework for real-world dehazing, is formulated as follows:

where P(x)P(x) and J(x)J(x) are the hazy image and the haze-free counterpart. AA signifies the global atmospheric light. t(x)t(x) characterizes the transmission map reflecting varying degrees of haze visibility across different regions.

Conventional methods are limited by fixed feature extractors, which struggle to handle the complexities of real haze. Although existing deep learning-based methods demonstrate improved performance, they face two significant challenges: (1) These methods do not accurately model the complex distribution of haze, leading to color distortion (as illustrated in fig. 1 DGUN ). (2) Real-world settings lack sufficient paired data for network training while optimizing the network with synthesized data brings a domain gap, limiting the generalizability of the models.

To overcome the first challenge, PDN first introduces unfolding network to the RID field. In specific, PDN unfolds the iterative optimization steps of an ASM-based solution into a deep network for end-to-end training, incorporating physical information into the deep network. However, PDN does not effectively leverage the complementary information between the dehazed image and the transmission map, bringing overfitting problems and resulting in detail blurring (see fig. 1).

In this paper, we introduce the COopeRative Unfolding Network (CORUN), also derived from the ASM-based function, to address PDN’s limitations and better model real hazy distribution. CORUN cooperatively models the atmospheric scattering and image scene by incorporating Transmission and Scene Gradient Descent Modules at each stage, corresponding to each iteration of the traditional optimization algorithm. To prevent overfitting, we introduce a global coherence loss, which constrains the entire pipeline to adhere to physical laws while alleviating constraints on the intermediate layers. These design choices collectively ensure that CORUN effectively integrates physical information into deep networks, thereby excelling in restoring haze-contaminated details, as depicted in fig. 1.

To enhance generalizability in real-world scenarios, we introduce the first RID-oriented iterative mean-teacher framework, named Coherence-based label generator (Colabator), designed to generate high-quality dehazed images as pseudo labels for training dehazing methods. Specifically, Colabator employs a teacher network, a dehazing network pretrained on synthesized datasets, to generate dehazed images on label-free real-world datasets. These restored images are stored in a dynamically updated label pool as pseudo labels for training the student network, which shares the same structure as the teacher network but with distinct weights. During network training, the teacher network generates multiple pseudo labels for a single real-world hazy image. We propose selecting the best labels to store in the label pool based on visual fidelity and dehazing performance.

To achieve this, we design a compound image quality assessment strategy tailored to the dehazing task, evaluating the global coherence of the dehazed images and selecting the most visually appealing ones without distortions for inclusion in the label pool. Additionally, we propose a patch-level certainty map to encourage the network to focus on well-restored regions of the dehazed pseudo labels, effectively constraining the local coherence between the outputs of the student model and the teacher model. As shown in fig. 1, Colabator, generating high-quality pseudo labels for network training, enhances the student dehazing network’s capacity for haze removal and color correction.

Our contributions are summarized as follows:

(1) We propose a novel dehazing method, CORUN, to cooperatively model the atmospheric scattering and image scene, effectively integrating physical information into deep networks.

(2) We propose the first iterative mean-teacher framework, Colabator, to generate high-quality pseudo labels for network training, enhancing the network’s generalization in haze removal.

(3) We evaluate our CORUN with the Colabator framework on real-world dehazing tasks. Abundant experiments demonstrate that our method achieves state-of-the-art performance.

Related Works

The dissonance between synthetic and real haze distributions often hinders existing Learning-based dehazing methods from effectively dehazing real-world images. Consequently, there’s a growing emphasis on tackling challenges specific to real-world dehazing .

Given the characteristics of real haze, RIDCP and Wang et al. proposed novel haze synthesis pipelines. However, relying solely on synthetic data limits models’ robustness in real-world dehazing scenarios. Recognizing the distributional disparities between synthetic and real haze, methods like CDD-GAN , D4 , Shao et al. , and Li et al. have utilized CycleGAN for dehazing. Despite this, the challenges inherent in GAN training often result in artifacts. Some approaches combine synthetic and real-world data, applying unsupervised loss to supervise real-world dehazing learning . However, these losses lack sufficient precision, leading to suboptimal results. Other methods leverage pseudo-labels , but the erroneous pseudo-labels cause degrade quality.

To address these challenges, we introduce a coherence-based pseudo labeling method termed Colabator. Our approach selectively identifies and prioritizes high-quality regions within pseudo labels, leading to enhanced robustness and superior generation quality for real-world image dehazing.

2 Deep Unfolding Image Restoration

Deep Unfolding Networks (DUNs) integrate model-based and learning-based approaches and thus offer enhanced interpretability and flexibility compared to traditional learning-based methods. Increasingly, DUNs are being utilized for various image tasks, including image super-resolution , compressive sensing , hyperspectral image reconstruction , and image fusion . DGUN proposes a general form of proximal gradient descent to learn degradation. However, it fails to decouple prior knowledge, relying solely on single-path DUN to model degradation and construct mappings, posing challenges in comprehending complex degradation. Yang and Sun first introduced DUNs to the image dehazing field and proposed PDN . However, PDN does not exploit the complementary information between the dehazed image and the transmission map, resulting in detail blurring. Our CORUN optimizes the atmospheric scattering model and the image scene feature through dual proximal gradient descent, thus preventing overfitting and facilitating detail restoration.

Methodology

We propose the Cooperative Unfolding Network (CORUN), the first Deep Unfolding Network (DUN) method utilizing Proximal Gradient Descent (PGD) to optimize image dehazing performance. CORUN leverages the Atmospheric Scattering Model (ASM) and neural image reconstruction in a cooperative manner. Each stage of CORUN includes Transmission and Scene Gradient Descent Modules (T&SGDM) paired with Cooperative Proximal Mapping Modules (T&S-CPMM). These modules work together to model atmospheric scattering and image scene features, enabling the adaptive capture and restoration of global composite features within the scene.

Where J\mathbf{J} means the clear image without hazy, I\mathbf{I} is the all-one matrix. Based on eq. 2, we can define our cooperative dehazing energy function like

where ψ(J)\psi(\mathbf{J}) and ϕ(T)\phi(\mathbf{T}) are regularization terms on T\mathbf{T} and J\mathbf{J}. We introduce two auxiliary variables T^\hat{\mathbf{T}} and J^\hat{\mathbf{J}} to approximate T\mathbf{T} and J\mathbf{J}, respectively. This leads to the following minimization problem:

Transmission optimization. Give the estimated coarse transmission map T\mathbf{T} and dehazed image J^k−1{\hat{\mathbf{J}}_{k-1}} at iteration k−1k-1, the variable T\mathbf{T} can be updated as:

We construct the proximal mapping between T^\hat{\mathbf{T}} and T\mathbf{T} by a encoder-decoder like neural network which we named T-CPMM and denoted as prox⁡ϕ\operatorname{prox}_{\phi}:

the auxiliary variables T^\hat{\mathbf{T}}, which we calculate by our proposed TGDM can be formulated as:

The variable λk\lambda_{k} is a learnable parameter, we enable CORUN to learn this parameter at each stage during the end-to-end learning process, allowing the network to adaptively control the updates in iteration.

Scene optimization. Give T^k{\hat{\mathbf{T}}_{k}} and J\mathbf{J}, the variable J\mathbf{J} can be updated as:

Same as the proximal mapping process in the transmission optimization, S-CPMM has the similar structure as T-CPMM but different inputs, we denote S-CPMM as prox⁡ψ\operatorname{prox}_{\psi}:

where the J^k\hat{\mathbf{J}}_{k} we process by our SGDM can be presented as:

as the λk\lambda_{k} in transmission optimization, μk\mu_{k} is also a learnable parameter to bring more generalization capabilities to the network.

Details about CPMM. T-CPMM and S-CPMM share the same structure, which is modified from MST for improved mapping quality. Each CPMM block uses a 4-channel convolution to embed T\mathbf{T} and J\mathbf{J} into a 30-dimensional feature map. The distinction between T-CPMM and S-CPMM lies in their outputs: T-CPMM produces a 1-channel result to aid TGDM in predicting a scene-compliant transmission map, whereas S-CPMM generates a 3-channel RGB image. This enables S-CPMM to learn additional scene feature information, such as atmospheric light and blur, assisting SGDM in generating higher-quality dehazed results with more details. For more efficient computation, each CPMM comprises only 3 layers with $$ blocks, doubling the dimensions with increasing depth.

2 Coherence-based Pseudo Labeling by Colabator

We generate and select pseudo labels using our proposed plug-and-play coherence-based label generator, Colabator. Colabator consists of a teacher network with weights θtea\mathcal{\theta}_{tea} shared with the student network θstu\mathcal{\theta}_{stu} via exponential moving average (EMA). It employs a tailored mean-teacher strategy with a trust weighting process and an optimal label pool to generate high-quality pseudo labels, addressing the scarcity of real-world data. Figure 3 illustrates the pipeline of our Colabator.

where Ψ\varPsi is compose sequence to map and resize as PHQ~R\mathbf{P}^{R}_{\widetilde{HQ}}, ,norm(⋅)\text{norm}(\cdot) means normalize scores from 0 to 1, that higher score means lower haze density and better image quality.

Optimal label pool. To ensure the use of optimal pseudo-labels and avoid domain adaptation collapse due to instability during training, we proposed an optimal label pool P\mathcal{P} to maintain the pseudo-labels in their optimal state. The overall procedure of our optimal label pool process is summarized in algorithm 1, compare pseudo-dehazed image PHQ~Ri{\mathbf{P}^{R}_{\widetilde{HQ}}}_{i} with previous pseudo-label PPseRi{\mathbf{P}^{R}_{Pse}}_{i} and update pseudo-dehazed image as pseudo-label if it better than previous. To summarize the algorithm 1 and eq. 11, the overall process of Colabator can be formalize as:

where C\mathcal{C} is our Colabator framework, PPseR\mathbf{P}^{R}_{Pse} is the paired pseudo label of As(PLQR)\mathcal{A}_{s}(\mathbf{P}^{R}_{LQ}), TPseR\mathbf{T}^{R}_{Pse} is the corresponding pesudo transmission map, wpsew_{pse} means the trusted weight of the pseudo label.

Weights update. The teacher network’s weights θtea\theta_{tea} are updated by exponential moving average (EMA) of the student network’s weights θstu\theta_{stu}, which is denoted as follows:

where η\eta is momentum and η∈(0,1)\eta\in(0,1). Using this update strategy, the teacher model can aggregate previously learned weights immediately after each training step, ensuring updating stability.

3 Semi-supervised Real-world Image Dehazing

To achieve success in real-world dehazing, we designed several loss functions for our CORUN and Colabator to constrain their learning process. We introduce a reconstruction loss using the L1L_{1} norm \|\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}\|_{1}. To enhance visual perception, we employ contrastive and common perceptual regularization to ensure the consistency of the reconstruction results with the ground truth in terms of features at different levels. The perceptual loss is defined as follows:

where PHQ\mathbf{P}_{HQ} is the dehazed result, φi(⋅)\varphi_{i}(\cdot) means the ithi_{th} hidden layer of pre-trained VGG-19 , τi\tau_{i} is the weight coefficient. Besides, to constrain the entire pipeline to obey physical laws while alleviating constraints on the intermediate layers, and prevent overfitting, we introduce a global coherence loss:

where ⊙\odot is the Hadamard product, I\mathbf{I} means the all-ones matrix as the same size of PLQS\mathbf{P}^{S}_{LQ}. The global coherence loss ensures that CORUN can more efficiently integrate physical information into the deep network to facilitate the recovery of more physically consistent details. In addition, we introduce a density loss LdensL_{dens} based on D(⋅)\mathcal{D}(\cdot) to score and constraint the model to dehaze in the semantic domain:

where Aw\mathcal{A}_{w} means weakly geometric data augment, PHQS\mathbf{P}^{S}_{HQ} means the dehazed result of synthetic hazy image, and THQS\mathbf{T}^{S}_{HQ} is the corresponding transmission map. In the pre-training phase, our CORUN is optimized end-to-end using two supervised loss functions. The overall loss of the pre-training phase:

where ρr\rho_{r} is the trade-off weight of LReccontraL^{contra}_{Rec}, ρc\rho_{c} is the trade-off weight of LCohL_{Coh}.

Fine-tuning phase. In fine-tuning phase, we adapt our CORUN pre-trained on synthetic data to the real-world domain by our Colabator framework. For more steady learning, in this phase, we train with both synthetic and real-world data. As eq. 13, we generate PHQR,THQR,PPseR,TPseR,wpse\mathbf{P}^{R}_{HQ},\mathbf{T}^{R}_{HQ},\mathbf{P}^{R}_{Pse},{\mathbf{T}^{R}_{Pse}},w_{pse} from PLQR\mathbf{P}^{R}_{LQ}, and we get PHQS,THQS\mathbf{P}^{S}_{HQ},\mathbf{T}^{S}_{HQ} use the eq. 18. The overall loss of the fine-tuning phase:

Experiments

Data Preparation. We use RIDCP500 dataset, comprising 500 clear images with depth maps estimated by , and follow the same way of RIDCP for generating paired data. During the fine-tuning phase, we incorporate the URHI subset of RESIDE dataset , which only consists of 4,807 real hazy images, for generating pseudo-labels and fine-tuning the network. We evaluate our framework qualitatively and quantitatively on the RTTS subset, which comprises over 4,000 real hazy images featuring diverse scenes, resolutions, and degradation. Fattal’s dataset , comprising 31 classic real hazy cases, serves as a supplementary source for cross-dataset visual comparison.

Implementation Details. Our framework is implemented using PyTorch and trained on four NVIDIA RTX 4090 GPUs. During the pre-training phase, we train the network for 30K iterations, optimizing it with AdamW using momentum parameters (β1=0.9,β2=0.999)(\beta_{1}=0.9,\beta_{2}=0.999) and an initial learning rate of 2×10−42\times 10^{-4}, gradually reduced to 1×10−61\times 10^{-6} with cosine annealing . In Colabator, the initial learning rate is set to 5×10−55\times 10^{-5} with only 5K iterations. Following , we employ random crop and flip for synthetic data augmentation. We use DA-CLIP as our haze density evaluator and MUSIQ as the image quality evaluator. Our CORUN consists of 4 stages and the trade-off parameters in the loss are set to βc,ρr,ρc\beta_{c},\rho_{r},\rho_{c} are set to 0.2,5,10−2,0.2,5,10^{-2}, respectively.

Metrics. We utilize the Fog Aware Density Evaluator (FADE) to assess the haze density in various methods. However, FADE focuses on haze density exclusively, overlooking other crucial image characteristics such as color, brightness, and detail. To address this limitation, we also employ Blind/Referenceless Image Spatial Quality Evaluator (BRISQUE) , and Neural Image Assessment(NIMA) for a more comprehensive evaluation of image quality and aesthetic. Higher NIMA scores, along with lower FADE and BRISQUE scores, indicate better performance. We use PyIQA for BRISQUE and NIMA calculations, and the official MATLAB code for FADE calculations. All of these metrics are non-reference because there is no ground-truth in RTTS .

2 Comparative Evaluation

We compare our method with 8 state-of-the-art methods: PDN , MBDN , DH , DAD , PSD , D4 , RIDCP , DGUN . The quantitative results, presented in table 1, show that our method achieved the highest performance, outperforming the second-best method (RIDCP) by 19.0%19.0\%. Specifically, our method improved FADE, BRISQUE, and NIMA scores by 20.4%20.4\%, 30.8%30.8\%, and 7.0%7.0\%, respectively. This demonstrates that our method surpasses current state-of-the-art techniques in both dehazing capability and the quality, and aesthetics of the generated images.

The visual comparisons of our proposed method and state-of-the-art algorithms are shown in figs. 5 and 4. We can observe that these methods have demonstrated some effectiveness in real-world dehazing tasks, but when images containing white objects, sky, or extreme haze, the results from PDN, DAD, PSD, and RIDCP exhibited varying degrees of dark patches and contrast inconsistencies. Conversely, D4 caused an overall reduction in brightness, leading to detail loss in darker areas. Under these conditions, DGUN produced relatively aesthetically pleasing results but lost significant local detail, impairing overall visual quality. Notably, PSD achieved higher brightness but suffered from severe oversaturation. CORUN+ consistently outperforms others by producing clearer images with natural colors and better contrast, effectively removing haze while preserving image details.

3 Ablation Study

Generalization and Effect of Colabator. We evaluates the performance and the impact of our proposed Colabator framework across different metrics. As shown in table 4, removing the fine-tuning phase of Colabator led to significant performance drops, highlighting its critical role in the dehazing process. To evaluate the generalizability of Colabator, we conducted additional experiments by replacing our CORUN with the DGUN , while maintaining consistent training settings. Results in table 4 and fig. 1 indicate that Colabator substantially enhances DGUN’s performance, demonstrating its effectiveness as a plug-and-play paradigm with strong generalization capabilities.

Effect of Colabator. We validate the effect of our Colabator. In table 4, we systematically removed critical components, such as iterative mean-teacher (IMD), trusted weighting, and the optimal label pool, from the model architecture. The outcomes indicate the performance deteriorates when these components are removed, highlighting their essential role in the system.

Ablations on stage number. The number of stages in a deep unfolding network significantly impacts its efficiency and performance. To investigate this, we experimented with different stage numbers for CORUN+, specifically choosing kk values from the set {1,2,4,6}\{1,2,4,6\}. The results detailed in table 4, indicate that CORUN+ achieves high-quality dehazing with 4 stages. Notably, increasing the number of stages does not necessarily improve outcomes. Excessive stages can increase the network’s complexity, hinder convergence, and potentially introduce errors in the results.

4 User Study and Downstream Task

User Study. We conducted a user study to evaluate the human subjective visual perception of our proposed method against other methods. We invited five experts with an image processing background and 16 naive observers as testers. These testers were instructed to focus on three primary aspects: (i) Haze density compared to the original hazy image, (ii) Clarity of details in the dehazed image, and (iii) Color and aesthetic quality of the dehazed image. The results for each method, along with the corresponding hazy images, were presented to the testers anonymously. They scored each method on a scale from 1 (worst) to 10 (best). The hazy images were selected randomly, with a total of 225 images from RTTS and 54 images from Fattal’s dataset. The user study scores are reported in table 6, showing that our method achieved the highest average score.

Downstream Task Evaluation. The performance of high-level vision tasks, e.g. object detection and semantic segmentation, is greatly affected by image quality, with severely degraded images often leading to erroneous results . To address this performance degradation, some methods have incorporated image restoration as a preprocessing step for high-level vision tasks. To validate the effectiveness of our approach for high-level vision, we utilized pretrained YOLOv3 , and tested it on the RTTS dataset, and evaluated the results using the mean Average Precision (mAP) metric. As shown in table 6 and fig. 6, our method demonstrates a substantial advantage over existing methods, verifying our efficacy in facilitating high-level vision understanding.

5 Limitations and Future Work

In fig. 7, our CORUN+ model struggles to maintain result quality and preserve texture details when dealing with severely degraded inputs, such as strong compression and extreme high-density haze. This challenge persists across existing methods and remains unresolved. We attribute this difficulty to the model’s struggle in reconstructing scenes from dense fog, where information is often severely lacking or entirely lost, affecting the reconstruction of both haze-free and low haze density areas. Moreover, the model solely focuses on defogging and lacks the capability to address other image degradations, such as image deblurring and low-light image enhancement , limiting its ability to achieve high-quality reconstruction results from complex degraded images. To address this limitation in future research, we propose not only focusing on environmental degradation but also considering additional information about image degradation when solving real-world dehazing problems. In addition to this, we can introduce more modalities as supplements to RGB images, enhancing the model’s ability to effectively recover details.

6 Broader Impacts

Real-world image dehazing is a crucial task in image restoration, aimed at removing haze degradation from images captured in real-world scenarios. In computer vision, dehazing can benefit downstream tasks such as object detection , image segmentation , and depth estimation , with applications ranging from autonomous driving to security monitoring. Our paper introduces a cooperative unfolding network and a plug-and-play pseudo-labeling framework, achieving state-of-the-art performance in real-world dehazing tasks. Notably, image dehazing techniques have yet to exhibit negative social impacts. Our proposed CORUN and Colabator methods also do not present any foreseeable negative societal consequences.

Conclusions

In this paper, we introduce CORUN to cooperatively model atmospheric scattering and image scenes and thus incorporate physical information into deep networks. Furthermore, we propose Colabator, an iterative mean-teacher framework, to generate high-quality pseudo-labels by storing the best-ever results with global and local coherence in a dynamic label pool. Experiments demonstrate that our method achieves state-of-the-art performance in real-world image dehazing tasks, with Colabator also improving the generalization of other dehazing methods. The code will be released.

References