RUN: Reversible Unfolding Network for Concealed Object Segmentation

Chunming He, Rihan Zhang, Fengyang Xiao, Chengyu Fang, Longxiang Tang, Yulun Zhang, Linghe Kong, Deng-Ping Fan, Kai Li, Sina Farsiu

Introduction

Concealed object segmentation (COS) aims to segment objects that are visually blended with their surroundings. It serves as an umbrella term with various applications, including camouflaged object detection (He et al., 2024b), polyp image segmentation (He et al., 2024a), and transparent object detection (Xiao et al., 2023), among others.

COS is a challenging problem due to the intrinsic similarity between the object and its background. Traditional methods address this challenge by relying on manually designed models with hand-crafted feature extractors tailored to subtle differences in textures, intensities, and colors (Wang et al., 2019). While offering clear interpretability, they often structure in complex scenarios. Deep learning advances COS by leveraging its strong generalization capabilities, driven by powerful feature extraction mechanisms. Early learning-based approaches, such as SINet (Fan et al., 2020a), primarily focus on foreground regions for segmentation, often overlooking discriminative cues in the background, leading to suboptimal performance (see Fig. 1). Recent algorithms, such as FEDER (He et al., 2023b), have sought to refine segmentation masks by reversibly modeling both foreground and background regions at the mask-level.

Reversible modeling enhances the network’s capacity to extract subtle discriminative cues by directing attention to uncertain regions—pixels with values that are neither 1 nor 0—thus improving segmentation results. However, current methods restrict the application of reversible strategies to the mask level, leaving the potential of the RGB domain underexplored. Such information can assist in identifying discriminative cues and enhancing segmentation quality. As shown in Fig. 2 (d) and (e), when reversibly separating the image into foreground and background regions based on the mask, the uncertainty regions in the mask tend to manifest as color distortion in the RGB space. Addressing these translates to a more precise separation of the foreground and background. In this case, two seemingly independent tasks—object segmentation and distortion restoration—share the same optimization goal. Existing research has shown that jointly optimizing such tasks helps guide the network toward an optimal solution (Xu et al., 2023).

To achieve this, we first introduce a deep unfolding network termed the Reversible Unfolding Network (RUN) for COS. RUN established a theoretical foundation to reversibly integrate the two aforementioned tasks, rather than directly combining them, to achieve more accurate segmentation. The COS task is formulated as a foreground-background separation process, and a new segmentation model is developed by incorporating a residual sparsity constraint to reduce segmentation uncertainties. The iterative optimization steps of the model-based solution are then unfolded into a multi-stage network, with each step corresponding to a stage. Each stage comprises two reversible modules: the Segmentation-Oriented Foreground Separation (SOFS) module and the Reconstruction-Oriented Background Extraction (ROBE) module. By integrating optimization solutions with deep networks, our RUN framework achieves an effective balance between interpretability and generalizability.

We implement the reversible strategy within the mask domain in SOFS and within the RGB domain in ROBE. In SOFS, the mask is initially updated strictly according to the optimization solution. Subsequently, the Reversible State Space (RSS) module, recognized for its strong capacity to extract non-local information, is employed to refine the segmentation mask using the previously estimated mask and background. In ROBE, the process begins with a mathematical update of the background. A lightweight network is then used to reconstruct the entire image, while simultaneously refining the background, based on the estimated foreground and background. Since the estimation of foreground and background regions is performed by distinct modules, their assessments of concealed content can differ (see Fig. 2 (d) and (f)). Regions of conflicting judgments are identified as distortion-prone areas during the reconstruction process (see Fig. 2 (g) and (h)). This auxiliary reconstruction task, which targets the resolution of such distortions, effectively directs the network’s attention to challenging regions where distinguishing between foreground and background is particularly difficult, improving segmentation performance.

As the stages progress, RUN incrementally facilitates reversible modeling of foreground and background in both the mask and RGB domains. This approach effectively focuses the network on uncertain regions, reducing false-positive and false-negative outcomes. Notably, RUN exhibits high flexibility, allowing seamless integration with existing methods to achieve further performance enhancements.

Our contributions are summarized as follows:

(1) We propose RUN for the COS task. To the best of our knowledge, this represents the first application of a deep unfolding network to address the COS problem, thereby balancing interpretability and generalizability.

(2) RUN proposes a novel COS model designed to reduce segmentation uncertainties and introduces SOFS and ROBE modules to integrate model-based optimization solutions with deep networks. By enabling reversible modeling of foreground and background across both the mask and RGB domains, RUN directs the network’s focus to uncertain regions, reducing false-positive and false-negative outcomes.

(3) Experiments on five COS tasks, as well as salient object detection, validate the superiority of our RUN method. Besides, its plug-and-play structure underscores the effectiveness and adaptability of unfolding-based frameworks for the COS task and other high-level vision tasks.

Related Works

Concealed object segmentation. Deep learning methods have advanced COS (Xiao et al., 2024). Among them, those using reversible techniques to segment from foreground and background aspects are gaining attention. PraNet (Fan et al., 2020b) introduced a parallel structure with reversible attention to enhance segmentation. FEDER (He et al., 2023b) used foreground and background masks to identify concealed objects with edge assistance. BiRefNet (Zheng et al., 2024) proposed a reconstruction module to refine the mask with gradient information. However, they only focus on the mask level, leaving the RGB domain underexplored. Hence, we propose the first deep unfolding network, RUN, for COS. RUN proposes a novel COS model and introduces SOFS and ROBE. By integrating optimization solutions with deep networks, RUN enables reversible modeling across mask and RGB domains, improving segmentation accuracy.

Deep unfolding network. The deep unfolding network, a well-established technique in low-level vision, integrates model-based and learning-based approaches (He et al., 2023a; Fang et al., 2025), offering enhanced interpretability compared to pure learning-based methods. However, its application in high-level vision remains underexplored, primarily due to the lack of explicit intrinsic models for high-level vision tasks. In this paper, we introduce a deep unfolding network, RUN, in COS and formulate a novel COS model. RUN achieves more accurate segmentation results by integrating optimization-based solutions with deep networks, verifying its potential for advancing COS.

Methodology

A concealed image C\mathbf{C} can be decomposed into its foreground region F\mathbf{F} and background region B\mathbf{B}, expressed as

Based on Eq. 1, the foreground and background regions can be obtained by optimizing the objective function:

where ψ(M)\psi(\mathbf{M}) and μ\mu are the regularization term and trade-off parameter for M\mathbf{M}. Due to the intrinsic ambiguity of foreground objects in concealed images and the diverse nature of their backgrounds, manually defining regularization terms for M\mathbf{M} and B\mathbf{B} can be challenging. To address this, we utilize deep neural networks to implicitly learn these constraints in a data-driven manner. Beyond the two intrinsic regularization terms above, we introduce an extra residual sparsity constraint \mathcal{S}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) to further refine segmentation and minimize uncertainties. This leads to the final objective function:

This design encourages the generation of segmentation masks with high certainty. Following the practice of (He et al., 2024a), pixels with values in the ambiguous range [0.4,0.6][0.4,0.6] are excluded from further consideration, while extreme values for M~\widetilde{\mathbf{M}} are set to 0.1 and 0.9 instead of 0 and 1 to allow greater flexibility for optimization.

2 RUN

We utilize the proximal gradient algorithm (Fang et al., 2025) to optimize Eq. 4, ultimately deriving the optimal mask M∗\mathbf{M}^{*} and background B∗\mathbf{B}^{*}:

The optimization process involves alternating updates of the two variables over iterations. Here we take the kthk^{th} stage (1≤k≤K1\leq k\leq K) to present the alternative solution process.

Optimizing Mk\mathbf{M}_{k}. First, the optimization function is partitioned to update the foreground mask Mk\mathbf{M}_{k}:

The solution comprises two terms: the gradient descent term and the proximal term. To address the proximal term, we introduce an auxiliary variable M^k\hat{\mathbf{M}}_{k}, resulting in:

where B0\mathbf{B}_{0} is initialized as zero. Having gotten Eqs. 7 and 8, M^k\hat{\mathbf{M}}_{k} can be solved by optimizing:

where M0\mathbf{M}_{0} is also initialized as zero. Note that wk\mathbf{w}_{k} and M~k\widetilde{\mathbf{M}}_{k} are constructed based on Mk−1{\mathbf{M}_{k-1}}. For the term \mathcal{S}(\mathbf{w}_{k}\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}(\hat{\mathbf{M}}-\widetilde{\mathbf{M}}_{k})), we employ a Taylor expansion rather than soft thresholding for flexibility in problem-solving. Following the practice of (Goldstein, 1977), we approximate \mathcal{S}(\mathbf{w}_{k}\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}(\hat{\mathbf{M}}-\widetilde{\mathbf{M}}_{k})) at the k−1thk-1^{th} iteration (for simplicity, we let \mathbf{R}=\mathbf{w}_{k}\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}(\hat{\mathbf{M}}-\widetilde{\mathbf{M}}_{k})), expressed as follows:

where LSL_{\mathcal{S}} is the Lipschitz constant. ∇S(Rk−1)\nabla\mathcal{S}(\mathbf{R}_{k-1}) is the Lipschitz continuous gradient function of S(Rk−1)\mathcal{S}(\mathbf{R}_{k-1}) with CSC_{\mathcal{S}}, a positive constant that can be omitted in optimization. Substituting into Eq. 9, we obtain the following equations:

where Qa=C2+LSwk2+μI\mathbf{Q_{a}}=\mathbf{C}^{2}+L_{\mathcal{S}}\mathbf{w}_{k}^{2}+\mu\mathbf{I}, I\mathbf{I} is an all-ones matrix, Qb=αLSwkwk−1+μI\mathbf{Q_{b}}=\alpha L_{\mathcal{S}}\mathbf{w}_{k}\mathbf{w}_{k-1}+\mu\mathbf{I}, \mathbf{Q_{c}}=\alpha L_{\mathcal{S}}\mathbf{w}_{k}(\mathbf{w}_{k}\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}\widetilde{\mathbf{M}}_{k}-\mathbf{Q_{d}})-\alpha\mathbf{w}_{k}\nabla\mathcal{S}(\mathbf{w}_{k-1}\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}{\mathbf{M}}_{k-1}-\mathbf{Q_{d}}), and \mathbf{Q_{d}}=\mathbf{w}_{k-1}\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}\widetilde{\mathbf{M}}_{k-1}.

Optimizing Bk\mathbf{B}_{k}. The optimization function of Bk\mathbf{B}_{k} is

Same as the optimization rule for Mk\mathbf{M}_{k}, the gradient descent term and the proximal term are correspondingly defined as:

The closed-form solution of B^k\hat{\mathbf{B}}_{k} can be acquired similarly:

2.2 Deep Unfolding Mechanism

We unfold the iterative optimization steps of the model-based solution into a multi-stage network, termed Reversible Unfolding Network (RUN), with each step corresponding to a stage. As shown in Fig. 3, each stage has two reversible modules: the Segmentation-Oriented Foreground Separation (SOFS) and Reconstruction-Oriented Background Extraction (ROBE) modules.

SOFS. SOFS, derived from Eqs. 8 and 13, utilizes \hat{\mathcal{M}}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) and \mathcal{M}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) to compute the optimization solution M^\hat{\mathbf{M}} and the refined mask M\mathbf{M} at each stage, respectively. Given Bk−1\mathbf{B}_{k-1}, Mk−1\mathbf{M}_{k-1}, and Mk−2\mathbf{M}_{k-2}, we define M^k\hat{\mathbf{M}}_{k} as follows:

Eq. 18 retains the same formulation as Eq. 13, but all originally fixed parameters, including \nabla\mathcal{S}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}), are relaxed to be learnable, improving the model’s generalizability.

To refine the initial mask M^k\hat{\mathbf{M}}_{k}, we introduce the Reversible State Space (RSS) module RSS(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}), which has a robust capacity for non-local information extraction. The RSS module incorporates two Visual State Space (VSS) VSS(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) modules (Liu et al., 2024b) with distinct perception fields. The VSS with a small perception field locally refines uncertain regions along the edges from the foreground perspective, while the VSS with a large perception field globally identifies missed segmented regions from the background perspective. This dual-perception mechanism ensures both accurate and comprehensive segmentation results. Following (He et al., 2023b), we also integrate an auxiliary edge output Ek\mathbf{E}_{k} to further enhance segmentation performance. Consequently, the computation of Mk\mathbf{M}_{k} and Ek\mathbf{E}_{k} is defined as:

where conv3conv3 is 3×33\times 3 convolution. VSS_{s}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) and VSS_{l}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) have small and large perception fields, incorporating convolutions with varying kernel sizes. For brevity, we omit the detailed description of VSS. Unlike low-level vision tasks, segmentation tasks, particularly inherently complex COS, strongly depend on semantic information. It is challenging to extract this fully using a shallow network. To address this, we adopt the common practice of leveraging deep features E(C)E(\mathbf{C}), extracted from an encoder (default: ResNet50 (He et al., 2016)). Rather than directly processing the concealed image, this approach enables the extraction of subtle discriminative features, achieving accurate segmentation.

ROBE. In ROBE, the calculation of B^k\hat{\mathbf{B}}_{k} relies on \hat{\mathcal{B}}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}), similar to Eq. 15 but with the fixed parameters made learnable:

This is essentially a dynamic fusion of the previously estimated background and the reversed foreground derived in the current stage. To refine B^k\hat{\mathbf{B}}_{k}, we propose a simple U-shaped network (Xu et al., 2023) with three layers, denoted as \mathcal{B}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}). However, as shown in Fig. 2, since separate modules estimate the foreground and background, their interpretations of the concealed content may differ. Hence, regions with conflicting interpretations are identified as distortion-prone areas in reconstruction. To address this, the network also generates a reconstructed result C^k\hat{\mathbf{C}}_{k}, formulated as:

C^k\hat{\mathbf{C}}_{k} is designed to be consistent with the concealed image, thereby mitigating distortions. This alignment fosters consistent judgments between SOFS and ROBE for foreground-background separation, improving segmentation accuracy. As the stages progress, RUN incrementally facilitates reversible modeling of the foreground and background in both the mask and RGB domains. This iterative process directs the network’s attention to regions of uncertainty, reducing false-positive and false-negative outcomes. Hence, RUN ensures robust and accurate segmentation performance.

Loss function. The loss function comprises a segmentation term and a reconstruction term. We adopt the training strategy from FEDER (He et al., 2023b) for the segmentation part. A mean square error loss governs the reconstruction component. The overall loss function is defined as

where KK is the number of stages. LBCEwL_{BCE}^{w} is the weighted binary cross-entropy loss, LIoUwL_{IoU}^{w} is the weighted intersection-over-union loss, and LdiceL_{dice} is the dice loss. GTsGT_{s} and GTeGT_{e} are the ground truth of the segmentation mask and edge.

Experiments

Implementation details. We implement our method using PyTorch on two RTX4090 GPUs. In line with (Fan et al., 2020a), we incorporate deep features from encoder-shaped networks into our framework. All images are resized to 352×352352\times 352 for the training and testing phases. During training, we use the Adam optimizer with momentum parameters (0.9,0.999)(0.9,0.999). The batch size is set to 36, and the initial learning rate is configured to 0.0001, which is reduced by 0.1 every 80 epochs. The stage number KK is set as 4. Additional parameters inherited from traditional methods are optimized in a learnable manner with random initialization.

We conduct experiments on various COS tasks and compare our performance with SOTA methods using standard metrics. Details on datasets and metrics are in Sec. A.1. Our superiority in concealed defect detection and salient object detection is verified in Secs. A.2 and A.3. For fairness, all results are evaluated with consistent task-specific evaluation tools. Except for COD, other tasks have few publicly open-sourced methods, limiting quantitative analysis.

Camouflaged object detection. As shown in Table 1, our method achieves SOTA performance across all three settings. In the common setting, it outperforms competing approaches across all three backbones: ResNet50 (He et al., 2016), Res2Net50 (Gao et al., 2019), and PVT V2 (Wang et al., 2022). This superior performance on four datasets, particularly on the largest dataset, COD10K, and the largest testing dataset, NC4K, underscores the robustness and generalization capabilities of our RUN framework. Furthermore, in the MIS and MS settings, our RUN adheres to the evaluation protocols of FEDER (He et al., 2023b) and delivers improved results over existing methods. As illustrated in Fig. 4, our method generates more complete and accurate segmentation maps. This is attributed to our jointly reversible modeling at both the mask and RGB levels.

Medical concealed object segmentation. We conducted experiments on two medical COS tasks, including polyp image segmentation (CVC-ColonDB and ETIS datasets) and medical tubular object segmentation (DRIVE and CORN datasets). Considering that recent SOTAs commonly use Transformer-based encoders, we adopt PVT V2 as our default encoder. As shown in Tables 2 and 3, our method achieves top performance across three tasks. Furthermore, the results in Fig. 4 confirm the effect of our approach in segmenting small polyps and fine vessels and nerves.

Transparent object detection. Accurately segmenting transparent objects is crucial for autonomous driving. As demonstrated in Tables 4 and 4, our RUN surpasses existing methods on two datasets, providing more precise segmentation of transparent objects compared to other approaches. These results highlight our potential to contribute to the advancement of autonomous driving.

2 Ablation Study

We conduct ablation studies on COD10K of the COD task.

Effect of SOFS. As presented in Table 7, replacing the deep features E(C)E(\mathbf{C}) with the concealed image results in performance decline, highlighting the critical role of incorporating deep features into the DUN-based framework. Additionally, we evaluate the impact of the state space-based structure by removing the RSS and VSS modules. The effectiveness of our reversible strategy is further validated by excluding the foreground prior M^k\hat{\mathbf{M}}_{k} and the background prior Bk−1\mathbf{B}_{k-1}. Finally, we confirm the utility of integrating the auxiliary edge output, contributing to performance improvements.

Effect of ROBE. As shown in Table 7, when replacing \mathcal{B}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) with other large-scale networks, i.e., the CNN-based network \mathcal{B}_{1}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) (Xu et al., 2023) and Transformer-based network \mathcal{B}_{2}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) (Fang et al., 2023), we observe no significant performance gains. This suggests that a simple network is sufficient for background extraction and image reconstruction. Furthermore, when the reconstructed output C^k\hat{\mathbf{C}}_{k} is removed, our RUN also produces suboptimal results.

3 Further Analysis, Applications, and Meanings

Performance on small objects or multiple objects. Small objects and multiple objects are challenging for lacking discriminative cues. To evaluate our performance on the two conditions, we filtered images from COD10K that satisfy these criteria, resulting in 1,0841,084 images having concealed objects smaller than a quarter of the entire image and 186186 images with multiple concealed objects. As shown in Tables 11 and 11, while the performance of all methods declines, our approach consistently outperforms the competition.

Performance on degraded COS scenarios. To assess the impact of environmental degradation, we followed (He et al., 2023a) to simulate haze on concealed images in COD10K and then evaluated the ability of the compared methods to resist degradation. As illustrated in Fig. 5, performance degrades as the haze concentration increases. However, our RUN demonstrates superior resilience to haze degradation, attributed to its multi-modality reversible modeling strategy. To enhance robustness, we replaced our reconstruction network \mathcal{B}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) with a more complex network from CoRUN (Fang et al., 2025), termed \mathcal{B}_{3}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}), which includes a pretrained dehazing model. This brought a novel unfolding network, RUN+, with \mathcal{B}_{3}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) incorporating the pretrained model. Fig. 5 indicates integrating \mathcal{B}_{3}(\mathchoice{\mathbin{\vbox{\hbox{\scalebox{0.5}{\displaystyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\textstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptstyle\bullet}}}}}{\mathbin{\vbox{\hbox{\scalebox{0.5}{\scriptscriptstyle\bullet}}}}}) enhances RUN’s robustness in resisting haze degradation. This underscores the potential of RUN in addressing degraded scenarios.

Potential applications of RUN. First, we test the effect of our RUN as a refiner, specifically by initializing M0\mathbf{M}_{0} with the results of existing methods. As shown in Table 11, our approach can enhance the performance of SOTA methods without requiring retraining. Furthermore, we incorporate the core structures of existing methods into our RUN framework, followed by retraining the entire network. This integration yields even greater improvements, demonstrating that the unfolding framework can function as a plug-and-play solution to enhance the performance of existing methods. For example, as shown in Fig. 6, we observe that while error predictions from FEDER influence FEDER-R, FEDER+ demonstrates better resilience to such errors.

Meanings of our framework. Beyond introducing the deep unfolding network to high-level vision for the first time and enabling reversible modeling across both mask and RGB domains, the proposed RUN framework offers the potential to establish a unified vision strategy. By combining image segmentation and image reconstruction, our RUN introduces a novel approach to unifying low-level and high-level vision. As shown in Fig. S2 in the appendix, unlike existing strategies, such as bi-level optimization (Xu et al., 2023), our unfolding-based combination strategy is underpinned by explicit theoretical guarantees with the two models better coupled. Moreover, as shown in Fig. 5 RUN+, using more complex low-level vision algorithms results in a strong ability to resist complex degradation. This motivates further exploration of unfolding-based combination strategies to enhance high-level vision algorithms’ resistance to environmental degradation and imaging interference. Simultaneously, it promotes low-level vision algorithms by integrating deep semantic information and high-level guidance. Together, these advancements ensure practical applicability in both real-world high-level and low-level vision tasks.

Conclusions

This paper proposes RUN to formulate the COS task as a foreground-background separation model. Its optimized solution is unfolded into a multistage network, where each stage comprises two reversible modules: SOFS and ROBE. SOFS applies the reversible strategy at the mask level and introduces RSS for non-local information extraction. ROBE employs a reconstruction network to address conflicting foreground and background regions in the RGB domain. Extensive experiments verify the superiority of RUN.

References

Appendix A Experiment

Camouflaged object detection. In this task, we follow the standard practice of SINet (Fan et al., 2020a) and perform experiments on four datasets: CHAMELEON (Skurowski et al., 2018), CAMO (Le et al., 2019), COD10K (Fan et al., 2021a), and NC4K (Lv et al., 2021). The CHAMELEON dataset comprises 76 images, while the CAMO dataset contains 1,250 images divided into 8 classes. The COD10K dataset includes 5,066 images categorized into 10 super-classes, and NC4K serves as the largest test set, with 4,121 images. For training, we use 1,000 images from CAMO and 3,040 images from COD10K. The remaining images from these two datasets, along with all images from the other datasets, constitute the test set. To evaluate performance, we employ four widely-used metrics: mean absolute error (MM), adaptive F-measure (FβF_{\beta}) (Margolin et al., 2014), mean E-measure (EϕE_{\phi}) (Fan et al., 2021b), and structure measure (SαS_{\alpha}) (Fan et al., 2017). Superior performance is indicated by lower values of MM and higher values of FβF_{\beta}, EϕE_{\phi}, and SαS_{\alpha}.

Medical concealed object segmentation. We evaluate the performance of our method on two specific tasks: polyp image segmentation and medical tubular object segmentation. For polyp image segmentation, we utilize two benchmarks: CVC-ColonDB (Tajbakhsh et al., 2015) and ETIS (Silva et al., 2014). The training protocol follows the setup of LSSNet (Wang et al., 2024). Quantitative evaluation is conducted using three commonly adopted metrics: mean Dice (mDice), mean Intersection over Union (mIoU), and structure measure (SαS_{\alpha}), where higher values indicate better performance. For medical tubular object segmentation, we evaluate our method on the DRIVEhttp://www.isi.uu.nl/Research/Databases/DRIVE/ and CORN (Ma et al., 2021) datasets, with training and inference conducted separately for each dataset. For the DRIVE dataset, training and inference adhere to the dataset’s predefined splits. For the CORN dataset, the last 70%70\% of the data is used for training, while the first 30%30\% serves as the test set. Following DSCNet (Qi et al., 2023), we employ three evaluation metrics: mDice, area under the ROC curve (AUC), and sensitivity (SEN), with higher values reflecting better performance. To ensure a fair comparison with state-of-the-art medical concealed object segmentation methods, which predominantly utilize transformer-based encoders, we adopt PVT V2 as the backbone for our encoder.

Transparent object detection. For a fair comparison, we use PVT V2 as our default backbone and conduct experiments on two datasets: GDD (Mei et al., 2020) and GSD (Lin & He, 2021). The training set consists of 2,980 images from GDD and 3,202 images from GSD, while the remaining images are reserved for inference. Consistent with GDNet-B (Mei et al., 2023), we evaluate performance using several metrics, including mIoU, and maximum F-measure (FβmaxF_{\beta}^{max}). Superior performance is indicated by lower values for MM, or higher values for mIoU and FβmaxF_{\beta}^{max}.

Concealed defect detection. In this task, we utilize PVT V2 as the default backbone. Consistent with established practices, we evaluate the generalization capacity of our RUN framework on the concealed defect detection task. Specifically, we use the model trained on the COD task to segment concealed objects in the CDS2K dataset (Fan et al., 2023a). Eight evaluation metrics are employed, where higher values indicate better performance for all metrics except MAE, for which lower values are preferred.

Salient object detection. We evaluate the performance of our method on five widely used benchmark datasets: DUT-OMRON (Yang et al., 2013), DUTS-test (Wang et al., 2017), ECSSD (Yan et al., 2013), HKU-IS (Li & Yu, 2015), and PASCAL-S (Li et al., 2014). The DUTS dataset contains 10,553 training images and 5,019 test images, referred to as DUTS-test. The DUT-OMRON and ECSSD datasets include 5,168 and 1,000 images, respectively. The HKU-IS dataset consists of 4,447 images featuring multiple foreground salient objects, while PASCAL-S comprises 850 images derived from the PASCAL VOC 2010 dataset (Everingham et al., 2010). For evaluation, we use the same metrics applied in COD.

A.2 Generalization on concealed defect detection

We evaluate the generalization capability of our RUN framework in the concealed defect detection task by directly using the model trained on the COD task. Detailed information about the experimental setup can be found in Sec. A.1. As shown in Table S2, our method achieves superior performance compared to existing state-of-the-art approaches, further highlighting the advancements and effectiveness of the RUN framework.

A.3 Generalization on salient object detection

We evaluate the generalization of our method in salient object detection. Details regarding the training configurations and datasets are provided in Sec. A.1. As shown in Table S2, our method outperforms existing state-of-the-art approaches, achieving a leading position. These results highlight the superiority of our approach and underscore the potential of unfolding-based frameworks for high-level vision tasks.

Appendix B Limitations and Future Work

As illustrated in Fig. 5, our RUN model, like other advanced methods, exhibits instability in degraded scenarios. This is primarily because environmental degradation exacerbates the challenge of extracting subtle discriminative information, bringing difficulties in capturing concealed objects. However, RUN+ demonstrates robustness to haze degradation, underscoring the potential of the RUN-based framework for addressing scenarios involving environmental degradation.

Future work will focus on integrating the RUN model with more advanced low-level vision algorithms to address increasingly complex scenarios involving diverse types of degradation, such as low light, blur, and noise. Additionally, incorporating large-scale algorithms into the unfolding-based multi-stage framework introduces significant computational and storage demands. Developing strategies to effectively integrate degradation-resistant models within this framework remains an important research direction.

Furthermore, as this work represents a pioneering effort in applying deep unfolding networks (DUNs) to high-level vision tasks, it opens the door for the development of more DUN-based algorithms in this domain. These future approaches are expected to better balance interpretability and generalizability, further advancing high-level vision tasks. Additionally, establishing an unfolding-based high-resolution COS method (Zheng et al., 2024) is also our goal.