FocusDiffuser: Perceiving Local Disparities for Camouflaged Object Detection

Jianwei Zhao, Xin Li, Fan Yang, Qiang Zhai, Ao Luo, Zicheng Jiao, Hong Cheng

Introduction

Camouflaged Object Detection (COD) is an important field within computer vision aimed at detecting objects that are intricately blended with their environments, rendering them nearly invisible. This technology is vital across various applications, including surveillance, search and rescue operations, and the analysis of medical imagery , where the ability to identify concealed items is crucial. At present, the foremost strategies in camouflaged object detection are centered around discriminative modeling . These techniques enhance conventional segmentation neural networks by weaving in a spectrum of supportive information, thereby sharpening the model’s ability to discern and detect camouflaged objects within complex scenes.

Despite the notable achievements of the de facto paradigm in camouflaged object detection, it is constrained by intrinsic shortcomings attributable to its discriminative nature. Firstly, the inherent design of discriminative models restricts their capability to grasp the fundamental probabilistic structures of data , complicating the management of intricate variations and advanced camouflage techniques. Secondly, their primary concentration on class differentiation, rather than a deep comprehension of data distribution, hampers their ability to generalize well to new or unseen camouflage patterns. Thirdly, these models are always vulnerable to noise and data irregularities, significantly undermining their precision in identifying camouflaged objects. Not to mention, discriminative models are notably prone to overfitting on training data, which diminishes their performance in practical, real-world situations. These limitations in discriminative models spark an intriguing question: Might the advanced representational abilities of generative models offer a more effective solution to camouflaged object detection?

With these considerations in mind, we explore generative models as a promising new direction for detecting camouflaged objects and introduce FocusDiffuser as our proposed solution. Fig. 1 illustrates the transition from discriminative mappings to a generative paradigm that effectively discerns and generates masks for camouflaged objects in scenes, guided by pertinent information as conditions. Our key concept involves detecting camouflaged objects through a ‘noise-to-mask’ strategy, methodically filtering out noise from randomly generated patterns, guided by conditions extracted from the image itself. Specifically, during the training process, Gaussian noise is methodically integrated into the ground truth camouflaged map to synthesize an initial noise. FocusDiffuser is adeptly calibrated to distill the noise, conditioned on meaningful information unearthed by built-in modules, thereby fostering the capability to infer the masks of camouflaged objects from stochastic beginnings. In the inference stage, FocusDiffuser adeptly backtracks along the diffusion pathway, systematically uncovering the camouflaged objects as heatmaps.

Diverging from the standard denoising diffusion model, our vision for FocusDiffuser is to specialize in understanding and detecting camouflaged objects, thus pioneering a novel model architecture equipped with specialized integrated modules. We believe that the key to detecting camouflaged objects lies in the details, and therefore, it is essential to amplify the model’s focus on intricate details. Thus, within our FocusDiffuser, we have configured two pivotal modules: the Boundary-Driven LookUp (BDLU) module, tasked with explicitly identifying discrepancies in details, and the Cyclic Positioning (CP) module, which adopts a broader perspective to cyclically pinpoint camouflaged objects. Our BDLU module draws inspiration from dense matching techniques, accomplishing the analysis of local discrepancies through targeted lookup operations. Additionally, we harness edge information for enhanced supervision, further refining the module’s precision in identifying subtle variations. As a complement, our CP module aligns with the denoising process of diffusion, continuously pinpointing targets from a broader scope, thus enhancing the overall detection capability within the diffusion framework. As a result, our FocusDiffuser boasts superior abilities in comprehending and identifying camouflaged objects, setting it apart from existing diffusion models across diverse application areas. Overall, our contributions include:

A novel generative approach for camouflaged object detection. FocusDiffuser represents one of the first forays into employing generative models to address the challenge of detecting camouflaged objects. We contend that this work provides a fresh viewpoint in the domain and lays the groundwork for future investigations.

A new diffusion model to enhance detail comprehension. FocusDiffuser introduces advanced capabilities for understanding fine details within the diffusion models. The integration of our BDLU and CP modules significantly augments the detail discernment capacity of diffusion models, pioneering their application in the field of camouflaged object detection.

State-of-the-art results on widely-used benchmarks. Our FocusDiffuser can accurately detect camouflaged objects, achieving state-of-the-art performance on benchmarks such as CAMO, COD10K and NC4K.

Related Work

Camouflaged Object Detection. Camouflaged object detection (COD) identifies hidden entities within their environments and has diverse applications. Initially, handcrafted features such as color and intensity variations , 3D convexity , and motion boundaries were used. Recently, deep learning methods have taken the lead, leveraging discriminative models to detect objects using edge , depth , and category information . Emerging research is exploring generative models for COD. While innovative training strategies inspired by generative models have been developed, they do not fully exploit generative capabilities. Our work introduces a novel diffusion model variant, pioneering the application of generative models in COD, offering new insights and advancements.

Diffusion Models. Diffusion models, a key member of the generative model family, have recently been widely applied to various AI tasks due to their scalability, stability, and superior understanding capabilities . In computer vision, their effectiveness in image and video generation, segmentation, and synthesis has been well-established by numerous studies . Our paper advances diffusion models with the novel FocusDiffuser for camouflaged object detection. Unlike the concurrent diffusion-based COD model , our approach significantly enhances detailed analysis capabilities in this domain.

Our Approach

Motivation. The potential of generative models like Stable Diffusion in comprehending objects within complex environments is well-explored. Their advanced pattern recognition and iterative refinement abilities offer a fresh perspective for revealing objects concealed within their environments. Inspired by these insights, we are motivated to explore the potential of diffusion models to enhance the field of camouflaged object detection by leveraging their powerful comprehension abilities.

2 FocusDiffuser

Fig. 2 provides the architecture of our FocusDiffuser, featuring an Image Conditional Encoder and a Denoising Diffusion Model with Boundary-Driven LookUp (BDLU) and Cyclic Positioning (CP) Modules.

Image Conditional Encoder. Given an image I\mathbf{I}, an image conditional encoder encodes it to produce hierarchical features {fs}s=14\{f_{s}\}_{s=1}^{4}. These features are then combined via a convolution module to create x0\boldsymbol{\mathit{x}}_{0}. Subsequently, a linear embedding layer is followed to upsample x0\boldsymbol{\mathit{x}}_{0}, yielding xc\boldsymbol{\mathit{x}}_{c}. After this, xc\boldsymbol{\mathit{x}}_{c} is concatenated with input noise ft\mathbf{f}_{t} and introduced into the diffusion processes.

Boundary-Driven LookUp. The Boundary-Driven LookUp (BDLU) module, receiving x0\boldsymbol{\mathit{x}}_{0} as its input, produces condition feature xl\boldsymbol{\mathit{x}}_{l} and edge map xe\boldsymbol{\mathit{x}}_{e} as outputs. Notably, inspired by , our BDLU features a unique lookup operation designed to enhance the exploration of local detail information. This extracted information then serves as conditions to guide the diffusion process.

Cyclic Positioning. At each denoising step, our Cyclic Positioning (CP) module generates a focus matrix region MM based on the intermediate result ft{\bf{f}}_{t} from the previous iteration. This process filters out noise and directs the model’s attention towards areas with camouflaged objects.

These modules combine to form our FocusDiffuser tailored for camouflaged object detection. It boasts enhanced detail analysis capabilities and aligns with the denoising process of diffusion models to dynamically update the focus areas within a scene. Next, we delve into the details of each module.

2.2 Image Conditional Encoder (ICE).

2.3 Boundary-Driven LookUp (BDLU).

where ⊗\otimes denotes element-wise multiplication. This method ensures a focused analysis on the vital local disparities required for effective boundary detection, optimizing the identification of camouflaged objects by leveraging the most relevant features.

It’s worth emphasizing that the core of our BDLU module is an advanced lookup operation, drawing inspiration from the latest developments in dense matching technologies . Yet, we’ve refined and adapted this concept to markedly improve the analysis of detail-rich information. This evolution ensures a more nuanced and effective methodology for detecting the intricacies of camouflaged objects, highlighting the diffusion model to enhance detail perception in complex visual environments. It primarily comprises two essential elements: Correlation Volume Pyramid and Correlation LookUp:

(a) Correlation Volume Pyramid: Given x1\boldsymbol{\mathit{x}}_{1}, we initiate the process by computing visual similarity through the construction of a comprehensive correlation volume. Unlike , this volume, denoted as V\mathbf{V}, is derived from the dot product of x1\boldsymbol{\mathit{x}}_{1} with itself, yielding a matrix that captures similarity across all spatial dimensions:

(b) Correlation LookUp: We refine the feature map by harnessing visual similarities from the correlation pyramid, denoted as LVL_{V}. For every pixel x=(u,v)x=(u,v) in the image I\mathbf{I}, a local grid with radius r\boldsymbol{r} outlines its neighborhood:

Furthermore, the feature map xl\boldsymbol{\mathit{x}}_{l} is concatenated with the latent U-net encoder feature um\boldsymbol{\mathit{u}}_{m}, and fed into a convolution block, C1\mathbf{C_{1}}. Simultaneously, um\boldsymbol{\mathit{u}}_{m} undergoes processing by a separate convolution block, C2\mathbf{C_{2}}. The outputs of both blocks are then directly combined, as shown:

This operation synergizes detailed feature maps with latent features for enhanced representation efficiency.

2.4 Cyclic Positioning (CP):

The loss of details from successive downsampling in conditional encoders significantly impacts the effectiveness of generative diffusion models, particularly in terms of segmentation precision. To complement the cyclical denoising steps inherent in diffusion models, we introduce the Cyclic Positioning module, an innovative approach aimed at accurately localizing camouflaged object regions with each denoising iteration.

Guided by intermediate results, we construct a bounding rectangle, indicating the predicted object from previous denoising iteration predictions, which allows us to create a binary mask Mi{M}_{i} for each denoising iteration ii, ranging from 1 to NN. This mask assigns 1 inside the rectangle and 0 outside, isolating the object of interest and eliminating redundant backgrounds. We leverage Mi{M}_{i} to isolate areas containing camouflaged objects through element-wise multiplication with the concatenated input image I\mathbf{I} and boundary predictions xe\boldsymbol{\mathit{x}}_{e}. This operation produces a new 4-channel image Io\mathbf{I}_{o}, which effectively removes unnecessary backgrounds, serving as enhanced detail input for the subsequent (i+1)(i+1)-th inference cycle.

It is important to highlight that initially, there is no pre-existing proposal mask for the first inference round, resulting in M0{M}_{0} being set to zero. We incorporate Io\mathbf{I}_{o} as an additional conditioning factor for the diffusion model, with a meticulously crafted integration process. The architecture is detailed in Fig. 3, where to enrich texture details, a three-tier structure featuring cascaded Texture Enhance Modules (TEM) is employed to generate a series of enhanced features xb∈{xbi}i=13\boldsymbol{\mathit{x}}_{b}\in\{\boldsymbol{\mathit{x}}_{b}^{i}\}_{i=1}^{3}. Inspired by the intricacies of the human visual system, TEM represents a highly specialized multi-branch and multi-scale convolution module. Specifically, the conditions {xbi}i=13\{\boldsymbol{\mathit{x}}_{b}^{i}\}_{i=1}^{3} are directly combined with the output features {ui}i=13\{\boldsymbol{\mathit{u}}^{i}\}_{i=1}^{3} from the initial three stages of the U-net encoder integrated into the diffusion model, as described by the following formula:

2.5 Forward Diffusion:

The Forward Diffusion (FD) step is a pivotal aspect of diffusion models, aimed at creating noisy masks from clean inputs by iteratively adding noise, a process vital for training. This step adheres to a Markov process, mathematically described as follows:

Here, βt{\beta}_{t} denotes the variance parameters of the Gaussian noise added at each step tt. As tt progresses from 1 to TT, ft\mathbf{f}_{t} evolves from the original camouflaged map f0\mathbf{f}_{0} through the process:

with αˉt=∏j=0tαj\bar{\alpha}_{t}=\prod_{j=0}^{t}{{\alpha}_{j}} and αj=1−βj{\alpha}_{j}=1-{\beta}_{j}. It’s essential to adjust f0\mathbf{f}_{0} to fit within the range [-b\boldsymbol{b}, b\boldsymbol{b}] before implementing FD:

The coefficient b\boldsymbol{b} plays a crucial role in modulating the signal-to-noise ratio of the diffusion (see Fig. 4), directly impacting the denoising outcome. Sec. 4.3.5 provides a detailed examination of bb and its effects.

2.6 Reverse Denoising:

2.7 Training Objective:

FocusDiffuser aims to predict the original camouflaged map f0\mathbf{f}_{0} directly, which necessitates measuring the loss between the conditional denoising outcome f0′\mathbf{f}_{0}^{{}^{\prime}} and the actual ground truth f0\mathbf{f}_{0} to guide the optimization of parameters Θ\Theta. To ensure high-quality generative camouflaged maps, Mean Squared Error (MSE) loss is employed. Moreover, to accentuate the separation between the camouflaged object and its background, we integrate Binary Cross-Entropy (BCE) and Intersection over Union (IoU) loss into the training objective. The overall training objective is represented as:

This formulation encapsulates the loss calculation, employing MSE for accuracy, BCE for binary classification of foreground and background, and IoU for spatial overlap, holistically enhancing the model’s performance on camouflaged object prediction.

Experiments

We evaluate our FocusDiffuser using three renowned Camouflaged Object Detection (COD) benchmarks: CAMO , COD10K , and NC4K . The CAMO dataset comprises 2,500 images, with an equal split between images containing camouflaged objects and those without. The COD10K dataset is more comprehensive, featuring 5,066 images with camouflaged objects, in addition to 3,000 background images and 1,934 non-camouflaged images, the latter often serving as negative controls. Consistent with prior studies , for training, we select subsets of 1,000 camouflaged images from CAMO and 3,040 from COD10K. For validation, we set aside 250 and 2,026 camouflaged images from these datasets, respectively. Additionally, the NC4K dataset, reserved exclusively for validation, tests the FocusDiffuser’s ability to generalize across different data scenarios.

1.2 Metrics:

To fairly evaluate FocusDiffuser, we employ four key metrics recognized in the COD community: structure measure (SαS_{\alpha}), mean E-measure (EϕE_{\phi}), weighted F-measure (FβwF_{\beta}^{w}), and mean absolute error (MM). Higher values denote better performance for SαS_{\alpha}, EϕE_{\phi}, and FβwF_{\beta}^{w}, indicated by an upward arrow (↑\uparrow), while lower values for MM signify improved accuracy, marked by a downward arrow (↓\downarrow).

1.3 Training Settings:

FocusDiffuser was built using the Pytorch and trained over approximately 200 epochs with batch sizes of 8 on an NVIDIA Tesla A100 GPU, ensuring full convergence. The SGD optimizer, set with a momentum of 0.9 and an initial learning rate of 0.001, facilitated the training process. This rate was progressively reduced following a cosine schedule. To optimize computational efficiency, input images were resized to 384×\times384 pixels, whereas the spatial noise resolution was adjusted to 288×\times288 pixels.

2 Main Results

To show the effectiveness of FocusDiffuser, we conducted a comprehensive comparison with 22 state-of-the-art (SOTA) models. This comparison was grounded in fairness, drawing results either directly from published studies or using their available trained models. As presented in Tab. 3, our model sets new benchmarks, leading across all metrics and datasets. Specifically, on the CAMO dataset, FocusDiffuser led the nearest competitor by 2.3% in SαS_{\alpha}, 2.4% in EϕE_{\phi}, 3.4% in FβwF_{\beta}^{w}, and 0.8% in MM, highlighting its superior detection capability in environments with complex camouflage patterns. Its proficiency was further demonstrated on the COD10K dataset, where it achieved the highest structure measure (Sα=0.875S_{\alpha}=0.875) and the lowest mean absolute error (M=0.020M=0.020), highlighting the model’s exceptional ability to discern details and focus on concealed edge information. Remarkably, on the NC4K dataset, FocusDiffuser exceeded the performance of existing state-of-the-art models without the need for further training, illustrating its strong generalization capabilities. This comprehensive performance not only proves FocusDiffuser’s adeptness at identifying camouflaged objects with high precision but also sets a new standard for the field, combining nuanced detection with robust adaptability.

2.2 Qualitative Analysis:

We benchmarked FocusDiffuser against four top SOTAs, highlighting its excellence in Fig. 5 with data from original studies. Challenges include blending objects with vague boundaries, subtle edges, and similar textures to the background. Existing models often underperform in these areas, whereas FocusDiffuser produces maps with clear object outlines, improved edges, and accurate object delineation. This demonstrates FocusDiffuser’s unmatched precision in detecting and enhancing camouflaged details for superior visual clarity.

3 Ablation Study

To highlight the Diffusion Model’s (DM) exceptional capabilities, we compared it against a baseline model with a standard decoder (2 transformer layers from Segmenter ), keeping all settings unchanged. Results in Tab. LABEL:tab:DM demonstrate the DM’s superiority, showing marked improvements in EϕE_{\phi}, FβwF^{w}_{\beta}, and MM. These metrics, essential for evaluating prediction mask quality, underscore the DM’s advanced representational efficacy.

3.2 Ablation for Boundary-Driven LookUp (BDLU):

BDLU is crafted to enhance edge discrimination by harnessing local grid similarities. We visualized the feature representation xl\boldsymbol{x}_{l} to clarify BDLU’s focus on areas where boundary distinction is most challenging, as illustrated in Fig. 6. Quantitative evaluations, shown in Tab. LABEL:tab:equipment, reveal a significant performance decline across all metrics without BDLU in FocusDiffuser, underscoring its essential role. Further adjustments to BDLU configurations were explored to assess its impact on detecting camouflaged objects. Removing boundary selection and lookup operation, directly incorporating x0edge\boldsymbol{x}^{edge}_{0} into the denoising process resulted in a marked performance drop. Solely feeding x0edge\boldsymbol{x}_{0}^{edge} without boundary selection into LookUp module further degraded all metrics.

Alternatively, discarding the lookup operation and utilizing the boundary selection feature x1\boldsymbol{x}_{1} for the denoising process achieved the second-best performance, highlighting the edge features’ significance in the lookup process. Visual comparisons between xl\boldsymbol{x}_{l} and x1\boldsymbol{x}_{1} (refer to Fig. 6, columns 3 and 4) show the former’s enhanced clarity around ambiguous boundaries. The lookup’s efficacy is determined by the local grid radius, rr, with an optimal rr being crucial for identifying similar features and clarifying indistinct boundaries. Optimum performance was observed with rr=2. As rr increases, it introduces extraneous features, diminishing the model’s capability to discern local variations, as documented in Tab. LABEL:tab:BDLU-CP-detail.

3.3 Ablation for Cyclic Positioning (CP):

CP is leveraged to pinpoint camouflaged objects with the aid of a coarse map, integrating prior object location knowledge to extract targeted rectangular areas from the original image and combine these with edge details. This approach emphasizes the object, enriching texture detail visibility and ensuring prediction precision and granularity. Fig. 7 depicts the CP mechanism, illustrating the iterative enhancement of prediction maps. Quantitative assessments in Tab. LABEL:tab:equipment suggest that omitting CP in FocusDiffuser leads to diminished outcomes. Moreover, three mask types, M, are evaluated to gauge the effect of background suppression on camouflage detection (see Tab. LABEL:tab:BDLU-CP-detail): ‘a’ normalizes the predicted camouflaged map f0′\mathbf{f}_{0}^{{}^{\prime}} for M, yet its subpar quality might exclude parts of camouflaged objects, yielding less optimal results than ‘b’. Contrarily, ‘b’ uses f0′\mathbf{f}_{0}^{{}^{\prime}} as a basis for a bounding box that encases camouflaged objects, creating a proposal mask M that better conserves semantic information and enhances detection robustness versus the coarse mask. ‘c’, setting M to 1 as a control, shows the least effectiveness without background exclusion.

3.4 Ablation for Denoising Steps 𝑵𝑵N:

We systematically investigate the influence of denoising steps NN during the inference phase. Tab. LABEL:tab:denoising displays the performance enhancement with increased NN from 2 to 4 for scale factors bb of 0.5 and 0.1. Nevertheless, we observed limited gain or even decline when NN exceeds 4.

3.5 Ablation Study for Scale Factor 𝒃𝒃b:

bb is employed to regulate the signal-to-noise ratio of FD. As depicted in Fig. 4, initially setting bb to 1 allows the noisy mask to remain distinguishable even at t=600 timestamps, suggesting a high signal-to-noise ratio. Consequently, this may lead to an abundance of straightforward training examples for FocusDiffuser, potentially resulting in a suboptimal model. In contrast, when bb is reduced to 0.1, the noisy mask becomes challenging to observe at t=600. Tab. LABEL:tab:denoising outlines the changes in metrics as bb decreases. Optimal performance for FocusDiffuser is achieved with bb=0.1.

Conclusion

We introduce FocusDiffuser, a conditional diffusion model featuring a Bound-ary-Driven LookUp (BDLU) and a Cyclic Positioning (CP) module, engineered to enhance the detection of camouflaged objects against their environments. The BDLU module aggregates similarity values within a local grid to refine feature representations, emphasizing areas where camouflaged objects blend with their surroundings. Concurrently, CP identifies and isolates camouflaged objects, creating a focus region matrix that separates them from the background. Our extensive experiments demonstrate these modules synergistically improve the identification and clarity of camouflaged objects in COD tasks. Furthermore, we posit FocusDiffuser’s methodology offers potential advancements in other computer vision areas requiring precise, high-quality predictions, such as high-resolution segmentation.

Acknowledgements

This research was supported by the National Natural Science Foundation of China (62203089, 62303092, 62103084); the National Key Research and Development Program of China (No. 2022YFB2503004); and the Sichuan Science and Technology Program (2022NSFSC0890, 2022NSFSC0865, 2021YFS0383, 2023YF G0024, 2022YFS0570, 2023YFS0213).

References