A Survey of Camouflaged Object Detection and Beyond
Fengyang Xiao, Sujie Hu, Yuqi Shen, Chengyu Fang, Jinfa Huang, Chunming He, Longxiang Tang, Ziyun Yang, Xiu Li
Introduction
Object detection, a fundamental task in computer vision, involves identifying and locating objects within images or videos. It comprises various fine-grained subfields: generic object detection (GOD), salient object detection (SOD), and camouflaged object detection (COD). GOD aims to detect general objects, while SOD identifies prominent objects that stand out from the background. In contrast, COD targets those objects that blend into their surroundings, making it an extremely challenging task. Fig. 1 illustrates the relationship between the target dog and its background across GOD, SOD, and COD tasks, sourced from the classical datasets of them.
COD has recently garnered increasing attention and rapid development for its advantages in facilitating the development of visual perception for nuance discrimination and promoting various valuable real-life applications, rang-
ing from concealed defect detection in industry and pest monitoring in agriculture to lesion segmentation in medical diagnosis and art, such as recreational art and photo-realistic blending.
However, unlike GOD and SOD, COD involves detecting objects that are purposefully designed to be inconspicuous, like the Dalmatian hidden in the forest on the far right of Fig.1 which is difficult to detect due to its camouflage with the surroundings, thus requiring more sophisticated detection strategies. COD can be further classified into image and video tasks. Normal COD, i.e., image-level COD, to detect camouflaged objects in static images, whereas video-level COD, dubbed VCOD, deals with detecting these objects within video sequences. The latter one introduces additional complexity due to temporal continuity and dynamic changes, necessitating models capable of extracting both spatial and temporal features effectively.
Traditional methods for COD and VCOD, including tex-
ture, intensity, color, motion, optical flow, and multi-modal analysis, have demonstrated their strengths in specific scenarios but also exhibit notable shortcomings. These approaches, relying on manually designed operators, suffer from limited feature extraction capacity, thus struggling with complex backgrounds and varying object appearances, constraining the accuracy and robustness.
In contrast, deep learning-based COD method, e.g., convolutional neural network (CNN), transformer, and diffusion model, offer significant advantages by automatically learning rich feature representations . In addition, these methods utilize various strategies to address such challenging task, e.g., aggregating multi-scale features , bio-inspired mechanism-simulation , fusing multi-source information , learning multi-task , jointing SOD , and setting novel-task . Despite their advantages, these methods also face intractable challenges, including high computational demands and the requirement for large annotated, clean, and paired datasets .
Several surveys have been conducted on COD, with three seminal works providing valuable overviews of the field. However, these surveys have limitations due to the narrow scope and limited number of papers they cover. For example, most of the methods discussed in these surveys are from before the first half of 2023, resulting in insufficient historical depth and domain breadth. As illustrated in Fig. 2, the COD field has seen rapid development in 2023. To address these gaps, we propose a more comprehensive survey that not only covers traditional and deep learning COD methods across both image and video domains but also benchmarks deep learning models in these areas. Furthermore, to the best of our knowledge, this survey is the first to deeply explore novel tasks such as referring-based COD and collaborative COD . We also provide a broader review of commonly used COD datasets and comprehensively cover recent advancements, challenges, and future trends.
The motivation for this paper stems from the critical importance of COD and the inadequacies of existing surveys. Our survey aims to provide a more thorough and detailed examination of COD, address gaps in the current literature, and highlight recent developments. We systematically categorize and analyze existing cutting-edge techniques, identify critical challenges, and suggest future research directions to advance the field.
Our contributions are summarized as follows:
We provide a comprehensive review of existing COD methods and related tasks in camouflaged scenario understanding (CSU), along with commonly used datasets and evaluation metrics. To the best of our knowledge, this work represents the most extensive investigation to date, encompassing approximately 180 CSU-relevant cutting-edge studies.
We methodically benchmark 40 representative image-level models and 8 representative video-level models based deep features on 6 characteristic datasets and 6 typical evaluation metrics, providing the quantitative and qualitative analysis of them.
We systematically identify the limitations of existing COD methods and propose potential directions for future research. By shedding light on these challenges and opportunities, our work serves to guide and inspire further research efforts to advance the state-of-the-art in COD technology.
We create a repository that houses a carefully curated collection of COD methods, datasets, and relevant resources, which will be consistently updated to ensure the latest information is accessible.
We hope that this survey on COD will not only enhance understanding of the field but also stimulate greater interest within the computer vision community, fostering further research initiatives in related areas.
Note. In developing our search strategy, we conducted a thorough investigation across a variety of databases, including DBLP, Google Scholar, and ArXiv Sanity Preserver. Our focus is particularly directed toward reputable sources, such as TPAMI and IJCV, as well as prominent conferences like CVPR, ICCV, and ECCV. We prioritized studies that provided official codes to enhance reproducibility, as well as those with higher citations and Github stars, indicative of significant recognition and adoption within the academic community. Following this initial screening, our literature selection process involved a rigorous evaluation of each paper’s novelty, contribution, and significance, and an assessment of its status as seminal work in the field. While we acknowledge the possibility of omitting some noteworthy papers, our aim is to present a comprehensive overview of the most influential and impactful research, promoting research advancement and suggesting potential future trends and directions.
Image-level COD models
Image-level COD refers to the process of identifying and distinguishing objects that are designed to blend into their surroundings 8 within static images, garnering significant attention currently. In this section, we categorize these methods into two primary approaches based on the features they utilize: traditional COD methods and deep learning COD methods. Traditional approaches typically rely on handcrafted features, whereas deep learning methods leverage neural networks to automatically learn and extract discriminative features from data. Given the rapid advancements in technology, we will focus primarily on deep learning COD methods, which have recently become the predominant approach.
COD is initially rooted in traditional image-level methods, leveraging hand-crafted low-level features tailored to capture nuances in textures, intensities, and colors. These approaches constituted the foundation of early efforts in this domain. Tab. I summarizes the relevant models and characteristics of these methods.
Texture feature-based approaches. Texture features capture the surface properties and distinctive patterns of images, usually manifested through grayscale distributions across pixels and their spatial neighborhoods. These features are utilized to distinguish camouflaged objects from their backgrounds based on texture differences.
Galun et al. introduce a bottom-up approach, TSMA, for texture segmentation. This method leverages the adaptive identification and characterization of texture elements–such as size, aspect ratio, orientation, and brightness–combined with filter response statistics for enhanced differentiation and noise reduction. Unlike TSMA, which extracts global texture features, Bhajantri et al. employed a co-occurrence matrix for block-based local texture analysis. Their method uses Watershed segmentation for each defective block, processed through a dendrogram to distinguish the target, and concludes with cluster analysis to further confirm detection results.
Building upon earlier works, Sengottuvelan et al. propose an unsupervised technique for decamouflaging that converts images to grayscale and then splits them into blocks for gray-level co-occurrence matrix analysis, capturing pixel-neighbor relationships. They use a dendrogram plot to identify camouflaged objects without requiring prior background knowledge. Traditional evaluation methods often rely on subjective assessments, which can be cumbersome and ambiguous processes. To address it, Song et al. leverage a weighted structural similarity index and intrinsic image feature analysis to assess camouflage textures. To be more objective, Feng et al. utilize human visual-based saliency maps to quantitatively assess texture differences.
Intensity feature-based approaches. These methods progressively advance from basic intensity-based techniques to more sophisticated exploitation of 3D convexity. CBDCE detected regions of interest by processing intensity images directly, distinguishing between 3D convex and concave regions. This approach demonstrates robustness against variations in illumination, orientation, and scale. To enhance the detection of 3D objects camouflaged in complex scenes, CVCB enhanced the D-Arg operator in CBDCE. It maximizes the response to curved 3D objects against flat backgrounds, effectively mitigating visual camouflage in both natural and artificial environments. CTDDC employed another 3D convexity-based operator that exploits image gray levels and median filtering to eliminate background noise, enabling effective and robust detection.
Color feature-based approaches. In certain scenarios, color contrast and distribution can provide significant distinctiveness for separating camouflaged objects from their surroundings. Siricharoen et al. optimized a statistical background subtraction and shadow detection algorithm by integrating color, edge, and intensity features for outdoor human segmentation, where strong shadows and low contrast are common. They employ vector median filtering to remove outlier pixels and combine color statistics with edge to generate initial coarse results, which are then refined using intensity features for accuracy. Kavitha et al. propose an image retrieval technique for COD, which involves segmenting images into blocks and extracting Hue, Saturation, Value (HSV) color, and gray-level co-occurrence matrix texture features from each block. They use the principle of matching images based on similarity and Euclidean distance to improve detection accuracy.
Handcrafted low-level features are specifically designed to be highly discriminative, making them effective in detecting and segmenting objects. By accentuating differences in texture and color, these features help isolate concealed objects. However, camouflage seeks to minimize these distinctions, reducing visibility and blending objects into the background. Hence, traditional methods fail in COD, succeeding only in simple scenes with uniform backgrounds. These methods tend to underperform with low-resolution images or in cases where the foreground and background share significant visual similarities.
2 Deep learning COD methods
While traditional methods rely on hand-crafted low-level features to capture key attributes of an image, deep learning methods extract complex and deep features directly from the data through automatic learning representations, demonstrating superior performance across various computer vision tasks. According to the existing surveys , deep learning methods for COD can be broadly categorized based on three fundamental criteria: network architecture, learning paradigm, and supervision level. A detailed description of the three criteria is shown in Fig. 3. What’s more, Tab. II and Tab. III outline the key characteristics of a total of 104 representative methods for image-level COD, published in 2019-2022 and 2023&2024, respectively.
Network architecture delineates how the input and output configurations are structured within the models. The linear employs bottom-up/top-down network, where the data flows in a single feed-forward pass. Aggregative architecture combines features from multiple input streams, while branched architecture is characterized by multiple pathways for multiple outputs, associated with the multi-task learning paradigm. The hybrid integrates elements and strategies from the aforementioned ones to leverage their strengths.
The learning paradigm pertains to the approach adopted by the models to learn and adapt. This includes single-task and multi-task learning, specifically the former only involving COD, while the latter often importing auxiliary tasks, e.g., localization/ranking , reconstruction and predicting associated cues, such as boundary , texture and uncertainty , leading to improved accuracy and performance.
The supervision level describes the degree and nature of supervision provided during the training phase. This can be classified into four categories: fully supervised, with complete ground truth data; weakly-supervised, which utilizes limited or imprecise annotations ; semi-supervised, which combines labeled and unlabeled data, generating pseudo labels for the latter ; and unsupervised, where no explicit labels are provided . Notably, the ”train-free” mode refers to models that do not require a traditional training process.
Beyond these categories, existing works also employ several strategies, such as aggregating multi-scale features, simulating bio-inspired mechanisms, fusing multi-source information, learning multiple tasks, jointing SOD, and establishing novel tasks, all aimed at enhancing COD performance. These strategies underscore the diverse and innovative approaches that researchers adopt to address the challenges in COD, revealing the underlying principles and techniques driving advancements. By focusing on these strategies, we can better understand the strengths and limitations of different methods, identify emerging trends, and provide a clearer roadmap for future research. However, there is currently a lack of detailed categorization and analysis of the various strategies employed. Hence, we aim to provide an elaborate introduction to this field. Fig. 4 presents a comprehensive breakdown of the various methods discussed in this paper, along with their respective contributions to the overall analysis as detailed in the following subsections, including 6 types as well as the sub-types of multi-task and multi-source strategies.
Besides, recent COD methods commonly utilize encoder-decoder architectures with the above strategies. The encoder is responsible for capturing rich semantic and structural information from input images, while the decoder reconstructs the image with the desired changes with the discriminative features extracted by the encoder. As shown in Tabs. II and III, they often extract discriminative features via convolution-based encoders, such as VGG , ResNet , Res2Net , and EfficientNet , or transformer-based encoders, such as ViT , CLIP , BLIP2 , PVT , and Swin .
This strategy captures the diverse appearances and varying scales of camouflaged objects with rich context information, and then aggregate cross-level features while gradually refining features , specifically in a hierarchical , residual , dual-branch , X-connection or iterative manner, to enhance the representation. Some are further guided by edge , frequency , querying or coarse prediction maps in fusion process.
ERRNet and HCM propose a reversible re-calibration mechanism that leverages prior prediction maps, specifically targeting low-confidence regions to detect previously missed parts. This approach refines detection by focusing on regions that are initially overlooked. To improve efficiency, TinyCOD introduces an adjacent scale feature fusion strategy using the lightweight TinyNet as the backbone. Inspired by the Transformer architecture , CamoFormer adopts masked separable attention, where multi-head self-attention is divided into three components. This approach allows for the simultaneous refinement of features at different levels in a top-down manner. While vision transformers excel in global context modeling, they often struggle with locality modeling and feature fusion. To address these issues, FSPNet introduces a non-local token enhancement mechanism for improved feature interaction and a feature shrinkage decoder. Building on the Swin transformer and its shifted window strategy, OWinCANet employs overlapped window cross-level attention. This method enhances low-level features with high-level guidance by sliding aligned window pairs across feature maps, ensuring a balance between local and global for superior performance.
Pixel-wise annotation of camouflaged objects is time-consuming and labor-intensive. Weakly supervised methods, using limited or scribble annotations, aim to reduce labeling effort and address boundary ambiguities. CRNet designs a local-context contrasted module to enhance image contrast and a logical semantic relation module to analyze semantic relations, combined with feature-guided and consistency losses to impose stability on the predictions. He et al. utilize the visual foundation model SAM with sparse annotations as prompts to achieve initial coarse segmentation. Enhanced with multi-scale feature grouping, they generate reliable pseudo-labels for training off-the-shelf methods, addressing intrinsic similarity issues in coherent segmentation.
2.2 Mechanism-simulation based strategy
This bio-inspired strategy simulates the behavior of predators in nature or the visual detection mechanisms of humans. This is a multi-stage, coarse-to-fine strategy to incrementally enhance the accuracy of outcomes. SINet simulates the search and identification process of predators, utilizing densely connected features and receptive fields, further enhanced by SINet-V2 with group-reversal attention. Semi-SINet addresses training data scarcity with semi-supervised learning, generating pseudo labels for unlabeled data and introducing edge attention to improve SINet. Such localization-segmentation two-stage process has also inspired many approaches . Besides, extra stages, e.g., restoring , matching and amplifying , have also been introduced for more precise detection and enhanced adaptability.
MirrorNet employs a mirror stream with embedded image flipping as a bio-inspired strategy to disrupt camouflage. ZoomNet emulates human visual patterns when observing ambiguous images, specifically through zooming in and out, and leverages scale integration and hierarchical units to capture mixed-scale semantics. In contrast to multi-scale input strategies, MFFN acquires complementary information through multi-view inputs from various angles, distances, and perspectives. This approach is designed to address the visually indistinguishable characteristics of camouflaged objects that arise under complex conditions, such as tiny object size, fuzzy objects, and blurred boundaries.
Moreover, incorporating auxiliary tasks could contribute to a better understanding of camouflage. PreyNet mimics the sensory and cognitive mechanism of predation combined with the auxiliary task, uncertainty estimation, through a bidirectional interaction module for feature aggregation, as well as a policy-and-calibration paradigm for feature calibration. Camouflageator , inspired by the prey-predator dynamics, presents an adversarial training framework with an auxiliary generator creating challenging camouflaged objects on the prey side. As for the predator side, Camouflageator employs a camouflaged feature coherence module and utilizes edge-guided calibration to enhance complete segmentation and boundary clarity.
2.3 Multi-source information fusion strategy
This strategy integrates diverse supplementary information sources to enhance COD, improving robustness and accuracy by leveraging complementary data from different domains, mainly frequency, depth, or text.
Frequency-domain integrated approaches. FEMNet is the first to extend COD into the spatial domain, incorporating frequency enhancement and high-order relation for effective digging and fusion of the frequency clues with RGB features. FEDER further decomposes the features into different frequency bands via learnable wavelets, using frequency attention and guidance-based aggregation to differentiate foreground-background, complemented by an ordinary differential equation-inspired edge reconstruction for precise boundaries.
Depth-perceptual integrated approaches. DCE introduce the first depth-guided COD network, which incorporates an auxiliary depth estimation branch and a multi-modal confidence-aware loss function via GAN for effective depth integration. Building on DCE, DaCOD also utilizes existing monocular depth estimation methods to generate depth maps and proposes a novel cross-modal asymmetric fusion strategy to blend these modalities asymmetrically. XMSNet implements all-around attentive fusion to facilitate explicit cross-modal semantic mining and consistency constraints across decoding layers, thereby mitigating bias from inherent noise in depth estimation. PopNet employs source-free depth for object pop-out, identifying contact surfaces under weak supervision and leveraging 3D priors for segmentation without source data, which enables efficient depth-to-semantics transfer. RISNet , designed for extreme agricultural application scenarios, uses multi-scale receptive fields and depth feature fusion to enhance dense and small COD in multiple stages.
Prompt-learning integrated approaches. These approaches utilize textual or visual prompts to adaptively guide and enhance COD models. CoVP introduces a chain of visual perception and language text prompts, which linguistically and visually enhance the camouflaged scene perception of large vision-language models (LVLM), thereby reducing hallucinations. In contrast, GenSAM employs cross-modal chains of thought prompting to generate visual prompts and progressively produce masks, iteratively refining the detection results. Both CoVP and GenSAM operate in a training-free manner, leveraging pre-trained models to avoid the need for extensive retraining, save computational resources, and enable rapid adaptation to COD. Rather than relying on text prompts, VSCode introduces 2D domain-specific and task-specific prompt learning, effectively disentangling domain and task peculiarities. Through joint training, VSCode emerges as the first generalist model capable of addressing multimodal SOD and COD tasks with remarkable zero-shot generalization ability.
2.4 Multi-task learning strategy
The fundamental premise of multi-task learning is that different tasks share common information for data processing. By leveraging these additional cues, multi-task learning is extensively employed to extract complementary insights from positively correlated tasks—such as reconstruction, classification, and localization—to enhance the detection accuracy of camouflaged objects.
Boundary-supervision integrated approaches. MGL integrates camouflaged object-aware edge extraction (COEE) with graph-based mutual learning, incorporating typed functions to enhance feature representation by capturing both semantic and spatial information, which benefits COD and boundary detail accuracy. To address interpolation accuracy loss and computational redundancy in MGL, MGL-V2 incorporates multi-source attention contextual recovery, iteratively leveraging pixel feature information. Inspired by human COD processes, BSA-Net employs a two-stream separated attention mechanism—reverse and normal attention streams—followed by a boundary guider to enhance performance by focusing on object boundaries. Beyond boundary enhancement, the extension of BSA-Net, FindNet , embeds texture information into feature representation, focusing on local patterns to perform effectively under complex COD conditions. To better integrate multiple features, ASBI introduces an attention-induced semantic and boundary interaction network, utilizing attention-induced interaction to completely fuse multivariate and heterogeneous information.
Category-prediction integrated approaches. ANet is the very classic method for COD, which includes a classification stream to predict whether they contain camouflaged objects and a segmentation stream to recognize them. To leverage the strong generalization ability and rich semantic knowledge of large-scale pre-trained foundation models, PAD utilizes the pre-trained ViT with lightweight parallel adapters by only tuning a small number of parameters for multi-task learning in a “pre-train, adapt, and detect” paradigm, achieving zero-shot task transferability, multi-task adaption, and cross-task generalization. As unseen classes are more general in real-world scenarios, Li et al. propose ZSCOD for zero-shot learning, which employs dynamic graph searching to adaptively capture edge details and a camouflaged visual reasoning generator for generating pseudo-features with the object-wise graph learning strategy to dynamically sample nodes to reduce background interference. Considering the limitation of data on COD, FS-CDIS leverages few-shot learning, simultaneously implementing instance segmentation task, with proposed instance triplet loss and instance memory storage, enhancing distinguishable features between background and foreground areas.
Localization/ranking integrated approaches. LSR introduces the first ranking-based COD network, which infers the detectability of different camouflaged objects through instance segmentation and classification branches. LSR-V2 builds on the pioneering work of LSR by further exploring the interdependencies among these tasks. It not only establishes a new baseline and benchmark for both individual and joint tasks within a triple-task learning framework but also revisits the role of the camouflaged object ranking (COR) task.
Reconstruction integrated approaches. FRINet uses Laplacian pyramid-like decomposition and transformer-CNN hybrid encoders with a reasoning module to capture and integrate high and low-frequency components, guided by an auxiliary image reconstruction task. Considering the strong performance of the transformer, Hao et al. uses ViT and regards both image reconstruction and binary segmentation tasks as training targets, with a local information capture module and dynamic weighted loss to enhance local modeling and handle complex cases.
Texture-detection integrated approaches. Motivated by the need to leverage complementary texture and camouflaged object cues, Zhu et al. utilizes texture labels and an interactive guidance framework. This framework consists of feature interaction guidance as well as texture and holistic perception decoders, aimed at refining segmentation with a particular focus on indefinite boundaries and texture differences. To better exploit the discriminative patterns within the objects, DGNet decouples COD into context and texture branches based on gradient generation. It employs a gradient-induced transition to softly group features from both branches, resulting in superior efficiency.
Uncertainty-estimation based approaches. By approximating the uncertainty across different areas, these methods enhance the model’s ability to focus on less identifiable regions, ultimately achieving high-confidence detection. UR-COD utilizes uncertainty-aware refinement to reduce the noise of pseudo-edge and pseudo-map labels. Yang et al. introduce Bayesian learning into transformer-based reasoning which revolutionizes the traditional deterministic mapping process employed in conventional COD by transitioning it into an uncertainty-guided context reasoning procedure. OCENet employs aleatoric uncertainty estimation for confidence-aware COD, using dynamic supervision for accurate maps and assessing pixel-wise accuracy without ground truth. PUENet integrates model and data uncertainty, implementing predictive uncertainty estimation and predictive uncertainty approximation for efficient test-time alongside SAM refining hierarchical features.
2.5 Joint-SOD based strategy
SOD and COD seem opposing but share common ground in the necessity to discern objects from a background based on contrast and contextual features. This strategy combines SOD with COD, leveraging the contradictory information or shared characteristics to gain more thorough comprehension of COD. Interestingly, many works are built in a multi-task manner, e.g., boundary detection and image reconstruction .
UJSCOD uses uncertainty-aware adversarial learning with a similarity measure module for modeling contradicting attributes and a data interaction strategy by defining simple COD samples as hard SOD samples for SOD data augmentation and higher robustness. Additionally, UJSCOD-V2 introduces contrastive learning to further investigate the cross-task correlations and random sampling-based foreground-cropping for COD data augmentation. CMNet proposes a new perspective, de-camouflaging, modeling task-conflicting and task-consistent attributes to destroy the camouflage.
2.6 Novel-task setting strategy
As COD continues to gain attention, its application scenarios have become increasingly diverse. Consequently, researchers have proposed various novel task settings to address COD challenges in different contexts, offering significant advantages, such as improved generalization to unseen data, better handling of complex scenes, and broadening the applicability of COD technologies.
Unsupervised camouflaged object segmentation (UCOS). Unsupervised learning is necessary due to the challenges in gaining extensive human labels for open-world applications, where supervised models often exhibit poor generalization. Domain adaptation is crucial because it allows models to effectively transfer knowledge from a source domain to a target domain. This adaptation helps bridge the gap between different data distributions, enhancing the model’s performance in real-world scenarios where data characteristics may vary significantly. Therefore, a novel task, termed UCOS-DA , is formulated for unsupervised COD where both source and target labels are absent in the training phase. UCOS-DA leveraging foreground-background contrastive self-adversarial for pseudo-labels and domain-specific adaptation to address the UCOS task.
Collaborative camouflaged object detection (CoCOD). Zhang et al. introduce this novel task to address the limitations of single-image analysis by jointly segmenting the same camouflaged object or objects belonging to the same class across multiple distinct images. This strategy leverages shared similarities and complementary cues inherent in the related images, thereby enhancing the accuracy of COD. To tackle CoCOD, BBNet extracts and integrates camouflaged object features through inter-image collaborative feature exploration and intra-image object feature search. Additionally, BBNet enhances the representation of co-camouflaged features by employing the strategy of local-global feature refinement.
Referring camouflaged object detection (RefCOD). Standard COD aims to detect all camouflaged objects in a given scene. However, in certain real-world scenarios, such as ecological species protection and discovery, explorers may be interested only in locating specific camouflaged objects. To address this need, target references are introduced into COD, where referring text or images containing salient targets guide the detection of specified camouflaged objects. R2CNet introduces reference and segmentation branches, combined with referring mask generation, to create pixel-level priors and enrich referring features, thereby enhancing detection capabilities. In contrast, MLKG uses text as a reference in a multi-level knowledge-guided multimodal approach. By leveraging multimodal large language models (MLLMs), it achieves deep alignment, improves performance, and enables zero-shot generalization on unimodal datasets.
Open-vocabulary camouflaged object segmentation (OVCOS). Open-vocabulary refers to models that recognize and segment objects from novel classes not seen during training, leveraging vision-language models(VLM) like CLIP . To fill in the gaps for COD in this field, Pang et al. introduce the novel task OVCOS and built the corresponding baseline OVCoser, which integrates semantic guidance and visual structure cues in an iterative refinement manner via a transformer-based architecture attached to a frozen CLIP . It also incorporates diverse sources of information, such as class semantic cues, spatial depth structures, object edge details, and iterative top-down guidance from the output space. To enhance task-relevant semantic context, OVCoser employs carefully crafted prompt templates, achieving robust and generalized performance.
Video-level COD models
Unlike image-level Camouflaged Object Detection (COD) techniques, which focus on single static images, video-level COD requires greater emphasis on motion cues to identify and localize camouflaged objects within continuous video frames. Video COD (VCOD) typically leverages temporal information, such as motion and changes across frames, to reveal objects that are difficult to detect in individual frames. However, this task presents significant challenges, including complex background noise, lighting variations, occlusions, and diverse camouflage strategies. Additionally, the high-dimensional nature of video data necessitates algorithms that are not only spatially accurate but also temporally consistent and stable. In this section, we categorize existing methods into two main types: traditional approaches and deep learning-based approaches.
Compared to image-level methods, the video-level ones include more techniques, like optical flow analysis and motion detection, making them practical in video surveillance and security. As illustrated in Tab. IV, this section will delve into 12 representative traditional VCOD models, and we also categorize and introduce these methods based on the types of features they rely on.
Texture feature-based approaches. Malathi et al. employ multi-camera codebooks for the detection of foreground objects. In their approach, texture pixels resembling the background are extracted and quantized into distinct codebooks. These codebooks are then integrated into a weighted framework to guide the abstraction of the foreground. To enhance the accuracy of predictions, disparity maps generated from the codebooks are used as supplementary information, particularly in cases where the color contrast between the foreground and background is weak. However, there remains the potential for shadows to be misclassified as targets.
Intensity feature-based approaches. When applied to video, these methods track dynamic changes between frames by leveraging temporal information, thereby improving the detection and localization of moving camouflaged objects. To tackle the challenge of foreground-background segmentation in video surveillance, particularly when complicated by camouflage, Guo et al. combine Bayesian classification with Gaussian mixture models and apply temporal averaging across multiple frames in video sequences to reduce the bias of background models. However, detecting subtle differences remains difficult when the foreground and background are highly similar. To address this issue, Li et al. extend COD to the wavelet domain, where subtle differences are emphasized in specific wavelet bands. They apply wavelet transforms to video sequences and estimate the probability of a wavelet coefficient belonging to the foreground by constructing foreground and background models within each individual wavelet band. This approach detects camouflaged moving foregrounds through wavelet-domain multi-scale fusion.
Color feature-based approaches. Zhang et al. employs computer-assisted detection of camouflaged targets, which first globally models the background of the input image, with the modeling of the foreground divided into global and local models. The global model captures the overall color and texture information of the foreground, while the local model focuses on details or changes that may exist in the foreground. Based on the models of the background and foreground, a factor measuring the degree of camouflage is introduced, determining whether a pixel is camouflaged by comparing the color differences between the background and foreground at that pixel. By comparing color contrast measurement results, true camouflaged regions are identified. Finally, the camouflage and identification models are fused under a Bayesian framework to perform complete target detection.
Motion feature-based approaches. These methods leverage the relative movement between the target object and the background between consecutive frames to identify camouflaged objects. Boult et al. propose a method for distinguishing background and foreground objects in visual surveillance systems using motion features. They introduce a conditional incremental model to update multiple background models in real-time and employed quasi-connected components to fill gaps caused by slight object movements, adapting to scene changes. However, background subtraction under camouflage conditions often results in fragmented, discontinuous object pieces, complicating subsequent classification or tracking processes. To address this, Conte et al. developed an algorithm for detecting camouflaged personnel by aggregating fragmented detection blocks to reconstruct the complete shape of the target object during post-processing. This algorithm relies on the consistent segmentation of parts of the human body, such as the head, torso, and legs, across video sequences. By defining specific parameters to model the bounding boxes of human figures, the algorithm merges two or more boxes according to a set of rules. Considering perspective effects, a semi-automatic calibration phase dynamically adjusts parameters to ensure the bounding boxes accurately reflect the actual size of the individuals.
Optical flow feature-based approaches. Optical flow represents the motion of objects in a sequence as a vector field, providing crucial information on object speed and direction. To address challenges in classifying and identifying objects from low spatial resolution images—particularly in security-related applications—Beiderman et al. introduce a technique that uses spatially coherent light beams to illuminate scenes and capture secondary speckle patterns from reflections. By tracking the temporal variations of these speckle patterns, the method extracts temporal feature signatures of objects. Comparing these signatures allows the algorithm to differentiate camouflaged objects from their surroundings, leveraging the unique physical properties of object surfaces for effective detection even under heavy camouflage. Building on this approach, Yin et al. propose a different method for detecting camouflaged objects in dynamic backgrounds using optical flow techniques. By simulating the motion patterns of objects and backgrounds and employing clustering analysis of optical flow features, this method accurately identifies objects. The algorithm further optimizes detection results through Kalman filtering, enhancing performance and accuracy while adapting to dynamic environments. Acknowledging the high data dependency of supervised methods and the difficulty of obtaining high-quality data in practical applications, Kim proposes an unsupervised method using a visible-near-infrared hyperspectral camera. This method selects spectral and spatial features online through statistical distance to generate candidate feature bands, followed by entropy-based analysis to remove ineffective features. This approach achieves precise detection of camouflaged objects while reducing computational load.
2 Deep learning VCOD methods
Compared to traditional VCOD techniques, deep learning methods verify a significant advantage by automatically learning complex feature representations from large datasets. These methods have a unique ability to capture intricate and subtle patterns, which in turn enhances the understanding of object dynamics within video sequences. However, video data presents greater complexity compared to image-based approaches. This added complexity stems from factors such as higher data dimensions, temporal continuity across frames, and the constantly evolving nature of dynamic changes in video content. These factors require models to possess not only strong spatial feature extraction capabilities but also a robust mechanism for accurately capturing subtle temporal variations. Despite promising advancements, existing work has largely focused on leveraging motion cues between different frames, and these approaches remain in their early stages of development, with considerable room for growth before they reach maturity.
In recent years, researchers have increasingly adopted a two-step framework, where optical flow maps or pseudo masks are pre-generated to serve as motion cues for video object detection. However, due to the challenges associated with cumulative errors and weak generalization , there is a growing trend towards employing end-to-end universal models to enhance reliability. The key characteristics of a total of 16 representative methods for VCOD are detailed in Tab. V. We categorize these methods based on various criteria, e.g., network architecture, the use of optical flow maps, supervision level, and synthetic dataset generation. Besides, some methods provide links to their open-source projects.
Since motion cues are crucial for distinguishing moving camouflaged objects from their backgrounds, researchers often incorporate a stage to pre-generate optical flow maps or pseudo masks for extracting motion information. Optical flow maps directly capture motion fields, i.e., the optical flow, between consecutive frames, providing compensation or registration to mitigate the camouflage effect. In contrast, pseudo masks implicitly learn temporal correspondences and motion patterns within the network, reducing reliance on external optical flow estimation and enabling the network to capture motion and maintain temporal consistency.
Explicit motion-based methods. Lamdouar et al. adopt optical flow and a different image as inputs, and propose a differentiable registration module for background alignment and a motion segmentation module with memory for moving object discovery. Facing the challenge of massive human annotation, Yang et al. introduce a self-supervised method without any manual supervision, which groups motion with similar optical flow according to perceptual grouping principles. Meunier et al. utilize unsupervised CNN-based motion segmentation from optical flow, leveraging the Expectation-Maximization framework for loss design and training, as well as designing data augmentation on the optical flow field, enabling real-time segmentation without annotations or iterative motion model estimation. Xie et al. introduce object-centric layered representation and generate synthetic data for multi-object segmentation and tracking. However, reliance on external optical flow estimation can introduce errors that accumulate over time, potentially compromising the final mask prediction.
Implicit motion-based methods. SLT-Net leverages short and long-term spatiotemporal relationships, specifically, utilizing the short-term motion capture between consecutive frames to produce pseudo masks and long-term temporal consistency to refine the former predictions, mitigating the flow estimation error. However, such implicit modeling may suffer from limited VCOD data.
2.2 End-to-end VCOD framework
This framework emerges as a solution to address the limitations of two-step approaches, particularly the cumulative errors stemming from intermediate stages and the weak generalization ability caused by limited training data. By integrating feature extraction, motion modeling, and segmentation into a unified network, this framework aims to minimize error propagation and maximize the utilization of scarce data, thereby enhancing robustness and performance. To extend the previous version which zooms in and out on images with a shared triplet feature encoder, ZoomNeXt implements image-video unified framework, and furtherly integrates multi-head scale integration and rich granularity perception, enhancing the structural representation and discrimination. IMEX unifies implicit and explicit motion learning in a cohesive framework, achieving inter-frame alignment and consistency preserving of camouflaged objects respectively. TSP-SAM and EMIP are both prompt learning-based methods for VCOD. The difference is that TSP-SAM utilizes frozen SAM embedded with temporal-spatial injection, motion-driven self-prompt learning and long-range consistency to learn reliable visual prompts, while EMIP adopts two-stream architecture, incorporating segmentation-to-motion and motion-to-segmentation prompts. Notably, although EMIP handles motion cues explicitly, it is not a two-step method, as it simultaneously conducts optical flow estimation and VCOD by interactive prompting. These methods demonstrate the evolving trend towards holistic, unified frameworks that not only address the shortcomings of traditional two-step approaches but also push the boundaries of performance and efficiency in VCOD.
Other camouflaged scenario tasks
Beyond the fundamental tasks of COD and VCOD, the domain of concealed scene understanding has evolved to encompass a wider array of high-level semantic tasks. These tasks aim to provide a deeper comprehension of camouflaged objects, extending from their classification to the generation of new camouflaged images. This section delves into the following advanced tasks, each addressing unique aspects of concealed scene understanding. The detail descriptions are illustrated in Fig. 6.
Camouflaged objects classification (COCls) focuses on distinguishing between different types of camouflaged objects. This task involves categorizing camouflaged objects into predefined or never seen classes, i.e., zero-shot and open-vocabulary learning, to handle the subtle differences and similarities among various camouflaged entities. Many COD methods integrate COCls , training with datasets that are labeled with specific categories , for better performance. Accurate classification of camouflaged objects can aid in better understanding and various ecological studies.
Camouflaged objects localization (COL) aims to identify the most detectable regions of camouflaged objects, and further localize the discriminative regions that make the camouflaged object stand out. Lv et al. is the first to propose this task and leverage an eye tracker to record human gaze patterns, pinpointing salient discriminative regions. Simultaneously, they relabel existing datasets with fixation annotation, providing a valuable resource for training and evaluating models. This task enhances the ability to pinpoint critical areas within a camouflaged scene, which is crucial for applications in wildlife monitoring, and search and rescue operations.
Camouflaged instance count (COCnt) is focused on quantifying the number of camouflaged objects within a given scene, even in complex environments where objects may overlap or partially obscure each other. Sun et al. introduce a correlated task, indiscernible object counting (IOC), to count objects that blend seamlessly with their surroundings. To tackle this challenge, they propose a unified framework, IOCFormer, integrating density-based and regression-based counting methods. Due to the scarcity of suitable datasets, they also created IOCfish5K, full of high-resolution images for underwater IOC, with dense annotations. This emerging and promising task helps in assessing population densities and ensuring comprehensive area surveillance.
Camouflaged instance rank (CIR) addresses the challenge of ranking multiple camouflaged instances based on specific criteria such as visibility, detectability, or relevance. With the introduction of the COL task, Lv et al. propose the CIR task, which is based on the difficulty level of camouflage. To this end, the proposed CAM-FR and CAM-LDR datasets also include ranking labels. Furthermore, they devise a triple-task learning framework that simultaneously localizes, segments and ranks camouflaged objects, efficiently utilizing the inner correlation among COD, COL, and CIR.
Camouflaged instance segmentation (CIS) is a more granular task that involves segmenting individual camouflaged objects, i.e., instances, from their background and from each other. CIS is proposed by Le et al. , who also introduce CIS dataset, i.e., CAMO++, by extending the previous CAMO dataset. They also conduct camouflage fusion learning, which fuses existing instance segmentation models, e.g., Cascade Mask RCNN , by learning to predict the best model per image. To better break the deceptive camouflage, Luo et al. propose a framework with a pixel-level camouflage decoupling module and an instance-level camouflage suppression module. Pei et al. introduce the first one-stage transformer in CIS by fusing local features and long-range context dependencies. Recently, Vu et al. leverage text-to-image diffusion and CLIP for CIS, while utilizing open-vocabulary capabilities to learn multi-scale textual-visual features. This level of detail allows for more precise identification and interaction with each object, which is essential for applications that require accurate object differentiation, such as advanced robotic vision, detailed ecological studies, and targeted medical imaging.
Experiments
Datasets play a pivotal role in the development and evaluation of COD algorithms. In this subsection, we provide an overview of prominent datasets relevant to COD tasks, categorized into image-level and video-level datasets. Tab. VI summarizes essential information along with their respective links for access. Additionally, Fig. 7 and Fig. 8 showcases exemplar images from these datasets, offering a visual insight into the challenges posed by COD.
CHAMELEON is a small-scale, unpeer-reviewed dataset consisting of 76 camouflaged images collected from the internet using the keyword “camouflaged animals”. Each image is manually annotated with labels focusing on animals camouflaged within complex ecological backgrounds. This dataset is typically used as a test dataset for model performance evaluation.
CAMO-COCO consists of 2,500 images across eight categories, created by merging the camouflaged dataset CAMO with the non-camouflaged dataset MS-COCO, each contributing 1,250 images. In CAMO-COCO, 80% of the images are designated for training and the remaining for testing. CAMO includes both natural camouflaged objects, such as animals, and artificial camouflaged objects, such as human beings, featuring seven challenging attributes that complicate detection and segmentation.
NC4K, the largest image-level COD test dataset currently available, contains 4,121 camouflaged scene images sourced from the internet, each annotated at both object and instance levels. The dataset covers predominantly natural scenes along with some artificial camouflage, making it a preferred choice for evaluating the generalization capabilities of COD models.
COD10K includes 10 superclasses and 78 subclasses, totaling 10,000 images, with 60% designated as training data. Within this largest image-level COD dataset to date, there are 5,066 camouflaged images (3,040 for training and 2,026 for testing), 1,934 non-camouflaged images, and 3,000 background images. All camouflaged images are densely annotated with categories, bounding boxes, and object and instance level labels, supporting a wide range of research tasks, e.g., COD, COS, and CIS. The high-quality annotations and diversity of COD10K make it an essential dataset for COD research.
S-COD is the first dataset created for weakly supervised learning based on scribble annotations. It uses a tagging process that relies on initial impressions to outline the rough structure of objects, including both foreground and background. The dataset comprises 3,040 images from the COD10K training set and 1,000 images from the CAMO training set, totaling 4,040 samples. Compared to pixel-level annotations, the annotations in S-COD are simpler and more efficient.
CAM-LDR facilitates research on COL and CIR by recording the time taken to detect camouflaged instances with an eye tracker, which correlates with the difficulty of detection. This dataset includes 4,040 training images, derived from the CAMO and COD10K training sets, and 2,026 testing images from the COD10K testing set. Detection times are used to rank camouflaged objects into six categories: background, easy, three medium levels, and hard. This approach offers a novel metric for understanding the detectability of camouflaged objects.
CoCOD8K is the first dataset for CoCOD, which comprises 8,528 images reorganized from four COD datasets—CHAMELEON, CAMO, COD10K, and NC4K. This dataset includes diverse natural and artificial scenes, classified into 5 superclasses and 70 subclasses, each annotated with object masks and category labels. Designed to foster research in detecting co-camouflaged objects across grouped images, CoCOD8K also filters images to fit specific criteria, supporting robust model training with 5,933 training and 2,595 test images.
R2C7K encompasses 6,615 images across 64 categories from real-world scenarios to facilitate RefCOD. It consists of a Camo-subset with 5,015 camouflaged images, primarily derived from COD10K, and a Ref-subset with 1,600 images of salient objects, uniformly sourced with 25 per category from Flickr and Unsplash webpages with no copyright disputes. For research, a referring split is provided where each category in the Ref-subset has 20 images for training and 5 for testing, while the distribution in the Camo-subset follows the original split from COD10K with additional samples from NC4K to ensure at least 6 samples in each category.
OVCamo advances the task of OVCOS by offering a robust dataset featuring 11,483 images across 75 object classes, derived from merged public datasets such as . This dataset uniquely tackles semantic ambiguities by redefining annotation standards, which ensures clear and distinct class definitions to improve segmentation accuracy. For realistic performance evaluation, OVCamo divides its dataset by allocating 14 seen classes to the training set, while the remaining 61 unseen classes are reserved for testing. This distribution maintains a training-to-testing sample ratio of 7:3, simulating real-world application challenges and ensuring a rigorous assessment of model generalization.
In current research practices, the most representative datasets typically selected for experiments of image-level COD models are CHAMELEON, CAMO-COCO, COD10K, and NC4K. The common setup for training includes 1,000 images from CAMO and 3,040 images from COD10K. The remaining parts of these datasets are used to test the generalization ability and viability of the models. To maintain consistency in our analysis, we will adopt this setup in our subsequent performance comparison of image-level COD models.
1.2 Characteristic VCOD Datasets
CAD2016 is composed of nine short video sequences sourced from YouTube, total having 836 frames. Each sequence is manually annotated every five frames. The camouflaged objects in these images are exclusively biological entities found in natural settings.
MoCA is currently the largest dataset for camouflaged animal detection in video format. It consists of 141 video sequences, also sourced from YouTube, representing 67 different categories of animals found in natural scenarios. The dataset spans over 37250 frames and 26 minutes of video content. Annotations include PWC-Net optical flow data for each frame and bounding boxes with motion labels provided every five frames, with linear interpolation used for the intervening frames.
MoCA-Mask is an extension of MoCA and includes 87 video sequences with a total of 22,939 frames, following the removal of irrelevant scenes. This dataset enhances MoCA by providing manually annotated masks every five frames, resulting in 4,691 bounding boxes and pixel-level masks. The dataset is divided into a training set that comprises 71 videos (19,313 frames), along with a testing set consisting of 16 videos (3,625 frames).
2 Evaluation metrics
We evaluate existing COD models using four common evaluation metrics as recommended in. These metrics, i.e., S-measure (), F-measure (), Mean Absolute Error (MAE), and E-measure (), provide a comprehensive assessment of performance. Here, we detail these evaluation metrics:
Precision-Recall (PR) curve is generated by transforming the input into a binary mask , which is segmented across a range of thresholds from 0 to 255. Precision () and Recall () are calculated by comparing the binary mask with the ground truth mask () at each threshold, producing the PR curve. and can be calculated as
where is the mask obtained by thresholding the non-binary prediction map at threshold .
S-measure () quantifies the spatial structural similarity between the predicted map () and the ground truth (). It combines both object-aware () and region-aware () assessments with the following definition:
where is a weighting factor, typically set at 0.5, that balances the contribution of and .
F-measure () is used to calculate the relationship between Precision () and Recall (). Initially, the input is transformed into a binary mask, , segmented over a range of thresholds from 0 to 255. and are calculated by comparing with across these thresholds. The formulas for and are defined as follows:
where represents the binary mask obtained by thresholding the non-binary prediction map at threshold , and denotes the total area of the mask. But further demonstrates the average harmonic mean value between them. The formula for is defined as
with typically set to 0.3. From the range of thresholds, three variants of are computed: the maximum (), the mean () and the adaptive (). Besides, () is also a widely used metric where the and are both weighted averages. is adopted in this paper.
E-measure () evaluates both the local and global similarity between and . It is defined as follows:
where represents an enhanced alignment matrix, and and are the width and height of the input image, respectively. also provides three indicative values: maximum (), mean (), and adaptive (). is adopted for evaluation in this survey.
Mean Absolute Error () quantifies the average absolute difference per pixel between the normalized predicted map and the ground truth map , where . The mean absolute error is formulated as follows:
where and denote the width and height of the input image, and represents the pixel coordinates. Unlike , and , a lower suggests a more accurate model.
For VCOD, mean Dice (mDice) for similarity evaluation and mean IoU (mIoU) for overlap measurement are also used for evaluation. Notice that larger mDice and mIoU scores indicate better performance.
3 Quantitative analysis
Results of deep COD models. In this section, we conduct experiments on 39 cutting-edge techniques, covering six strategies mentioned earlier, and categorize them based on their backbones into two types: convolution-based and transformer-based. As shown in Tab.VII.
PopNet stands out as a top performer, achieving the highest and scores across multiple datasets. Its success is attributed to its innovative integration of source-free depth estimation and object-popping techniques, followed by the precise separation of objects from their contact surfaces. Overall, methods utilizing transformer-based backbones, such as FSNet and HitNet , exhibit superior performance compared to those using convolution-based backbones. The self-attention mechanism inherent in transformers is highly effective in modeling long-range dependencies, which are crucial for accurate camouflaged object detection (COD). Furthermore, the results reveal that certain strategies are particularly effective for specific datasets. For example, methods incorporating multi-scale context information, such as C2F-Net-V2 , tend to perform better on datasets with varying object sizes, such as COD10K-test. Similarly, mechanism simulation strategies, such as LSR-V2 , are beneficial for datasets with complex backgrounds, like CAMO-test. Interestingly, GenSAM , despite incorporating advanced models like CLIP and BLIP2, does not consistently outperform other methods. This underscores that while powerful pre-trained models can be beneficial, their effectiveness also depends on their integration into the overall COD framework. Additionally, models designed for more challenging settings, such as UCOS-DA for UCOS and MLKG for RefCOD, do not always achieve outstanding performance. These models require more sophisticated approaches and larger datasets to generalize effectively. Consequently, some researchers are developing new datasets to train proposed models, such as R2CNet , OVCoser , and BBNet .
Results of deep VCOD models. We compare 17 cutting-edge methods across two camouflaged video datasets, as detailed in Tab. VIII. This includes 7 image-based methods, 2 related video-oriented object segmentation methods, and 8 deep VCOD methods.
ZoomNeXt stands out as the most effective VCOD approach, achieving state-of-the-art results in , and . This is mainly due to the zoom-in-and-out operation and scale integration for robust feature fusion and efficient temporal modeling in a unified framework. Video-based methods overall outperform their image-based counterparts across several key metrics, which is attributed to their ability to exploit temporal consistency across frames, which helps distinguish camouflaged objects from their backgrounds even when they are moving at a high speed. The end-to-end frameworks, e.g., TMNet , IMEX and EMIP , demonstrate the efficacy of integrating temporal modeling directly into the network architecture. In contrast, two-stage frameworks, like SLT-Net and MG , perform slightly worse, which follow an explicit or implicit motion-based approach, first estimating optical flow or pseudo masks before performing COD. While effective, this decoupled strategy may introduce errors that propagate through the pipeline. What’s more, based on the powerful SAM, TSP-SAM and SAM-PM also perform well, which showcases the remarkable spatial segmentation versatility and potential of SAM when applied to the challenging task of VCOD.
4 Qualitative analysis
As depicted in Fig. 9, we present a comprehensive visual comparison of 10 cutting-edge image-level COD methods. We selected various challenging camouflaged images across 7 typical complex scenarios, including background matching, variable shape, multi-object environments, degraded scenarios, tiny objects, blurred boundaries, and severe occlusion. Additional, more challenging scenarios for COD are discussed in Section 6, and typical examples can also be seen in Fig. 10.
In the background-matching scenario, where insects blend seamlessly with their surroundings, we observe that SINet struggles to identify insects, often resulting in missed detections. Conversely, SegMaR demonstrates robustness by effectively outlining insect contours, even under low-contrast conditions. In the variable shape scenario, such as with the leafy sea dragon, whose appendages resemble leaves, the adaptability of models is tested. While ZoomNet misidentifies parts of these large appendages as separate objects, FPNet exhibits a superior ability to segment more complete instances, illustrating its robustness to shape variations.
Under multi-object configurations, where the scene is crowded with numerous camouflaged entities, the models generally succeed in locating all targets but struggle with accurate segmentation. However, PopNet performs commendably, achieving better coverage without excessive false positives. Degraded scenarios, including poor lighting or blurring, significantly impact the localization and identification processes of the models. The detection results across all models fall short of expectations, indicating considerable room for improvement in these challenging conditions.
For tiny objects, prediction precision degrades, with most methods failing to detect them accurately. However, ZoomNet, with its zoom-in-and-out operation, shows potential in highlighting even the smallest targets. In blurred boundary scenarios, defining precise contours becomes challenging. Here, FEDER stands out with its ODE-inspired edge reconstruction for complete edge prediction, whereas other methods often produce incomplete boundary predictions. Lastly, in severe occlusion scenarios, where two insects partially obscure each other, the results from all methods are generally acceptable. Nonetheless, BGNet displays an improved capacity through its edge-guidance feature module, though further enhancement in detail detection is needed.
In conclusion, while all cutting-edge methods demonstrate acceptable performance across various challenging COD scenarios, there are still notable limitations, particularly in degraded conditions, tiny objects, and blurred boundaries. Further research focusing on enhancing robustness and precision under these extreme conditions is crucial for advancing the field.
Future directions
Deep generative models for data scarcity. To mitigate the scarcity of data, leveraging deep generative models to synthesize diverse, realistic camouflaged images will bolster training effectiveness by dataset augmentation , enhancing model robustness in dealing with camouflaged scenarios. With the rise of image generation models represented by GANs and diffusion models, diverse and high-quality images can now be controlled and generated using other multimodal inputs such as text. This advancement can effectively address the longstanding issue of dataset scarcity in this field. By training with both generated camouflaged object data and the original datasets, performance can be significantly improved. Adopting an adversarial manner, where the camouflaged sample generator and the segmentation network are pitted against each other, may help unlock further potential . However, there are no quantitative metrics to evaluate the camouflage degree and sample quality in deep datasets, making the widespread use of generated samples for training a topic of debate. For VCOD datasets, the generation of such data still has a long way to go due to the current immaturity and lack of a general framework in video generation technology. Diffusion models have become a research hotspot in the field of computer vision due to their stability in algorithm training and high-quality sample generation. CamDiff innovates by synthesizing salient objects within camouflage scenes, mitigating the scarcity of multi-pattern training data and enhancing robustness to salient misclassifications. The authors also use CamDiff to propose Diff-COD dataset from the original COD datasets to enhance the robustness to saliency. Meanwhile, LAKE-RED tackles the limitations of dataset diversity and expensive data collection by automatically generating camouflage images without manual background specification, fostering scalability and interpretability. Both approaches emphasize the potential of diffusion models to address data-scarce vision tasks.
Tackling complex scenarios & challenging samples. Addressing the intricacies of COD requires grappling with various challenges posed by complex scenes and difficult samples. Key issues include extremely complex backgrounds and advanced camouflage techniques that render objects nearly invisible. Fig. 10 depicts some extremely concealed scenarios with challenging samples. Multi-object, multi-scale scenarios present difficulties in capturing objects of varying sizes and scales, while severely occluded objects demand robust algorithms capable of disentangling overlapping structures. Objects with fuzzy appearances, particularly small ones, challenge the limits of detection frameworks due to their minimal visual cues. Disruptive patterns, transparency, and background matching further complicate detection by seamlessly blending objects into their surroundings. Blurred boundaries and structural ambiguity add to the difficulty of delineating object edges, while variable shapes require flexible recognition capabilities. Moreover, detecting extremely rare camouflaged objects necessitates specialized methods to distinguish them from the vast majority of non-camouflaged instances. Objects concealed in darkness, affected by drastic illumination variations, pose unique challenges that require innovative approaches to handle varying light conditions. Overcoming these obstacles demands the development of advanced detection models capable of adapting to diverse and extreme scenarios, ensuring continued progress and real-world applicability of COD research. Additionally, performance degradation in challenging environments, such as low-light conditions and foggy settings, exacerbates the difficulty of detecting camouflaged objects . Low light intensifies contrast issues, making camouflaged objects nearly invisible against their surroundings, while fog introduces noise and occlusion, further complicating feature extraction and localization. Enhancing robustness in such extreme conditions requires innovations in image enhancement techniques , coupled with adaptive feature learning strategies specifically tailored for degraded images. In actual application scenarios, the above problems are very common, for example, in the agricultural domain, detecting numerous small and concealed crops amidst severe occlusions highlights the need for precision and robustness. To address this, Wang et al. propose a depth-aware concealed crop detection method for dense agricultural scenes, along with the corresponding dataset ACOD-12K, specifically designed to tackle COD-related agricultural tasks.
Annotation-efficient learning in limited conditions. In the context of COD under constrained conditions, annotation-efficient learning emerges as a crucial approach to addressing the scarcity of labeled data. Strategies such as few/zero-shot learning, weakly supervised learning, and open-world learning aim to alleviate the challenge of exhaustive annotation requirements. Specifically, few/zero-shot learning addresses the detection of unseen object classes by leveraging knowledge transfer from related categories. Weakly supervised learning, which relies on image-level labels to identify object locations, provides a less precise but still valuable alternative. Unsupervised learning, through techniques like clustering and representation learning, uncovers patterns in unlabeled data, thereby enhancing model generalization. Self-supervised learning further contributes by exploiting inherent data properties to generate pseudo-labels, fostering the development of robust representations. The limitation of training data also necessitates innovative data augmentation and domain adaptation techniques to combat overfitting. Open-world applications introduce unique challenges, requiring models to detect novel objects while maintaining performance on known classes. Collectively, these approaches strive to overcome the difficulties posed by scarce annotations and unseen object classes, thereby pushing the boundaries of concealed object detection in practical, annotation-constrained scenarios.
Camouflage loss functions. The introduction of camouflage-specific loss functions enhances the training process for COD by directing model optimization toward accurately distinguishing camouflaged regions from salient ones. For instance, SCLoss shows promise by incorporating spatial coherence into the learning process for ambiguous regions. This approach significantly improves the network’s ability to discern boundaries and subtle nuances in camouflaged objects. By focusing on both individual pixel responses and their mutual interactions, SCLoss addresses the limitations of single-response loss functions, providing a more comprehensive solution to the challenges posed by camouflaged objects.
Real-time performance constraints. A significant challenge in COD is the lack of real-time performance, primarily due to the high computational and memory requirements of current models. This limitation impedes deployment on resource-constrained devices, restricting real-world applications such as surveillance and autonomous navigation. Train-free learning offers a solution by enabling rapid adaptation without retraining, thereby facilitating quicker deployment. Additionally, green learning-based gradient boosting emphasizes the development of energy-efficient models that reduce computational costs by selectively focusing on challenging samples, thereby optimizing real-time performance. The development of efficient networks with lightweight architectures aims to minimize latency and memory usage, making real-time COD feasible in practical scenarios. Overcoming these constraints is essential for the broader adoption and practical impact of COD technologies.
2 Exploring expansive potentials
Embracing Novel Tasks. CoCOD, a collaborative approach to COD, holds the potential to significantly enhance performance by leveraging multi-source data and cross-modal interactions. However, challenges persist in efficiently fusing heterogeneous information and ensuring robust performance across diverse scenarios. Another promising avenue, RefCOD, demands advancements in techniques for multi-modal alignment and the comprehension of complex textual or visual references, particularly for rare or ambiguous species. The primary difficulties include bridging cross-modal gaps and extracting fine-grained visual cues that are relevant to the provided textual or visual descriptions. Beyond CoCOD and RefCOD, there are numerous unexplored opportunities for novel tasks that could further empower COD. For instance, Interactive COD (I-COD) could incorporate user feedback during detection, enabling iterative refinement and personalization. This would require the development of robust interaction mechanisms and ensuring the system’s responsiveness to user inputs . By continually exploring and innovating in these novel tasks, we can substantially enhance the capabilities and applicability of COD, thereby pushing the boundaries of what is possible in this exciting field.
Differentialting SOD and COD. At the feature level, delving into the nuances between SOD and COD is crucial, as both, though under the umbrella of abnormal segmentation, target distinct abnormalities . Current COD techniques struggle with salient objects , as they misinterpret prominence as camouflage, thereby necessitating robustness enhancement. Transferring salient objects to concealed scenarios, or vice versa, could alleviate data scarcity and boost model generalization. The partial positive correlation between SOD and COD highlights opportunities for conversion strategies, increasing sample diversity. Integrating generative adversarial mechanisms between the two tasks fosters innovation. Ultimately, a multi-task unified learning framework that encapsulates domain self-generalization for saliency and camouflage detection within the broader AI landscape represents a compelling frontier.
Multimodal information fusion. Integrating multiple modalities, such as text, audio, video, optical, depth, infrared, and 3D, presents a promising yet challenging direction in COD. Each modality offers unique insights but also introduces additional complexities. For instance, RGB-T and thermal infrared modalities can enhance detection in low-light or obscured environments, yet they require robust fusion strategies to effectively combine visual cues with temperature variations. Audio cues, while informative for certain concealed objects, present challenges in accurately localizing and interpreting sounds amidst background noise. Video data, though rich in temporal information, necessitates efficient processing techniques to handle high dimensionality and real-time requirements. The key challenges lie in designing robust fusion mechanisms that can leverage the complementary strengths of these modalities while mitigating their individual limitations. Furthermore, annotating multi-modal datasets for concealed objects is both time-consuming and expensive, exacerbating the scarcity of labeled data. Addressing these challenges will require innovative approaches in multi-modal representation learning, fusion algorithms, and data augmentation techniques that are specifically tailored for COD.
Efficiency-oriented large-scale vision-language models. Large vision-language models (VLMs) have demonstrated superior performance across various tasks, attracting significant attention from researchers. However, training VLMs specifically tailored for camouflaged object detection (COD) is not ideal due to limited data availability and the heavy computational demands involved. In this context, training-free prompt engineering has emerged as a promising direction to address the resource-intensive challenges associated with training large models. The primary challenge lies in designing effective prompts that can elicit nuanced understanding from these models without requiring extensive training. Recent advancements, such as SAM and LVLMs, offer compelling avenues for integrating visual and linguistic modalities, yet efficiency remains a major concern. Developing relatively lightweight models, combined with advanced prompt engineering techniques, could potentially mitigate efficiency issues while still handling diverse and challenging samples. SAM 2, as a pioneering effort, significantly advances the frontier of real-time video processing by leveraging a streaming memory transformer architecture. Its success in reducing interactions by 3x in video segmentation underscores the potential for enhancing efficiency in COD tasks, particularly within the video domain. To balance performance with efficiency, we could leverage the strengths of both large and lightweight models, potentially through techniques such as knowledge distillation, transfer learning, model compression, domain adaptation, meta-learning, or modularized architectures tailored for specific tasks. Additionally, when there is a significant domain gap between training data and test data, designing a plug-and-play component is a promising approach. For instance, adapters enable efficient fine-tuning of large pre-trained models, allowing them to adapt to the nuances of camouflaged objects without the full cost of retraining and thereby significantly reducing computational overheads.
Conclusion
We present a comprehensive and exhaustive overview of the rapidly evolving field of COD, encompassing both traditional and deep learning methods in image and video domains. By reviewing approximately 150 relevant COD studies, we offer the most extensive and detailed survey to date, providing a concise yet thorough perspective for both newcomers and established scholars. Additionally, we provide both quantitative and qualitative benchmarks for representative image and video models across 6 characteristic datasets and 6 evaluation metrics. Through rigorous benchmarking, we have identified key limitations and challenges in current COD methods, thereby paving the way for future research directions. We hope this survey will inspire innovative solutions that push the boundaries of COD technology. Additionally, we have established a dedicated GitHub repository to house COD techniques, datasets, and resources, ensuring that the latest developments and insights are readily accessible to the research community for further exploration.
Acknowledgments
The authors express their sincere appreciation to Dr. Dengping Fan for his insightful comments, which greatly improved the quality of this paper.