Concealed Object Detection

Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng, Ling Shao

Introduction

Can you find the concealed object(s) in each image of Fig. 1 within 10 seconds? Biologists refer to this as background matching camouflage (BMC) , where one or more objects attempt to adapt their coloring to match “seamlessly” with the surroundings in order to avoid detection . Sensory ecologists have found that this BMC strategy works by deceiving the visual perceptual system of the observer. Naturally, addressing concealed object detection (CODWe define COD as segmenting objects or stuff (amorphous regions ) that have a similar pattern, e.g., texture, color, direction, etc., to their natural or man-made environment. In the rest of the paper, for convenience, the concealed object segmentation is considered identical to COD and used interchangeably.) requires a significant amount of visual perception knowledge. Understanding COD has not only scientific value in itself, but it also important for applications in many fundamental fields, such as computer vision (e.g., for search-and-rescue work, or rare species discovery), medicine (e.g., polyp segmentation , lung infection segmentation ), agriculture (e.g., locust detection to prevent invasion), and art (e.g., recreational art ).

In Fig. 2, we present examples of generic, salient, and concealed object detection. The high intrinsic similarities between the targets and non-targets make COD far more challenging than traditional object segmentation/detection . Although it has gained increased attention recently, studies on COD still remain scarce, mainly due to the lack of a sufficiently large dataset and a standard benchmark like Pascal-VOC , ImageNet , MS-COCO , ADE20K , and DAVIS .

In this paper, we present the first complete study for the concealed object detection task using deep learning, bringing a novel view to object detection from a concealed perspective.

COD10K Dataset. With the goal mentioned above, we carefully assemble COD10K, a large-scale concealed object detection dataset. Our dataset contains 10,000 images covering 78 object categories, such as terrestrial, amphibians, flying, aquatic, etc. All the concealed images are hierarchically annotated with category, bounding-box, object-level, and instance-level labels (Fig. 3), benefiting many related tasks, such as object proposal, localization, semantic edge detection, transfer learning , domain adaption , etc. Each concealed image is assigned challenging attributes (e.g., shape complexity-SC, indefinable boundaries-IB, occlusions-OC) found in the real-world and matting-level labeling (which takes ∼\sim60 minutes per image). These high-quality labels could help provide deeper insight into the performance of models.

COD Framework. We propose a simple but efficient framework, named SINet (Search Identification Net). Remarkably, the overall training time of SINet takes 4 hours and it achieves the new state-of-the-art (SOTA) on all existing COD datasets, suggesting that it could offer a potential solution to concealed object detection. Our network also yield several interesting findings (e.g., search and identification strategy is suitable for COD), making various potential applications more feasible.

COD Benchmark. Based on the collected COD10K and previous datasets , we offer a rigorous evaluation of 12 SOTA baselines, making ours the largest COD study. We report baselines in two scenarios, i.e., super-class and sub-class. We also track the community’s progress via an online benchmark (http://dpfan.net/camouflage/).

Downstream Applications. To further support research in the field, we develop an online demo (http://mc.nankai.edu.cn/cod) to enable other researchers to test their scenes easily. In addition, we also demonstrate several potential applications such as medicine, manufacturing, agriculture, art, etc.

Future Directions. Based on the proposed COD10K, we also discuss ten promising directions for future research. We find that concealed object detection is still far from being solved, leaving large room for improvement.

This paper is based on and extends our conference version in terms of several aspects. First, we provide a more detailed analysis of our COD10K, including the taxonomy, statistics, annotations, and resolutions. Secondly, we improve the performance our SINet model by introducing neighbor connection decoder (NCD) and group-reversal attention (GRA). Thirdly, we conduct extensive experiments to validate the effectiveness of our model, and provide several ablation studies for the different modules within our framework. Fourth, we provide an exhaustive super-class and sub-class benchmarking and a more insightful discussion regarding the novel COD task. Last but not least, based on our benchmark results, we draw several important conclusions and highlight several promising future directions, such as concealed object ranking, concealed object proposal, concealed instance segmentation.

Related Work

In this section, we briefly review closely related works. Following , we roughly divide object detection into three categories: generic, salient, and concealed object detection.

Generic Object Segmentation (GOS). One of the most popular directions in computer vision is generic object segmentation . Note that generic objects can be either salient or concealed. Concealed objects can be seen as difficult cases of generic objects. Typical GOS tasks include semantic segmentation and panoptic segmentation (see Fig. 2 b).

Salient Object Detection (SOD). This task aims to identify the most attention-grabbing objects in an image and then segment their pixel-level silhouettes . The flagship products that make use of SOD technology are Huawei’s smartphones, which employ SOD to create what they call “AI Selfies”. Recently, Qin et al. applied the SOD algorithm to two (near) commercial applications: AR COPY & PASTE https://github.com/cyrildiagne/ar-cutpaste and OBJECT CUThttps://github.com/AlbertSuarez/object-cut. These applications have already drawn great attention (12K github stars) and have important real-world impacts.

Although the term “salient” is essentially the opposite of “concealed” (standout vs. immersion), salient objects can nevertheless provide important information for COD, e.g., images containing salient objects can be used as the negative samples. Giving a complete review on SOD is beyond the scope of this work. We refer readers to recent survey and benchmark papers for more details. Our online benchmark is publicly available at: http://dpfan.net/socbenchmark/.

Concealed Object Detection (COD). Research into COD, which has had a tremendous impact on advancing our knowledge of visual perception, has a long and rich history in biology and art. Two remarkable studies on concealed animals from Abbott Thayer and Hugh Cott are still hugely influential. The reader can refer to the survey by Stevens et al. for more details on this history. There are also some concurrent works that are accepted after this submission.

COD Datasets. CHAMELEON is an unpublished dataset that has only 76 images with manually annotated object-level ground-truths (GTs). The images were collected from the Internet via the Google search engine using “concealed animal” as a keyword. Another contemporary dataset is CAMO , which has 2.5K images (2K for training, 0.5K for testing) covering eight categories. It has two sub-datasets, CAMO and MS-COCO, each of which contains 1.25K images. Unlike existing datasets, the goal of our COD10K is to provide a more challenging, higher quality, and more densely annotated dataset. COD10K is the largest concealed object detection dataset so far, containing 10K images (6K for training, 4K for testing). See Table I for details.

Types of Camouflage. Concealed images can be roughly split into two types: those containing natural camouflage and those with artificial camouflage. Natural camouflage is used by animals (e.g., insects, sea horses, and cephalopods) as a survival skill to avoid recognition by predators. In contrast, artificial camouflage is usually used in art design/gaming to hide information, occurs in products during the manufacturing process (so-called surface defects , defect detection ), or appears in our daily life (e.g., transparent objects ).

COD Formulation. Unlike class-aware tasks such as semantic segmentation, concealed object detection is a class-agnostic task. Thus, the formulation of COD is simple and easy to define. Given an image, the task requires a concealed object detection algorithm to assign each pixel ii a label Labeli∈Label_{i}\in {0,1}, where LabeliLabel_{i} denotes the binary value of pixel ii. A label of 0 is given to pixels that do not belong to the concealed objects, while a label of 1 indicates that a pixel is fully assigned to the concealed objects. We focus on object-level concealed object detection, leaving concealed instance detection to our future work.

COD10K Dataset

The emergence of new tasks and datasets has led to rapid progress in various areas of computer vision. For instance, ImageNet revolutionized the use of deep models for visual recognition. With this in mind, our goals for studying and developing a dataset for COD are: (1) to provide a new challenging object detection task from the concealed perspective, (2) to promote research in several new topics, and (3) to spark novel ideas. Examples from COD10K are shown in Fig. 1. We will provide the details on COD10K in terms of three key aspects including image collection, professional annotation, and dataset features and statistics.

As discussed in , the quality of annotation and size of a dataset are determining factors for its lifespan as a benchmark. To this end, COD10K contains 10,000 images (5,066 concealed, 3,000 background, 1,934 non-concealed), divided into 10 super-classes (i.e., flying, aquatic, terrestrial, amphibians, other, sky, vegetation, indoor, ocean, and sand), and 78 sub-classes (69 concealed, 9 non-concealed) which were collected from multiple photography websites.

Most concealed images are from Flickr and have been applied for academic use with the following keywords: concealed animal, unnoticeable animal, concealed fish, concealed butterfly, hidden wolf spider, walking stick, dead-leaf mantis, bird, sea horse, cat, pygmy seahorses, etc. (see Fig. 4) The remaining concealed images (around 200 images) come from other websites, including Visual Hunt, Pixabay, Unsplash, Free-images, etc., which release public-domain stock photos, free from copyright and loyalties. To avoid selection bias , we also collected 3,000 salient images from Flickr. To further enrich the negative samples, 1,934 non-concealed images, including forest, snow, grassland, sky, seawater and other categories of background scenes, were selected from the Internet. For more details on the image selection scheme, we refer to Zhou et al. .

2 Professional Annotation

Recently released datasets have shown that establishing a taxonomic system is crucial when creating a large-scale dataset. Motivated by , our annotations (obtained via crowdsourcing) are hierarchical (category ↣\rightarrowtail bounding box ↣\rightarrowtail attribute ↣\rightarrowtail object/instance).

∙\bullet Categories. As illustrated in Fig. 6, we first create five super-class categories. Then, we summarize the 69 most frequently appearing sub-class categories according to our collected data. Finally, we label the sub-class and super-class of each image. If the candidate image does not belong to any established category, we classify it as ‘other’.

∙\bullet Bounding boxes. To extend COD10K for the concealed object proposal task, we also carefully annotate the bounding boxes for each image.

∙\bullet Attributes. In line with the literature , we label each concealed image with highly challenging attributes faced in natural scenes, e.g., occlusions, indefinable boundaries. Attribute descriptions and the co-attribute distribution is shown in Fig. 7.

∙\bullet Objects/Instances. We stress that existing COD datasets focus exclusively on object-level annotations (Table I). However, being able to parse an object into its instances is important for computer vision researchers to be able to edit and understand a scene. To this end, we further annotate objects at an instance-level, like COCO , resulting in 5,069 objects and 5,930 instances.

3 Dataset Features and Statistics

We now discuss the proposed dataset and provide some statistics.

∙\bullet Resolution distribution. As noted in , high-resolution data provides more object boundary details for model training and yields better performance when testing. Fig. 8 presents the resolution distribution of COD10K, which includes a large number of Full HD 1080p resolution images.

∙\bullet Object size. Following , we plot the normalized (i.e., related to image areas) object size in Fig. 9 (top-left), i.e., the size distribution from 0.01%∼\sim 80.74% (avg.: 8.94%), showing a broader range compared to CAMO-COCO, and CHAMELEON.

∙\bullet Global/Local contrast. To evaluate whether an object is easy to detect, we describe it using the global/local contrast strategy . Fig. 9 (top-right) shows that objects in COD10K are more challenging than those in other datasets.

∙\bullet Center bias. This commonly occurs when taking a photo, as humans are naturally inclined to focus on the center of a scene. We adopt the strategy described in to analyze this bias. Fig. 9 (bottom-left/right) shows that our COD10K dataset suffers from less center bias than others.

∙\bullet Quality control. To ensure high-quality annotation, we invited three viewers to participate in the labeling process for 10-fold cross-validation. Fig. 10 shows examples that were passed/rejected. This matting-level annotation costs ∼\sim 60 minutes per image on average.

∙\bullet Super/Sub-class distribution. COD10K includes five concealed super-classes (i.e., terrestrial, atmobios, aquatic, amphibian, other) and 69 sub-classes (e.g., bat-fish, lion, bat, frog, etc). Examples of the word cloud and object/instance number for various categories are shown in Fig. 5 & Fig. 11, respectively.

∙\bullet Dataset splits. To provide a large amount of training data for deep learning algorithms, our COD10K is split into 6,000 images for training and 4,000 for testing, randomly selected from each sub-class.

∙\bullet Diverse concealed objects. In addition to the general concealed patterns, such as those in Fig. 1, our dataset also includes various other types of concealed objects, such as concealed body paintings and conceale in daily life (see Fig. 12).

COD Framework

Fig. 13 illustrates the overall concealed object detection framework of the proposed SINet (Search Identification Network). Next, we explain our motivation and introduce the network overview.

Motivation. Biological studies have shown that, when hunting, a predator will first judge whether a potential prey exists, i.e., it will search for a prey. Then, the target animal can be identified; and, finally, it can be caught.

Introduction. Several methods have shown that satisfactory performance is dependent on the re-optimization strategy (i.e., coarse-to-fine), which is regarded as the composition of multiple sub-steps. This also suggests that decoupling the complicated targets can break the performance bottleneck. Our SINet model consists of the first two stages of hunting, i.e., search and identification. Specifically, the former phase (Section 4.2) is responsible for searching for a concealed object, while the latter one (Section 4.3) is then used to precisely detect the concealed object in a cascaded manner.

Next, we elaborate on the details of the three main modules, including a) the texture enhanced module (TEM), which is used to capture fine-grained textures with the enlarged context cues; b) the neighbor connection decoder (NCD), which is able to provide the location information; and c) the cascaded group-reversal attention (GRA) blocks, which work collaboratively to refine the coarse prediction from the deeper layer.

2 Search Phase

Texture Enhanced Module (TEM). Neuroscience experiments have verified that, in the human visual system, a set of various sized population receptive fields helps to highlight the area close to the retinal fovea, which is sensitive to small spatial shifts . This motivates us to use the TEM to incorporate more discriminative feature representations during the searching stage (usually in a small/local space). As shown in Fig. 13, each TEM component includes four parallel residual branches {bi,i=1,2,3,4}\{b_{i},i=1,2,3,4\} with different dilation rates d∈{1,3,5,7}d\in\{1,3,5,7\} and a shortcut branch (gray arrow), respectively. In each branch bib_{i}, the first convolutional layer utilizes a 1×11\times 1 convolution operation (Conv1×\times1) to reduce the channel size to 32. This is followed by two other layers: a (2i−1)×(2i−1)(2i-1)\times(2i-1) convolutional layer and a 3×33\times 3 convolutional layer with a specific dilation rate (2i−1)(2i-1) when i>1i>1. Then, the first four branches {bi,i=1,2,3,4}\{b_{i},i=1,2,3,4\} are concatenated and the channel size is reduced to CC via a 3×\times3 convolution operation. Note that we set C=32C=32 in the default implementation of our network for time-cost trade-off. Finally, the identity shortcut branch is added in, then the whole module is fed to a ReLU function to obtain the output feature fk′f_{k}^{\prime}. Besides, several works (e.g., Inception-V3 ) have suggested that the standard convolution operation of size (2i−1)×(2i−1)(2i-1)\times(2i-1) can be factorized as a sequence of two steps with (2i−1)×1(2i-1)\times 1 and 1×(2i−1)1\times(2i-1) kernels, speeding-up the inference efficiency without decreasing the representation capabilities. All of these ideas are predicated on the fact that a 2D kernel with a rank of one is equal to a series of 1D convolutions . In brief, compared to the standard receptive fields block structure , TEM add one more branch with a larger dilation rate to enlarge the receptive field and further replace the standard convolution with two asymmetric convolutional layers. For more details please refer to Fig. 13.

However, there are still two key issues when aggregating multiple feature pyramids; namely, how to maintain semantic consistency within a layer and how to bridge the context across layers. Here, we propose to address these with the neighbor connection decoder (NCD). More specifically, we modify the partial decoder component (PDC) with a neighbor connection function and get three refined features fknc=FNC(fk′;WNCu)f_{k}^{nc}=F_{NC}(f_{k}^{\prime};\mathbf{W}_{NC}^{u}), k∈{3,4,5}k\in\{3,4,5\} and u∈{1,2,3}u\in\{1,2,3\}, which are formulated as:

where g[⋅;WNCu]g[\cdot;\mathbf{W}_{NC}^{u}] denotes a 3×\times3 convolutional layer followed by a batch normalization operation. To ensure shape matching between candidate features, we utilize an upsampling (e.g., 2 times) operation δ↑2(⋅)\delta^{2}_{\uparrow}(\cdot) before element-wise multiplication ⊗\otimes. Then, we feed fknc,k∈{3,4,5}f_{k}^{nc},k\in\{3,4,5\} into the neighbor connection decoder (NCD) and generate the coarse location map C6\mathbf{C_{6}}.

3 Identification Phase

Reverse Guidance. As discussed in Section 4.2, our global location map C6C_{6} is derived from the three highest layers, which can only capture a relatively rough location of the concealed object, ignoring structural and textural details (see Fig. 13). To address this issue, we introduce a principled strategy to mine discriminative concealed regions by erasing objects . As shown in Fig. 14 (b), we obtain the output reverse guidance r1kr_{1}^{k} via sigmoid and reverse operation. More precisely, we obtain the output reverse attention guidance r1kr_{1}^{k} by a reverse operation, which can be formulated as:

where δ↓4\delta^{4}_{\downarrow} and δ↑2\delta^{2}_{\uparrow} denote a ×\times4 down-sampling and ×\times2 up-sampling operation, respectively. σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}) is the sigmoid function, which is applied to convert the mask into the interval . ⊝\circleddash is a reverse operation subtracting the input from matrix E\mathbf{E}, in which all the elements are 11.

Group-Reversal Attention (GRA). Finally, we introduce the residual learning process, termed the GRA block, with the assistance of both the reverse guidance and group guidance operation. According to previous studies , multi-stage refinement can improve performance. We thus combine multiple GRA blocks (e.g., GikG^{k}_{i}, i∈{1,2,3},k∈{3,4,5}i\in\{1,2,3\},k\in\{3,4,5\}) to progressively refine the coarse prediction via different feature pyramids. Overall, each GRA block has three residual learning processes:

We combine candidate features pikp_{i}^{k} and r1kr_{1}^{k} via the group guidance operation and then use the residual stage to produce the refined features pi+1kp_{i+1}^{k}. This is formulated as:

where Wv\mathbf{W}^{v} denotes the convolutional layer with a 3×\times3 kernel followed by batch normalization layer for reducing the channel number from C+miC+m_{i} to CC. Note that we only reverse the guidance prior in the first GRA block (i.e., when i=1i=1) in the default implementation. Refer to Section 5.3 for detailed discussion.

Then, we get a single channel residual guidance:

which is parameterized by learnable weights WGRAw\mathbf{W}^{w}_{GRA}.

Finally, we only output the refined guidance, which serves as the residual prediction. It is formulated as:

where δ(⋅)\delta(\cdot) is δ↑2\delta^{2}_{\uparrow} when k={3,4}k=\{3,4\} and δ↓4\delta^{4}_{\downarrow} when k=5k=5.

4 Implementation Details

Our loss function is defined as L=LIoUW+LBCEW\textit{L}=\textit{L}_{IoU}^{W}+\textit{L}_{BCE}^{W}, where LIoUW\textit{L}_{IoU}^{W} and LBCEW\textit{L}_{BCE}^{W} represent the weighted intersection-over-union (IoU) loss and binary cross entropy (BCE) loss for the global restriction and local (pixel-level) restriction. Different from the standard IoU loss, which has been widely adopted in segmentation tasks, the weighted IoU loss increases the weights of hard pixels to highlight their importance. In addition, compared with the standard BCE loss, LBCEW\textit{L}_{BCE}^{W} pays more attention to hard pixels rather than assigning all pixels equal weights. The definitions of these losses are the same as in and their effectiveness has been validated in the field of salient object detection. Here, we adopt deep supervision for the three side-outputs (i.e., C3C_{3}, C4C_{4}, and C5C_{5}) and the global map C6C_{6}. Each map is up-sampled (e.g., C3upC_{3}^{up}) to the same size as the ground-truth map GG. Thus, the total loss for the proposed SINet can be formulated as: Ltotal=L(C6up,G)+∑i=3i=5L(Ciup,G)\textit{L}_{total}=\textit{L}(C_{6}^{up},G)+\sum_{i=3}^{i=5}\textit{L}(C_{i}^{up},G).

4.2 Hyperparameter Settings

SINet is implemented in PyTorch and trained with the Adam optimizer . During the training stage, the batch size is set to 36, and the learning rate starts at 1e-4, dividing by 10 every 50 epochs. The whole training time is only about 4 hours for 100 epochs. The running time is measured on an Intel® i9-9820X CPU @3.30GHz ×\times 20 platform and a single NVIDIA TITAN RTX GPU. During inference, each image is resized to 352×\times352 and then fed into the proposed pipeline to obtain the final prediction without any post-processing techniques. The inference speed is ∼\sim45 fps on a single GPU without I/O time. Both PyTorch and Jittor verisons of the source code will be made publicly avaliable.

COD Benchmark

Mean absolute error (MAE) is widely used in SOD tasks. Following Perazzi et al. , we also adopt the MAE (MM) metric to assess the pixel-level accuracy between a predicted map and ground-truth. However, while useful for assessing the presence and amount of error, the MAE metric is not able to determine where the error occurs. Recently, Fan et al. proposed a human visual perception based E-measure (EϕE_{\phi}) , which simultaneously evaluates the pixel-level matching and image-level statistics. This metric is naturally suited for assessing the overall and localized accuracy of the concealed object detection results. Note that we report mean EϕE_{\phi} in the experiments. Since concealed objects often contain complex shapes, COD also requires a metric that can judge structural similarity. We therefore utilize the S-measure (SαS_{\alpha}) as our structural similarity evaluation metric. Finally, recent studies have suggested that the weighted F-measure (FβwF_{\beta}^{w}) can provide more reliable evaluation results than the traditional FβF_{\beta}. Thus, we further consider this as an alternative metric for COD. Our one-key evaluation code is also available at the project page.

1.2 Baseline Models

We select 12 deep learning baselines according to the following criteria: a) classical architectures, b) recently published, and c) achieve SOTA performance in a specific field.

1.3 Training/Testing Protocols

For fair comparison with our previous version , we adopt the same training settings for the baselines.To verify the generalizability of SINet, we only use the combined training set of CAMO and COD10K without EXTRA (i.e., additional) data. We evaluate the models on the whole CHAMELEON dataset and the test sets of CAMO and COD10K.

2 Results and Data Analysis

This section provides the quantitative evaluation results on CHAMELEON, CAMO, and COD10K datasets, respectively.

Performance on CHAMELEON. From Table II, compared with the 12 SOTA object detection baselines and ANet-SRM, our SINet achieves the new SOTA performances across all metrics. Note that our model does not apply any auxiliary edge/boundary features (e.g., EGNet , PFANet ), pre-processing techniques , or post-processing strategies such as .

Performance on CAMO. We also test our model on the CAMO dataset, which includes various concealed objects. Based on the overall performances reported in Table II, we find that the CAMO dataset is more challenging than CHAMELEON. Again, SINet obtains the best performance, further demonstrating its robustness.

Performance on COD10K. With the test set (2,026 images) of our COD10K dataset, we again observe that the proposed SINet is consistently better than other competitors. This is because its specially designed search and identification modules can automatically learn rich diversified features from coarse to fine, which are crucial for overcoming challenging ambiguities in object boundaries. The results are shown in Table II and Table III.

Per-subclass Performance. In addition to the overall quantitative comparisons on our COD10K dataset, we also report the quantitative per-subclass results in the Table IV to investigate the pros and cons of the models for future researchers. In Fig. 15, we additionally show the minimum, mean, and maximum S-measure performance of each sub-class over all baselines. The easiest sub-class is “Grouse”, while the most difficult is the “LeafySeaDragon”, from the aquatic and terrestrial categories, respectively.

Qualitative Results. We present more detection results of our conference version model (SINet_cvpr) for various challenging concealed objects, such as spider, moth, sea horse, and toad, in the supplementary materials. As shown in Fig. 16, SINet further improves the visual results compared to SINet_cvpr in terms of different lighting (1st1^{st} row), appearance changes (2nd2^{nd} row), and indefinable boundaries (3rd3^{rd} to 5th5^{th}). PFANet is able to locate the concealed objects, but the outputs are always inaccurate. By further using reverse attention module, PraNet achieves a relatively more accurate location than PFANet in the first case. Nevertheless, it still misses the fine details of objects, especially for the fish in the 2nd2^{nd} and 3rd3^{rd} rows. For all these challenging cases, SINet is able to infer the real concealed object with fine details, demonstrating the robustness of our framework.

GOS vs. SOD Baselines. One noteworthy finding is that, among the top-3 models, the GOS model (i.e., FPN ) performs worse than the SOD competitors, CPD , EGNet , suggesting that the SOD framework may be better suited for extension to COD tasks. Compared with both the GOS and the SOD models, SINet significantly decrease the training time (e.g., SINet: 4 hours vs. EGNet: 48 hours) and achieve the SOTA performance on all datasets, showing that they are promising solutions for the COD problem. Due to the limited space, fully comparing them with existing SOTA SOD models is beyond the scope of this paper. Note that our main goal is to provide more general observations for future work. More recent SOD models can be found in our project page.

Generalization. The generalizability and difficulty of datasets play a crucial role in both training and assessing different algorithms . Hence, we study these aspects for existing COD datasets, using the cross-dataset analysis method , i.e., training a model on one dataset, and testing it on others. We select two datasets, namely CAMO , and our COD10K. Following , for each dataset, we randomly select 800 images as the training set and 200 images as the testing set. For fair comparison, we train SINet_cvpr on each dataset until the loss is stable.

Table V provides the S-measure results for the cross-dataset generalization. Each row lists a model that is trained on one dataset and tested on all others, indicating the generalizability of the dataset used for training. Each column shows the performance of one model tested on a specific dataset and trained on all others, indicating the difficulty of the testing dataset. Please note that the training/testing settings are different from those used in Table II, and thus the performances are not comparable. As expected, we find that our COD10K has better generalization ability than the CAMO (e.g., the last column ‘Drop↓\downarrow: -6.0%’). This is because our dataset contains a variety of challenging concealed objects (Section 3). We can thus see that our COD10K dataset contains more challenging scenes.

3 Ablation Studies

We now provide a detailed analysis of the proposed SINet on CHAMELEON, CAMO, and COD10K. We verify the effectiveness by decoupling various sub-components, including the NCD, TEM, and GRA, as summarized in Table VI. Note that we maintain the same hyperparameters mentioned in Section 4.4 during the re-training process for each ablation variant.

Effectiveness of NCD. We explore the influence of the decoder in the search phase of our SINet. To verify its necessity, we retrain our network without the NCD (No.#1) and find that, compared with #OUR (last row in Table VI), the NCD is attributed to boosting the performance on CAMO, increasing the mean EϕE_{\phi} score from 0.869 to 0.882. Further, we replace the NCD with the partial decoder (i.e., PD of No.#2) to test the performance of this scheme. Comparing No.#2 with #OUR, our design can enhance the performance slightly, increasing it by 1.7% in terms of FβwF^{w}_{\beta} on the CHAMELEON.

As shown in Fig. 17, we present a novel feature aggregation strategy before the modified UNet-like decoder (removing the bottom-two high-resolution layers), termed the NCD, with neighbor connections between adjacent layers. This design is motivated by the fact that the high-level features are superior to semantic strength and location accuracy, but introduce noise and blurred edges for the target object.

Instead of broadcasting features from densely connected layers with a short connection or a partial decoder with a skip connection , our NCD exploits the semantic context through a neighbor connection, providing a simple but effective way to reduce inconsistency between different features. Aggregating all features by a short connection increases the parameters. This is one of the major differences between DSS (Fig. 17 a) and NCD. Compared to CPD (Fig. 17 b), which ignores feature transparency between f5′f_{5}^{\prime} and f4′f_{4}^{\prime}, NCD is more efficient at broadcasting the features step by step.

Effectiveness of TEM. We provide two different variants: (a) without TEM (No.#3), and (b) with symmetric convolutional layers (No.#4). Comparing with No.#3, we find that our TEM with asymmetric convolutional layers (No.#OUR) is necessary for increasing the performance on the CAMO dataset. Besides, replacing the standard symmetric convolutional layer (No.#4) with an asymmetric convolutional layer (No.#OUR) has little impact on the learning capability of the network, while further increasing the mean EϕE_{\phi} from 0.866 to 0.882 on the CAMO dataset.

Effectiveness of GRA. Reverse Guidance. As shown in the ‘Reverse’ column of Table VI, {*,*,*} indicates whether the guidance is reversed (see Fig. 14 (b)) before each GRA block GikG^{k}_{i}. For instance, {1,0,0} means that we only reverse the guidance in the first block (i.e., r1kr^{k}_{1}) and the remaining two blocks (i.e., r2kr^{k}_{2} and r3kr^{k}_{3}) do not have a reverse operation.

We investigate the contribution of the reverse guidance in the GRA, including three alternatives: (a) without any reverse, i.e., {0,0,0} of No.#5, (b) reversing the first two guidances rik,i∈{1,2}r^{k}_{i},i\in\{1,2\}, i.e., {1,1,0} of No.#6, and (c) reversing all the guidances rik,i∈{1,2,3}r^{k}_{i},i\in\{1,2,3\}, i.e., {1,1,1} of No.#7. Compared to the default implementation of SINet (i.e., {1,0,0} of No.#OUR), we find that only reversing the first guidance may help the network to mine diversified representations from two perspectives (i.e., attention and reverse attention regions), while introducing reverse guidance several times in the intermediate process may cause confusion during the learning procedure, especially for setting #6 on the CHAMELEON and COD10K datasets.

Group Size of GGO. As shown in the ‘Group Size’ column of Table VI, {∗;∗;∗}\{*;*;*\} indicates the number of feature slices (i.e., group size gig_{i}) from the GGO of the first block G1kG^{k}_{1} to last block G3kG^{k}_{3}. For example, {32;8;1}\{32;8;1\} indicates that we split the candidate feature pik,i∈{1,2,3}p^{k}_{i},i\in\{1,2,3\} into 32, 8, and 1 group sizes at each GRA block Gik,i∈{1,2,3}G^{k}_{i},i\in\{1,2,3\}, respectively. Here, we discuss two ways of selecting the group size, i.e., the uniform strategy (i.e., {1;1;1}\{1;1;1\} of #8, {8;8;8}\{8;8;8\} of #9, {32;32;32}\{32;32;32\} of #10) and progressive strategy (i.e., {1;8;32}\{1;8;32\} of #11 and {32;8;1}\{32;8;1\} of #OUR). We observe that our design based on the progressive strategy can effectively maintain the generalizability of the network, providing more satisfactory performance compared with other variants.

Downstream Applications

Concealed object detection systems have various downstream applications in fields such as medicine, art, and agriculture. Here, we envision some potential uses due to the common feature of these applications where the target objects share similar appearance with the background. Under such circumstances, COD models are very suitable to act as a core component of these applications to mine camouflaged objects. Note that these applications are only toy examples to spark interesting ideas for future research.

As we all know, early diagnosis through medical imaging plays a key role in the treatment of diseases. However, the early disease area/lesions usually have a high degree of homogeneity with the surrounding tissues. As a result, it is difficult for doctors to identity the lesion area in the early stage from a medical image. One typical example is the early colonoscopy to segment polyps, which has contributed to roughly 30% decline in the incidence of colorectal cancer . Similar to concealed object detection, polyp segmentation (see Fig. 18) also faces several challenges, such as variation in appearance and blurred boundaries. The recent state-of-the-art polyp segmentation model, PraNet , has shown promising performance in both polyp segmentation (Top-1) and concealed object segmentation (Top-2). From this point of view, embedding our SINet into this application could potentially achieve more robust results.

1.2 Lung Infection Segmentation

Another concealed object detection example is the lung infection segmentation task in the medical field. Recently, COVID-19 has been of particular concern, and resulted in a global pandemic. An AI system equipped with a COVID-19 lung infection segmentation model would be helpful in the early screening of COVID-19. More details on this application can be found in the recent segmentation model and survey paper . We believe retrain our SINet model using COVID-19 lung infection segmentation datasets will be another interesting potential application.

2 Application II: Manufacturing

In industrial manufacturing, products (e.g., wood, textile, and magnetic tile) of poor quality will inevitably lead to adverse effects on the economy. As can be seen from Fig. 20, the surface defects are challenging, with different factors including low contrast, ambiguous boundaries and so on. Since traditional surface defect detection systems mainly rely on humans, major issues are highly subjective and time-consuming to identify. Thus, designing an automatic recognition system based on AI is essential to increase productivity. We are actively constructing such a data set to advance related research. Some related papers can be found at: https://github.com/Charmve/Surface-Defect-Detection/tree/master/Papers.

3 Application III: Agriculture

Since early 2020, plagues of desert locusts have invaded the world, from Africa to South Asia. Large numbers of locusts gnaw on fields and completely destroy agricultural products, causing serious financial losses and famine due to food shortages. As shown in Fig. 21, introducing AI-based techniques to provide scientific monitoring is feasible for achieving sustainable regulation/containment by governments. Collecting relevant insect data for COD models requires rich biological knowledge, which is also a difficulty faced in this application.

3.2 Fruit Maturity Detection

In the early stages of ripening, many fruits appear similar to green leaves, making it difficult for farmers to monitor production. We present two types of fruits, i.e., Persea Americana and Myrica Rubra, in Fig. 22. These fruits share similar characteristics to concealed objects, so it is possible to utilize a COD algorithm to identify them and improve the monitoring efficiency.

4 Application IV: Art

Background warping to concealed salient objects is a fascinating technique in the SIGGRAPH community. Fig. 23 presents some examples generated by Chu et al. in . We argue that this technique will provide more training data for existing data-hungry deep learning models, and thus it is of value to explore the underlying mechanism behind the feature search and conjunction search theory described by Treisman and Wolfe .

4.2 From Concealed to Salient Objects

Concealed object detection and salient object detection are two opposite tasks, making it convenient for us to design a multi-task learning framework that can simultaneously increase the robustness of the network. As shown in Fig. 24, there exist two reverse objects (a) and (c). An interesting application is to provide a scroll bar to allow users to customize the degree of salient objects from the concealed objects.

5 Application V: Daily Life

Transparent objects, such as glass products, are commonplace in our daily life. These objects/things, including doors and walls, inherit the appearance of their background, making them unnoticeable, as illustrated in Fig. 25. As a sub-task of concealed object detection, transparent object detection and transparent object tracking have shown promise.

5.2 Search Engines

Fig. 26 shows an example of search results from Google. From the results (Fig. 26 a), we notice that the search engine cannot detect the concealed butterfly, and thus only provides images with similar backgrounds. Interestingly, when the search engine is equipped with a concealed detection system (here, we just simply change the keyword), it can identify the concealed object and then feedback several butterfly images (Fig. 26 b).

Potential Research Directions

Despite the recent 10 years of progress in the field of concealed object detection, the leading algorithms in the deep learning era remain limited compared to those for generic object detection and cannot yet effectively solve real-world challenges as shown in our COD10K benchmark (Top-1: Fβw<0.7F_{\beta}^{w}<0.7). We highlight some long-standing challenges, as follows:

Concealed object detection under limited conditions: few/zero-shot learning, weakly supervised learning, unsupervised learning, self-supervised learning, limited training data, unseen object class, etc.

Concealed object detection combined with other modalities: Text, Audio, Video, RGB-D, RGB-T, 3D, etc.

New directions based on the rich annotations provided in the COD10K, such as concealed instance segmentation, concealed edge detection, concealed object proposal, concealed object ranking, among others.

Based on the above-mentioned challenges, there are a number of foreseeable directions for future research:

(1) Weakly/Semi-Supervised Detection: Existing deep-based methods extract the features in a fully supervised manner from images annotated with object-level labels. However, the pixel-level annotations are usually manually marked by LabelMe or Adobe Photoshop tools with intensive professional interaction. Thus, it is essential to utilize weakly/semi (partially) annotated data for training in order to avoid heavy annotation costs.

(2) Self-Supervised Detection: Recent efforts to learn representations (e.g., image, audio, and video) using self-supervised learning have achieved world-renowned achievements, attracting much attention. Thus, it is natural to setup a self-supervised learning benchmark for the concealed object detection task.

(3) Concealed Object Detection in Other Modalities: Existing concealed data is only based on static images or dynamic videos . However, concealed object detection in other modalities can be closely related in domains such as pest monitoring in the dark night, robotics, and artist design. Similar to in RGB-D SOD , RGB-T SOD , CoSOD , and VSOD , these modalities can be audio, thermal, group image, or depth data, raising new challenges under specific scenes.

(4) Concealed Object Classification: Generic object classification is a fundamental task in computer vision. Thus concealed object classification will also likely gain attention in the future. By utilizing the class and sub-class labels provided in COD10K, one could build a large scale and fine-grain classification task.

(5) Concealed Object Proposal and Tracking: In this paper, the concealed object detection is actually a segmentation task. It is different from traditional object detection, which generates a proposal or bounding boxes as the prediction. As such, concealed object proposal and tracking is a new and interesting direction for future work.

(6) Concealed Object Ranking: Currently, most concealed object detection algorithms are built upon binary ground-truths to generate the masks of concealed objects, with only limited works analyzing the rank of concealed objects . However, understanding the level of concealment could help to better explore the mechanism behind the models, providing deeper insights into them. We refer readers to for some inspiring ideas.

(7) Concealed Instance Segmentation: As described in , instance segmentation is more crucial than object-level segmentation for practical applications. For example, we can push the research on camouflaged object segmentation into camouflaged instance segmentation.

(8) Universal Network for Multiple Tasks: As studied by Zamir et al. in Taskonomy , different visual tasks have strong relationships. Thus, their supervision can be reused in one universal system without piling up complexity. It is natural to consider devising a universal network to simultaneously localize, segment and rank concealed objects.

(9) Neural Architecture Search: Both traditional algorithms and deep learning-based models for concealed object detection require human experts with strong prior knowledge or skilled expertise. Sometimes, the hand-crafted features and architectures designed by algorithm engineers may not optimal. Therefore, neural architecture search techniques, such as the popular automated machine learning , offer a potential direction.

(10) Transferring Salient Objects to Concealed Objects: Due to space limitations, we only evaluated typical salient object detection models in our benchmark section. There are several valuable problems that deserve further studying, however, such as transferring salient objects to concealed objects to increase the training data, and introducing a generative adversarial mechanism between the SOD and COD tasks to increase the feature extraction ability of the network.

The ten new research directions listed for concealed object remain far from being solved. However, there are many famous works that can be referred to, providing us a solid basis for studying the object detection task from a concealed perspective.

Conclusion

We have presented the first comprehensive study on object detection from a concealed vision perspective. Specifically, we have provided the new challenging and densely annotated COD10K dataset, conducted a large-scale benchmark, developed a simple but efficient end-to-end search and identification framework (i.e., SINet), and highlighted several potential applications. Compared with existing cutting-edge baselines, our SINet is competitive and generates more visually favorable results. The above contributions offer the community an opportunity to design new models for the COD task. In the future, we plan to extend our COD10K dataset to provide inputs of various forms, such as multi-view images (e.g., RGB-D SOD ), textual descriptions, video (e.g., VSOD ), among others. We also plan to automatically search the optimal receptive fields and employ improved feature representations for better model performance.

Acknowledgments

We thank Guolei Sun and Jianbing Shen for insightful feedback. This research was supported by the National Key Research and Development Program of China under Grant No. 2018AAA0100400, NSFC (61922046), and S&T innovation project from Chinese Ministry of Education.

References