Frequency Perception Network for Camouflaged Object Detection
Runmin Cong, Mengyao Sun, Sanyi Zhang, Xiaofei Zhou, Wei Zhang, Yao Zhao
Introduction
In nature, animals use camouflage to blend in with their surroundings to avoid detection by predators. The camouflaged object detection (COD) task aims to allow computers to automatically recognize these camouflaged objects that blend in with the background, which can be used in numerous downstream applications, including medical segmentation (Fan et al., 2020b; Huang et al., 2020; Cong et al., 2022c, b), unconstrained face recognition (Chen et al., 2023), and recreational art (Feng and Prabhakaran, 2013; Dean et al., 2017). However, the COD task is very challenging due to the low contrast properties between the camouflaged object and the background. Furthermore, camouflaged objects may have multiple appearances, including shapes, sizes, and textures, which further increases the difficulty of detection.
At the beginning of the research, the COD task was also regarded as a low-contrast special case of the salient object detection (SOD) task, but simple SOD model (Cong et al., 2019; Chen et al., 2020; Cong et al., 2023a; Zhang et al., 2021; Cong et al., 2023c, 2022a, 3883, b; Jing et al., 2021; Cong et al., 6476) retraining cannot obtain satisfactory COD results, and usually requires some special positioning design to find camouflaged objects. Recently, with the development of deep learning (Fan et al., 2020c; Ni et al., 2017; Hu et al., 2021; Yue et al., 2022), many customized networks for COD tasks have gradually emerged (Lv et al., 2023; Fan et al., 2020a; Chen et al., 2022). However, current solutions still struggle in challenging situations, such as multiple camouflaged objects, uncertain or fuzzy object boundaries, and occlusion, as shown in Figure 1. In general, these methods mainly design modules in the RGB color domain to detect camouflaged objects, and complete the initial positioning of camouflaged objects by looking for areas with inconsistent information such as textures (called breakthrough points). However, the concealment and confusion of the camouflaged objects itself make this process very difficult. In the image frequency domain analysis, the high-frequency and low-frequency component information in the frequency domain describes the details and contour characteristics of the image in a more targeted manner, which can be used to improve the accuracy of the initial positioning. Inspired by this, we propose a Frequency Perception Network (FPNet) that employs a two-stage strategy of search and recognition to detect camouflaged objects, taking full advantage of RGB and frequency cues.
On the one hand, the main purpose of the frequency-guided coarse positioning stage is to use the frequency domain features to find the breakthrough points of the camouflaged object position. We first adopt the transformer backbone to extract multi-level features of the input RGB image. Subsequently, in order to realize the extraction of frequency domain features, we introduce a frequency-perception module to decompose color features into high-frequency and low-frequency components. Among them, the high-frequency features describe texture features or rapidly changing parts, while the low-frequency features can outline the overall contour of the image. Considering that both texture and contour are important for camouflaged object localization, we fuse them as a complete representation of frequency domain information. In addition, a neighbor interaction mechanism is also employed to combine different levels of frequency-aware features, thereby achieving coarse detection and localization of camouflaged objects. On the other hand, the detail-preserving fine localization stage focuses on progressively prior-guided correction and fusion across layers, thereby generating the final finely camouflaged object masks. Specifically, we design the correction fusion module to achieve the cross-layer high-level feature interaction by integrating the prior-guided correction and cross-layer feature channel association. Finally, the shallow high-resolution features are further introduced to refine and modify the boundaries of camouflaged objects and generate the final COD result.
The main contributions are summarized as follows:
We propose a novel two-stage framework to deeply exploit the advantages of RGB and frequency domains for camouflaged object detection in an end-to-end manner. The proposed network achieves competitive performance on three popular benchmark datasets (i.e., COD10K, CHAMELEON, and CAMO).
A novel fully frequency-perception module is designed to enhance the ability to distinguish camouflaged objects from backgrounds by automatically learning high-frequency and low-frequency features, thereby achieving coarse localization of camouflaged objects.
We design a progressive refinement mechanism to obtain the final refined camouflaged object detection results through prior-guided correction, cross-layer feature channel association, and shallow high-resolution boundary refinement.
Related Work
The COD task aims to localize objects that have a similar appearance to the background, which makes it extremely challenging. Early methods employed hand-crafted low-level features to achieve this goal, such as color (Huerta et al., 2007), expectation-maximization statistics (Liu et al., 2012), convex intensity (Tankus and Yeshurun, 2001), optical flow (Hou and Li, 2011), and texture (Bhajantri and Nagabhushan, 2006; Kavitha et al., 2011). However, due to the imperceptible differences between objects and backgrounds in complex environments, and the limited expressive power of hand-crafted features, they do not perform satisfactorily.
Recently CNN-based methods (Le et al., 2019; Fan et al., 2020a; Mei et al., 2021) have achieved significant success in the COD task. In general, CNN-based methods often employ one or more of the following strategies, such as two-stage strategy (Fan et al., 2020a; Chen et al., 2022), multi-task learning strategy (Lv et al., 2021), and incorporating other guiding cues such as frequency (Gueguen et al., 2018). For instance, Fan et al. (Fan et al., 2020a) proposed a two-stage process named SINet, which represents the new state-of-the-art in existing COD datasets and created the largest COD10K dataset with 10K images. Mei et al. (Mei et al., 2021) imitated the predator-prey process in nature and developed a two-stage bionic framework called PFNet.
In terms of frequency domain studies, Gueguen et al. (Gueguen et al., 2018) directly used the Discrete Cosine Transform (DCT) coefficients of the image as input to CNN for subsequent visual tasks. Ehrlich et al. (Ehrlich and Davis, 2019) presented a general conversion algorithm for transforming spatial domain networks to the frequency domain. Interestingly, both of these works delve deep into the frequency domain transformation of the image JPEG compression process. Subsequently, Zhong et al. (Zhong et al., 2022) modeled the interaction between the frequency domain and the RGB domain, introducing the frequency domain as an additional cue to better detect camouflaged objects from backgrounds. Unlike these methods, on the one hand, we use octave convolution to realize the online learning of frequency domain features, instead of offline extraction methods (e.g., DCT); on the other hand, frequency domain features are mainly used for coarse positioning in the first stage, that is, by making full use of high-frequency and low-frequency information to find the breakthrough point of camouflaged object positioning in the frequency domain.
In addition, some methods (Liu et al., 2021b; Owens et al., 2014; Sun et al., 2022) also try to combine edge detection to extract more precise edges, thereby improving the accuracy of COD. It is worth mentioning that in order to exploit the power of the Transformer model in the COD task, many Transformer-based methods have emerged. For example, Yang et al. (Yang et al., 2021) proposed to incorporate Bayesian learning into Transformer-based reasoning to achieve the COD task. The T2Net proposed by Mao et al. (Mao et al., 2021) in 2021 used a Swin-Transformer as the backbone network, surpassing all CNN-based approaches at that time.
Our Approach
Our goal is to exploit and fuse the inherent advantages of the RGB and frequency domains to enhance the discrimination ability to discover camouflaged objects in the complex background. To that end, in this paper, we propose a Frequency Perception Network (FPNet) for camouflaged object detection, as shown in Figure 2, including a feature extraction backbone, a frequency-guided coarse localization stage, and a detail-preserving fine localization stage.
2. Frequency-guided Coarse Positioning
Inspired by predator hunting systems, frequency information is more advantageous than RGB appearance features for a specific predator in the wild environment. This point of view has also been verified in (Zhong et al., 2022), and then a frequency domain method for camouflaged object detection is proposed. Specifically, this work (Zhong et al., 2022) used offline discrete cosine transform to convert the RGB domain information of an image to the frequency domain, but the offline frequency extraction method limits its flexibility. As described in (Chen et al., 2019), octave convolution can learn to divide an image into low and high frequency components in the frequency domain. The low-frequency features correspond to pixel points with gentle intensity transformations, such as large color blocks, that often represent the main part of the object. The high-frequency components, on the other hand, refer to pixels with intense brightness changes, such as the edges of objects in the image. Inspired by this, we propose a frequency-perception module to automatically separate features into high-frequency and low-frequency parts, and then form a frequency-domain feature representation of camouflaged objects, the detailed process is shown in Figure 3.
Specifically, we employ octave convolution (Chen et al., 2019) to automatically perceive high-frequency and low-frequency information in an end-to-end manner, enabling online learning of camouflaged object detection. The octave convolution can effectively avoid blockiness caused by the DCT and utilize the advantage of the computational speed of GPUs. In addition, it can be easily plugged into arbitrary networks. The detailed process of output of the octave convolution could be described in the following:
Considering that both high-frequency texture attribute and low-frequency contour attribute are important for camouflaged object localization, we fuse them as a complete representation of frequency domain information:
Then, the Neighbor Connection Decoder (NCD) (Deng-Ping et al., 2022), as shown in the top region (the part above the three FPMs) of Figure 2, is adopted to gradually integrate the frequency-domain features of the top-three layers, fully utilizing the cross-layer semantic context relationship through the neighbor layer connection, which can be represented as:
3. Detail-preserving Fine Localization
In the previous section, we introduced how to use frequency-domain features to achieve coarse localization of camouflaged objects. But the first stage is more like a process of finding and locating breakthrough points, the integrity and accuracy of results are still not enough. To this end, we propose a detail-preserving fine localization mechanism, which not only achieves a progressive fusion of high-level features through prior correction and channel association but also considers high-resolution features to refine the boundaries of camouflaged objects, as shown in Figure 2.
To achieve the above goals, we first design a correction fusion module (CFM), which effectively fuses adjacent layer features and a coarse camouflaged mask to produce fine output. The module includes three inputs: the current and previous layer features and , and the coarse mask . In addition, we first reduce the number of input feature channels to , denoted as and , which helps to improve computational efficiency while still retaining relevant information for detection. As shown in Figure 4, our CFM consists of two parts. In order to make full use of the existing prior guidance map , we purify the features of the previous layer and select the features most related to the camouflaged features to participate in the subsequent cross-layer interaction. Mathematically, the feature map is first multiplied with the coarse mask to obtain the output features :
It is well known that high-level features possess very rich channel-aware cues. In order to achieve more sufficient cross-layer feature interaction and effectively transfer the high-level information of the previous layer to the current layer, we design the channel-level association modeling. We perform channel attention by taking the inner product between each pixel point on and , which calculates the similarity between different feature maps in the channel dimension of the same pixel. To further reduce computational complexity, we also employ a convolution that creates a bottleneck structure, thereby compressing the number of output channels. This process can be described as:
where is the matrix multiplication. Then, we learn two weight maps, and , by using two convolution operations on the features . They are further used in the correction of the features of the current layer in a modulation manner. In this way, the final cross-level fusion features can be generated through the residual processing:
In addition to the above-mentioned prior correction and channel-wise association modeling on the high-level features, we also make full use of the high-resolution information of the first layer to supplement the detailed information. Specifically, we use the receptive field block (RFB) module (Liu et al., 2018) and the spatial attention module (Woo et al., 2018) on the first-layer features () to enlarge the receptive field and highlight the important spatial information of the features, and then fuse with the output of the CFM module () to generate the final prediction map:
where and are the receptive field block and the spatial attention module, respectively. represents the convolution layer along with the batch normalization and ReLU.
4. Loss Function
Following (Wei et al., 2020), We compute the weighted binary cross-entropy loss and IoU loss on three COD maps (i.e., , , and ) to form our final loss function:
where , , denotes the loss between the coarse prediction map and ground truth, denotes the loss about the prediction map after the first CFM, and denotes the loss between the final prediction map and ground truth.
Experiment
Datasets. We conduct experiments and evaluate our proposed method on three popular benchmark datasets, i.e., CHAMELEON (Skurowski et al., 2018), CAMO (Le et al., 2019), and COD10K (Fan et al., 2020a). CHAMELEON (Skurowski et al., 2018) dataset has 76 images. CAMO (Le et al., 2019) contains 1,250 camouflaged images covering different categories, which are divided into 1,000 training images and 250 testing images, respectively. As the largest benchmark dataset currently, COD10K (Fan et al., 2020a) includes 5,066 images in total, 3,040 images are chosen for training and 2,026 images are used for testing. There are five concealed super-classes (i.e., terrestrial, atmobios, aquatic, amphibian, other) and 69 sub-classes. And the pixel-level ground-truth annotations of each image in these three datasets are provided. Besides, for a fair comparison, we follow the same training strategy of previous works (Zhong et al., 2022), our training set includes 3,040 images from COD10K datasets and 1,000 images from the CAMO dataset.
Evaluation Metrics. We use four widely used and standard metrics to evaluate the proposed method, i.e., structure-measure (Fan et al., 2017), mean E-measure (Fan et al., 2018), weighted F-measure (Margolin et al., 2014), and mean absolute error (Li et al., 2020; Zhang et al., 2020; Li et al., 2021). Overall, a better COD method has larger , , and scores, but a smaller score.
Implementation Details. In this paper, we propose a frequency-perception network (FPNet) to address the challenge of camouflaged object detection by incorporating both RGB and frequency domains. Specifically, a frequency-perception module is proposed to automatically separate frequency information leading the model to a good coarse mask at the first stage. Then, a detail-preserving fine localization module equipped with a correction fusion module is explored to refine the coarse prediction map. Comprehensive comparisons and ablation studies on three benchmark COD datasets have validated the effectiveness of the proposed FPNet. The proposed method is implemented with PyTorch and leverages Pyramid Vision Transformer (Wang et al., 2021) pre-trained on ImageNet (Krizhevsky et al., 2017) as our backbone network. We also implement our network by using the MindSpore Lite toolhttps://www.mindspore.cn/. To update the network parameters, we use the Adam optimizer, which is widely used in transformer-based networks (Wang et al., 2021, 2022; Liu et al., 2021a). The initial learning rate is set to 1e-4 and weight decay is adjusted to 1e-4. Furthermore, we resize the input images to , the model is trained with a mini-batch size of 4 for 100 epochs on an NVIDIA 2080Ti GPU. We augment the training data by applying techniques such as random flipping, random cropping, and so on.
2. Comparison with State-of-the-art Methods
We conduct a comparison of our proposed method with 12 state-of-the-art mthods, including FPN (Lin et al., 2016), MaskRCNN (He et al., 2017), CPD (Wu et al., 2019), SINet (Fan et al., 2020a), LSR (Lv et al., 2021), PraNet (Fan et al., 2020b), C2FNet (Sun et al., 2021), UGTR (Yang et al., 2021), PFNet (Mei et al., 2021), ZoomNet (Pang et al., 2022), SINet-V2 (Deng-Ping et al., 2022), and FreNet (Zhong et al., 2022). The visualization comparisons and quantitative results are shown in Figure 5, and Table 1 summarizes the quantitative results of the COD methods on three benchmark datasets.
Quantitative Evaluation. Table 1 presents a detailed comparison of evaluation metrics, we can observe that our proposed model (FPNet) outperforms all SOTA models on all datasets. For example, our FPNet achieves obvious performance gains over other state-of-the-art ones on the CAMO-Test dataset. According to Table 1, our proposed FPNet achieves the best weighted F-measure score of 0.806 on the CAMO-Test dataset, and the MAE score outperforms the second-best method ZoomNet (Pang et al., 2022) by 15.2%. Moreover, the proposed FPNet outperforms ZoomNet (Pang et al., 2022) by an obvious margin in terms of the on all three datasets. For example, compared with the second-best method, the percentage gain of the reach 2.6%, 7.2%, and 1.3% on the COD10K-Test, CAMO-Test and CHAMELEON datasets, respectively. While we observe the frequency-guided method FreNet (Zhong et al., 2022), it can achieve better performance than most state-of-the-art methods. However, our proposed FPNet outperforms FreNet comprehensively in terms of all evaluation metrics, indicating that the proposed learnable frequency-guided solution is superior in discerning discriminative cues of camouflaged objects.
Qualitative Evaluation. As shown in Figure 5, whether the camouflaged object in the image is terrestrial or aquatic, or a camouflaged human, the proposed FPNet method is capable of accurately predicting the region of the camouflaged object. When the camouflaged object is extremely similar to the background, as illustrated in the first row in Figure 5, other SOTA methods fail to accurately distinguish the camouflaged object, especially on the edge regions. By contrast, the proposed FPNet, benefiting from a frequency-aware learning mechanism, can clearly predict the mask of objects with clear and sharp boundaries. When tackling complex background interference, including the salient but non-camouflaged objects (see the third row of Figure 5), our proposed FPNet is capable of effectively separating the camouflaged object from the background, with a more complete and clear structure description ability. For the approximate appearance of similar objects, as shown in the fourth row of Figure 5, the camouflaged human face is hard to distinguish from other pineapples. Most methods fail to recognize it, but our proposed FPNet can discern it clearly. Additionally, our proposed FPNet is also effective in detecting some challenging situations (as displayed in Figure 1), such as indefinable boundaries, multiple objects and occluded objects. The impressive prediction results further highlight the usefulness of the frequency-perception mechanism which connects RGB-aware and frequency-aware clues together to arrive at a unified solution that can adaptly address challenging scenarios.
3. Ablation Studies
To verify the effectiveness of the proposed network, we separate FPNet into a series of ablation parts, i.e., frequency perception module, high-resolution preserving, and correction fusion module, where ‘baseline’ is the PVT backbone for camouflaged object detection. The comparison results are shown in Table 2.
Effectiveness of Frequency Perception Module. The proposed frequency perception module incorporates Octave convolution (Chen et al., 2019) that minimizes spatial redundancy and automatically learns high-frequency and low-frequency features. As shown in Table 2, if we add the frequency perception module (i.e., baseline+FPM), all metrics can obtain performance gains compared with the PVT-alone without the octave convolution. The good performance lies that FPM learns rich frequency-aware information especially the high-frequency clues that are useful for camouflaged object coarse positioning. The other advantage of the FPM lies in that it is automatically online learning without any other extra offline operations. Thus, the flexibility and the high performance of the FPM make it suitable for accurately detecting camouflaged objects in real-world scenes. And FPM can also be easily integrated into other frameworks to assist in distinguishing the obscure boundaries of objects that are similar to the background.
Effectiveness of High-resolution Preserving Module. Although The first frequency-guided coarse positioning stage (i.e., PVT+FPM) has achieved good target prediction maps, the object boundaries are still unsatisfactory. Thus, we adopt the high-resolution preserving mechanism for further detail refining. As shown in Table 2, we conduct the detail-preserving fine localization stage without the correction fusion module (CFM) upon the first coarse positioning stage, i.e., PVT+FPM+High-res Preserving. If we introduce the low-level RGB-aware feature with high resolution to guide the refining process, we can find that the network outperforms the one PVT+FPM. The reason why we need a high-resolution preserving module for fine localization lies in two aspects, i.e., 1) the scales of camouflaged objects are various, and 2) the boundaries of camouflaged objects are usually meticulous which are hard to discern through the high-level semantic features. Inspired by the human visual perception system, humans usually need to zoom in on subtle details in a clear, high-resolution view to recognize camouflaged objects. If the scale is small in the image, we need to leverage the low-level edge-aware or shape-aware information to help the network obtain a fine localization. For the obscure boundary problem, multi-scale features fused in a step-by-step manner will give more help to the boundary separating from the complex background. Thus, we design the refining mechanism to integrate the high-resolution information and gradually fuse deep features together to solve these problems. The experimental results also show that the high-resolution preserving part can provide more performance gains for detail refining. And we can conclude that the detail refinement strategy is not only significant but also effective in localizing the camouflaged objects.
Effectiveness of Correction Fusion Module. Though the high-resolution preserving mechanism for detail refining achieves good performance, the coarse camouflaged mask from the frequency-guided stage is still not effectively exploited enough. Thus, we propose a correction fusion module to further improve the quality of the camouflaged mask by completely mining the ability of the coarse map and the neighbor features. Specifically, we implement the CFM on the detail-preserving fine localization stage, the results are shown in the last row of Table 2. While we update the detail-preserving with the CFM mechanism, all metric scores can be further improved especially in terms of the and scores. The good target detection ability indicates that CFM plays an essential role in improving the detection performance of camouflaged objects. The main reason lies that CFM takes the prior coarse camouflaged mask and the neighbor layer interaction into account. First, the prior coarse prediction mask can provide us with an accurate region of the highlighted camouflaged objects which can extract object-centric related features well. Then, the channel-wise correlation correlates and combines neighbor layers to enhance the object representation which is more distinguishable for perceiving the camouflaged objects. Since CFM learns the channel correlation between adjacent features to obtain learnable weight maps and adjust original features, the dynamic mechanism achieves superior performance compared to the simple concatenation method (the third-row result of Table 2). The good performance reflects that progressively fusing the prior coarse mask and cross-layer interaction is beneficial for camouflaged object refining.
In summary, the frequency-guided coarse positioning stage mainly highlights the important regions of the camouflaged objects under the guidance of hierarchy frequency-aware semantic information, and the detail-preserving fine localization stage further assists in separating the camouflaged objects from the obscure boundaries of the complex background by integrating the high-resolution clue, adjacent correlation features, and the coarse prediction mask. Finally, the proposed FPNet leads us to an accurate and effective solution for detecting camouflaged objects.
3.2. Detail Analysis of the Frequency-aware Information.
In order to verify the effectiveness of the frequency perception mechanism, we analyze different frequency fusing types through quantitative results, as shown in Table 3. We also provide some visualization comparison results from the prediction mask and the learned frequency features, as shown in Figures 6 and 7.
The proposed frequency perception module can automatically separate the features into high-frequency and low-frequency related features. However, how to choose a suitable way to integrate the prominent frequency-aware features to help obtain satisfactory camouflaged object masks needs further discussion. To verify it, we design different comparisons, i.e., only using high-frequency or low-frequency branches for following camouflaged object mask prediction. The detailed comparisons on the COD10k-Test dataset are shown in Table 3. We can observe that the only high-frequency method gives more help than the low-frequency method. All metrics scores of only high-frequency outperform the low-frequency to a great extent. This reflects that high-frequency information is more robust and distinguishable for recognizing camouflaged objects. It also meets the human visual system, we usually employ high-frequency clues to discern the target object from the uncertain region. However, since the octave convolution is an unsupervised operation, that is, no frequency-labeled maps by humans are used for optimization. Thus, some features learned from the low-frequency branch may be useful for camouflaged object detection. Moreover, we adopt a simple addition operation fusing the high-frequency and low-frequency together, the result is shown in the last row of Table 3. The simple addition of high and low frequencies achieves the best performance over only single-frequency ones. Based on these observations, we suggest combining the high-frequency and low-frequency into a single addition to obtain further improvements.
In Figure 6, we analyze the influence of different frequency-aware feature types. In particular, the prediction masks of only high-frequency, only low-frequency, and ours (high-frequency+low-frequency) are shown. The high-frequency method can predict the key part of the camouflaged object, the low-frequency method can obtain an intact region but with some interference background regions. Our proposed method can obtain an accurate object mask compared with the high-frequency or low-frequency ones. The comparison results indicate that the frequency features are meaningful for camouflaged object detection. And fusing the high-frequency and low-frequency will further assist the model in obtaining a relatively complete object mask.
We also visualize the learned frequency-aware features via the octave convolution to further explain the effectiveness of the proposed frequency perception mechanism, as shown in Figure 7. First, our proposed frequency perception mechanism can automatically separate the frequency features into high and low frequency groups without any frequency supervision information. Second, we can observe that the high-frequency and low-frequency groups in the learning process of octave convolution extract the edge information and the main part of the image, respectively. The low-frequency group (Figure 7(c)) focuses more on the overall composition of the image, while the high-frequency group (Figure 7(d)) portrays the edge part of the camouflaged object in the image. While combing the low-frequency and high-frequency groups (Figure 7(e)), our model can focus on the crucial regions of the camouflaged object despite it is similar to the surrounding region.
In conclusion, the proposed frequency perception network has been verified by analyzing the qualitative and quantitative comparison results that the frequency information can give more help to camouflaged object detection. And the proposed frequency perception module can be plugged and played into arbitrary frameworks.
Conclusion
In this paper, we propose a frequency-perception network (FPNet) to address the challenge of camouflaged object detection by incorporating both RGB and frequency domains. Specifically, a frequency-perception module is proposed to automatically separate frequency information leading the model to a good coarse mask at the first stage. Then, a detail-preserving fine localization module equipped with a correction fusion module is explored to refine the coarse prediction map. Comprehensive comparisons and ablation studies on three benchmark COD datasets have validated the effectiveness of the proposed FPNet. This work will benefit more sophisticated algorithms exploiting frequency clues pursuing appropriate solutions in various areas of the multimedia community. In addition, the long-tail problem also exists in COD, this motivates us to explore reasonable solutions referring to the typical methods of long-tail recognition (Yang et al., 2022, 2023).
References
Appendix A Appendix
In Figure 8, we provide more visual examples of different methods. We can see that our proposed network is still competitive in challenging and difficult scenarios, such as multiple objects, fine objects and complex background distractions.
When there are multiple camouflaged objects, as shown in the last two rows of Figure 8, other SOTA methods either cannot identify all the camouflaged objects well or the boundaary of the recognized object are not clear. In contrast, our proposed FPNet network can predict all the camouflaged objects with clear and sharp boundary.
The detection results of our model have a clear advantage in describing the fine details of camouflaged objects. For example, the camouflaged object in the fifth image of Figure 8 has many small, thin, burr-like structures. Compared with other methods, only our method not only accurately detects the camouflaged object, but also fully characterizes these trivial details. Similar advantages are also reflected in the situation that the camouflaged target is occluded. For example, in the images seventh, eighth, eleventh rows of Figure 8, the camouflaged object is partially occluded, but our method is still robust to this case. It is worth mentioning that we also accurately exclude non-camouflaged occluded object regions from the prediction results.
Furthermore, our method is also able to handle challenging complex background scenes, such as the first, second, sixth, ninth and tenth rows of Figure 8. Taking the sixth row image as an example, we should detect a ghost pipefish from the image. Other methods treat the indistinguishable shadows of the ghost pipefish as camouflaged object, while only our method can effectively detect the ghost pipefish and eliminate interference from shadows.
A.2. Visualization of Ablation Studies
We also supplement the visualization results of the ablation experiment in Figure 9. Taking the first image as an example, our baseline model (i.e., Figure 9(f)) can only roughly determine the main part of the camouflaged object, and there is still much room for improvement in terms of details and accuracy. Further, after the introduction of the frequency-perception module in the baseline model, the left shoulder area of the camouflaged human has been significantly improved, but the problem of leg integrity remains unresolved. Then, we add a high-resolution preserving design to our network, which makes the leg details more complete but introduces some noise. Finally, using our designed correction fusion module in our network allows us to achieve the best performance with accurate location, complete structure, and sharp boundary.