Frequency-Spatial Entanglement Learning for Camouflaged Object Detection
Yanguang Sun, Chunyan Xu, Jian Yang, Hanyu Xuan, Lei Luo
Introduction
“Camouflage” is a natural defense mechanism used by certain animals, such as chameleons, grasshoppers, and caterpillars, to blend into their surroundings and protect themselves. The study of camouflaged object detection (COD) focuses on identifying concealed targets in real-world situations. This research is crucial in developing robust visual perception models in computer vision. COD has a wide range of applications, including medical image analysis , species conservation , and industrial defect detection .
In the early stages, some COD methods relied on manually crafted features to detect camouflaged objects. However, due to the extremely challenging appearance of these objects, the results obtained were often unsatisfactory. Later, with the advancement of deep learning and the availability of large-scale datasets , many COD methods based on deep learning have been proposed. With a large amount of data available for training, these methods have the potential to automatically extract features and detect camouflaged objects, resulting in impressive performance. More recently, researchers have introduced various techniques for exploring useful information of input features from the spatial domain using boundary-guided , multi-scale strategy , uncertainty-aware , distraction mining , and .
We have observed that the COD methods mentioned above primarily focus on single spatial features. While these spatial features are advantageous for COD tasks, they are often susceptible to interference from complex backgrounds. This vulnerability arises from their reliance on pixel-level information, with a primary emphasis on the local intensity and spatial position of individual pixels. Furthermore, spatial features possess local properties, meaning that pixels within a feature may only exhibit certain correlations with surrounding pixels. That being said, relying solely on spatial features can make it challenging to distinguish subtle variations within concealed objects and backgrounds. Therefore, it is crucial to find ways to overcome the limitations of spatial features to achieve accurate COD results. Recently, frequency features generated through Fourier Transform have been shown to have global characteristics and have been proven to be beneficial for understanding image contents . This can help break the bottleneck of spatial features.
Some recent COD methods have begun to incorporate frequency clues in their approach. These methods can be divided into two categories based on their objectives. The first category (, FDNet and EVP ) is to act directly on input images through different frequency transforms to extract frequency features, which are then combined with spatial features. However, camouflaged images often contain a lot of background noise, making the frequency features obtained from the image unreliable. When aggregated with spatial features, this may introduce some unnecessary background noises, resulting in under-segmented results (as depicted EVP in Fig. 1). The second category focuses on initial features from the encoder. For example, Cong designed a frequency-perception module to improve the detection of camouflaged objects by utilizing both high-frequency features and low-frequency features. And He proposed frequency attention modules that obtain important parts of corresponding features by considering both high-frequency and low-frequency components. Although these methods have shown promising results, they only focus on high-frequency and low-frequency features, overlooking some information that falls between these two frequencies. This can be seen in FPNet and FEDER in Fig. 1, where significant information within the frequency domain may be missed.
Based on the above discussion, we propose a novel method called Frequency-Spatial Entanglement Learning for accurate camouflaged object detection. Our method combines global frequency features and local spatial features to optimize the initial input features and enhance their discriminative ability. Specifically, we first establish a Frequency Self-attention to obtain discriminative global frequency features, which models the correlation between each frequency band and learns the dependency relationships between different frequencies in input bands. Moreover, we introduce entanglement learning between the frequency and spatial features in the Entanglement Transformer Block, allowing them to mutually learn and collaborate for optimization. Furthermore, we extend the applicability of global frequency features by utilizing the Joint Domain Perception Module and the Dual-domain Reverse Parser to optimize the input features and generate powerful representations that incorporate both frequency and spatial information. Extensive experiments on three widely-used benchmark datasets ( CAMO , COD10K , and NC4K ) demonstrate that FSEL consistently outperforms 21 state-of-the-art COD methods across different backbones.
The main contributions can be summarized as follows:
(1) We propose a Frequency-Spatial Entanglement Learning (FSEL) framework that utilizes both global frequency and local spatial features to enhance the detection of camouflaged objects.
(2) To improve the representation capability of frequency and spatial features, we have designed an Entanglement Transformer Block (ETB). This block allows for entanglement learning of frequency-spatial features, resulting in a more comprehensive understanding of the data.
(3) To reduce the sensitivity and locality limitation of spatial features, we have incorporated frequency domain transformations into both the Joint Domain Perception Module (JDPM) and the Dual-domain Reverse Parser (DRP).
Related work
Camouflaged Object Detection. Recently, with the public availability of datasets (, CAMO , COD10K , and NC4K ), deep learning-based COD methods have started to surface in large numbers, which can be broadly categorized into, including multi-scale strategies , edge-guidance , uncertainty-aware , multi-graph learning , iterative manner , transformer , and so on. Besides, other COD methods considered frequency clues to help with reasoning camouflaged objects. Particularly, Zhong processed directly the camouflaged image through discrete cosine transform to obtain frequency information. He . proposed frequency attention modules to filter out the noteworthy parts of corresponding features. After that, Cong designed a frequency-perception module by learning different frequency features to achieve coarse localization of camouflaged objects. However, these methods often focus on the high- and low-frequency information, ignoring the relationship between all bands in the frequency domain. Therefore, we conduct frequency analysis on different spectral features to achieve all frequency band interactions and importance allocation. Furthermore, we perform entanglement learning on global frequency and local spatial features, which is beneficial for obtaining powerful representations.
Vision Transformer. Transformer utilized self-attention to model global semantic and long-range dependencies, which helps to understand the correlation between different regions in an image, and therefore it has been widely used in some computer vision tasks, including object detection , image classification , semantic segmentation , . For example, Yuan obtained long-range relationships from the sequence of image patches to perform the image classification. Next, Liu split input maps into non-overlapping local windows, and then transferred the information through shift operations between the windows to improve the efficiency of the model. In addition, other transformer models have been successful in computer visions, such as Restormer , CrossFormer , EfficientViT , MPFormer , and among others. Unlike these methods, which always model relationships based on the spatial domain, we transform spatial features into the frequency domain and combine them to perform dual-domain feature optimization.
Frequency Learning. The frequency domain is very important for signal analysis, and recently it has been gradually applied in computer vision tasks. Particularly, Qin assumed channel attention as a compression problem and introduced frequency transformation in the channel attention. Yun handled the balancing problem of different frequency components of visual features. Wang proposed a frequency shortcut perspective in image classification. In addition, some frequency domain-based methods have achieved great performance. In this paper, we extend global frequency features to different applications, involving multi-receptive fields perception, transformer, and reverse attention.
Method
Camouflaged objects exhibit a high level of visual similarity to their backgrounds, achieved through adaptive changes in color, texture, and shape. This creates challenges in distinguishing between object and background pixels in the spatial domain. Additionally, the locality of features in the spatial domain is limited in understanding camouflaged objects. To address this issue, we have implemented several strategies: 1) We have expanded beyond the spatial domain and utilized Fourier transformation to map features to the frequency domain, allowing for a more global perspective; 2) We have analyzed the relationships between all frequency bands to combine global frequency features with local spatial features; 3) We have extended frequency features to multiple components to fully utilize the global understanding of the object.
2 Joint Domain Perception Module
Multi-scale information is beneficial for contextual understanding in different regions. We observe that these methods often generate multi-scale features through different convolutions with multiple receptive fields in the spatial domain. However, the receptive field of convolution operations in the spatial domain is limited, and in the process of data processing, tiny fluctuations may be overlooked, resulting in sub-optimized outcomes.
Therefore, we propose a Joint Domain Perception Module (JDPM) that reconstructs multi-receptive field information by introducing frequency transformation in multi-scale features. As depicted in Fig. 3, our JDRM uses the hierarchical structure to extract frequency-spatial information of different receptive fields. Technically, we use feature as input and first reduce its channel numbers using 11 convolution (), , . Then, we construct a set of 33 atrous convolutions () with filling rate to capture local multi-scale spatial feature , , , where and . Next, we transform local spatial features into the frequency domain using the Fast Fourier Transform () and perform redundancy filtering. We then use the Inverse Fast Fourier Transform () and the modulus of complex features to obtain global frequency features , which is defined as:
where and “” denote the modulus operation and the element-wise multiplication. presents a set of weight coefficients, which sequentially contains a convolution, a batch normalization, a ReLU, a convolution, and a sigmoid function. and denote frequency domain coordinates and spatial domain coordinates, represents the imaginary part. After that, we aggregate global frequency features with local spatial features to generate intermediate multi-scale features , that is, .
Finally, we concatenate all multi-scale features and introduce residual connections to generate a coarse feature map with 1-channel through a 3 3 and a 1 1 convolutions, which can be expressed as:
where presents convolution. and “+” denote concatenation and element-wise addition.
3 Entanglement Transformer Block
Unlike previous methods , which only model long-range dependencies based on local features in the spatial domain, our ETB incorporates different relationships from the frequency and spatial domains. In addition, we propose entanglement learning for different domain features in the ETB, allowing for the integration of information such as color, texture, edge, spectral, amplitude, and energy. This approach is beneficial for learning discriminative representations by considering various types of information. As depicted in Fig. 4, our ETB consists of three key components: frequency self-attention (FSA), spatial self-attention (SSA), and entanglement feed-forward network (EFFN).
where presents a combination function that combines the imaginary and real parts into a complex number. denotes a Softmax function. Subsequently, we use the attention map to optimize the weights on the frequency feature and then use the Inverse Fast Fourier Transform () to convert it to an original domain and employ the modulus operation to obtain the frequency attention feature. In addition, we introduce a frequency residual connection to increase frequency information (, =), and finally fuse features to produce the frequency feature , which is formulated as:
where and are concatenation and matrix multiplication. presents the modulus operation. is the reshaped .
Spatial self-attention. Considering the unfixed size of camouflaged objects, we embed abundant contextual information into spatial self-attention. As shown in the bottom right of Fig. 4, similar to the FSA, we take the feature as the input and encode the position information using a 11 convolution (), and then we obtain the , , and required by the self-attention by utilizing two depth-wise separable convolution with 33 () and 55 (). After that, we generate the attention map (, () ) through the reconstructed and and activate it using Softmax function. Subsequently, the activated attention map is used to correct the weights of . Besides, to increase the spatial local information (, ), we perform a residual connection to generate spatial feature , as shown in:
where , , and are the same as in Eq. (4). is the reshaped .
Entanglement feed-forward network. Frequency and spatial features usually contain different information. The frequency domain focuses on the global energy distribution and variation of signals, while spatial information acts on local pixel-level details and spatial structures, all of which are crucial for comprehending camouflaged objects. In our EFFN, these features are considered as two kinds of states that can perform entanglement learning to obtain more robust and powerful representations during the entanglement process.
Specifically, we first entangle the global frequency feature and the local spatial feature to adapt them to each other, followed by the residual connection to acquire the comprehensive feature , that is, , which performs the layer normalization to improve the stability, and then the normalized feature (=) is subjected to non-linearity entanglement learning in the EFFN. Technically, the EFFN consists of two phases, the first stage projects the feature to the frequency and spatial domains, and utilizes the GELU function for nonlinear activation and a gate mechanism to obtain global frequency feature and local spatial feature , which can be written as follows:
where denotes the GELU function. Subsequently, in the second stage, the frequency and spatial features from the first stage are again entangled by interacting with each other by transferring information from different domains, and the entangled frequency-spatial features are optimized independently. They are then aggregated and reduced channels to generate comprehensive feature , which can be formulated as follows:
where , and are the same as in Eq. (4). and present the Fast Fourier Transform and the Inverse Fast Fourier Transform. denotes the depth-wise separable convolution with 33 kernel. Finally, we introduce residual connections to obtain the final feature with 128-channel in the ETB, , =. Through multiple aggregation interactions, global frequency and local spatial features interact and depend on each other, leading to the entanglement of features from different states, forming rich and comprehensive representations.
4 Dual-domain Reverse Parser
Different from these methods that integrate multi-level features based on the spatial domain, we propose the dual-domain reverse parser (DRP), which optimizes and aggregates diverse information from multi-level feature in both frequency and spatial domains. As depicted in Fig. 2, we first take the feature from the ETB as the optimization objective and use the higher-level semantic feature as the auxiliary objective.
The DRP consists of two branches (as shown in Fig. 5), in the first branch, we first expand the channel of auxiliary feature to match the dimension of the optimization objective and aggregate these feature to obtain feature , , , where denotes to expand the channel to 128, presents a 11 convolution and two 33 convolutions. And then feature is separated into the spatial and frequency domains. We perform the Fast Fourier Transform () and Inverse Fast Fourier Transform () in the frequency domain and adopt a series of convolution operations () to optimize features in the spatial domain. Subsequently, they are aggregated to obtain the fused feature , that is,
where denotes the modulus operation. “” presents element-wise addition. In the second branch, we produce the hybrid reverse attention map (, ) ) using the auxiliary feature, where denotes the Sigmoid function. Unlike other methods , our the reverse attention map () contains abundant the frequency-spatial information to efficiently obtain the reverse feature , , . Next, we integrate the features and to generate final feature , that is, +. Subsequently, will continue to optimize features as an auxiliary objective in the proposed DRP. Note that there must be at least one auxiliary feature used for optimizing feature to generate feature , and auxiliary features are input in the dense connection manner.
5 Loss function
In the proposed FSEL method, we supervise multi-level feature to produce an accurately predicted map. Specifically, we adopt the weighted binary cross-entropy (BCE) and the weighted intersection over union (IoU) as the overall loss function to optimize the model based on ground truth (). The loss function can be defined as:
where and denote the weighted BCE and IoU functions. is the feature from the JDPM.
Experiment
Datasets. We evaluate our FSEL model on three benchmark datasets: CAMO , COD10K , and NC4K . CAMO is an early dataset that contains 1,250 camouflaged images with 1,000 training images and 250 testing images. COD10K is a currently large dataset of camouflaged objects, consisting of 3,040 training images and 2,026 testing images. NC4K is the largest COD dataset for testing, containing 4,121 images of camouflaged objects. We use 4,040 images from CAMO and COD10K as training samples to train the FSEL.
Implementation details. The proposed FSEL model is implemented in the PyTorch framework on four NVIDIA GTX 4090 GPUs with 24GB. We utilize the pre-trained PVTv2 /ResNet50 /Res2Net as the encoder to extract initial features. Following , we also employ data augmentation techniques such as random flipping and random clipping to enhance training data. We use the Adam optimizer with an initial learning rate of 1e-4 and decay the rates by 10 every 60 epochs. All input images are resized to 416416, and the batch size is set to 40 for 180 epochs of training progressing.
Evaluation metrics. We use six well-known evaluation metrics, including Mean Absolute Error (), Maximum F-measure (), Average F-measure (), Weighted F-measure (), S-measure (), and E-measure ().
2 Comparisons with the SOTAs
We conduct a comparison of our FSEL with twenty-one COD methods, including SINet , C2FNet , UGTR , JSOCOD , MGL-S , LSR , PFNet , VST , FAPNet , SINetv2 , BSANet , SegMaR , ZoomNet , BGNet , PreyNet , FEDER , EVP , FPNet , HitNet , FSPNet , and SAM . Note that the predicted maps from all methods are provided by the authors or obtained from open-source codes.
Quantitative Evaluation. Table 1 summarizes the quantitative result of our FSEL and other 21 SOTA models. From Table 1, we can observe that the FSEL model achieves excellent performance across different backbone networks. Particularly, compared to the recently proposed FEDER method, our FSEL with ResNet50 backbone overall surpasses 5.97%, 3.23%, and 4.76% on three public datasets under the metric. Besides, in the Res2Net backbone, our FSEL achieves average performance gains of 9.23%, 2.30%, 2.16%, 2.94%, 1.34%, and 1.24% over the second-best method in terms of six public evaluation metrics on CAMO dataset. Moreover, compared to the frequency-based FPNet and EVP methods, FSEL method with PVTv2 backbone achieves average performance gains of 38.10%, 4.41%, 4.05%, 5.96%, 3.07%, and 2.09% over FPNet and 52.38%, 6.36%, 12.43%, 10.19%, 4.55%, and 5.82% over EVP in terms of , , , , , and on the COD10K dataset. Furthermore, FSEL achieves excellent performance when adopting the same input strategy with ZoomNet . The superiority in performance benefits from the joint optimization of the ETB, JDPM, and DRP for input features in the frequency and spatial domains. In addition, we provide the parameters and FLOPs in Table 2. It can be seen that the proposed FSEL method parameters and FLOPs are at a medium to high level, however, our performance far exceeds that of methods with similar parameters and FLOPs.
Qualitative Evalation. Fig. 6 gives the visual comparisons between our FSEL and several COD method in different scenarios. As depicted in Fig. 6, the proposed FSEL method exhibits accurate and complete segmentation for camouflaged objects with different sizes compared to current COD methods (, HitNet , FSPNet , and FPNet ). These visual results demonstrate the superiority of the FSEL method for detecting camouflaged objects through the frequency-spatial domain optimization strategy.
3 Ablation Study
Effectiveness of proposed each component. We provide the quantitative results of different components in the proposed FSEL model, shown in Table 3. Specifically, we first adopt “ResNet50 - FPN ” as “Baseline” (Tab. 3(a)) to detect camouflaged objects. And then we independently validate the effectiveness of “ETB” (Tab. 3(b)), “DRP” (Tab. 3(c)) and “JDPM” (Tab. 3(d)), and it can be seen that the performance of the predicted map increases significantly when the proposed component is embedded in the “Baseline” (Tab. 3(a)). Additionally, we validate the compatibility among all modules. From Tab. 3(e), Tab. 3(f), and Tab. 3(g), it can be observed that the three components are compatible with each other. Subsequently, all components are integrated, and the performance of the model is improved once again, as shown in Tab. 3(h). Additionally, in Fig. 7, we show the visual results obtained by progressively adding the proposed components (, ETB, JDPM, and DRP), generating that the predicted map gradually approaches the ground truth (GT). The above results demonstrate the effectiveness of our proposed modules in detecting camouflaged objects.
Effectiveness of frequency-spatial information within the ETB. Do we really need frequency information? To answer this question, we perform a series of experiments in the internal part of the ETB. Specifically, the ETB is first divided into two parts, with “ETB-S” (Tab. 4(a)) containing only spatial information, and “ETB-F” (Tab. 4(b)) presenting that it includes only frequency information. Based on Table 4, the performance of the separate frequency and spatial domains exhibits certain differences compared to the complete ETB (Tab. 4(g)). Besides, we investigate the entanglement learning of frequency-spatial information in the proposed ETB. In Tab. 4(c)-(f), it can be seen that the frequency and spatial features interact fusion to achieve entanglement between two states, enhancing the model’s reasoning ability of camouflaged objects.
4 Expanded application
To demonstrate the generalization ability of our FSEL model, we extend the FSEL model to salient object detection and polyp segmentation tasks. As shown in Fig. 8, the proposed FSEL method achieves highly accurate segmentation for both salient objects and polyps, benefiting from the complementary utilization of frequency domain and spatial information. More details and data are presented in the supplementary materials.
Conclusion
In this paper, we introduce a new approach for detecting camouflaged objects called Frequency-Spatial Entanglement Learning (FSEL). The key to FSEL is to extract important information from both the frequency and spatial domains. To achieve this, we have developed a Joint Domain Perception Module that combines multi-scale information from frequency-spatial features to accurately localize regions. Additionally, we have created an Entanglement Transformer Block that can be easily integrated into existing methods to improve their performance by modeling long-range dependencies in the hybrid domain. Furthermore, we have designed a Dual-Domain Reverse Parser that interacts with diverse information in multi-layer features to achieve more precise segmentation. Our extensive comparison experiments demonstrate that FSEL outperforms 21 state-of-the-art COD methods on three popular benchmark datasets.
This work was supported in part by the National Science Fund of China (No. 62276135, 62361166670, 62372238, and 62302006).