Salient Object Detection via Integrity Learning

Mingchen Zhuge, Deng-Ping Fan, Nian Liu, Dingwen Zhang, Dong Xu, Ling Shao

Introduction

Salient object detection (SOD) aims to imitate the human visual perception system to capture the most significant regions in a given image . As SOD is widely used in the field of computer vision, it plays a vital role in many downstream tasks, such as object detection , image retrieval , co-salient object detection , multi-modal matching , VR/AR applications and semantic segmentation .

Traditional SOD methods predict saliency maps in a bottom-up manner, and are mainly based on handcrafted features, such as color contrast , boundary backgrounds , or center priors . To improve the representation capacity of the features used in SOD, current models employ convolutional neural network (CNN) or fully convolutional network architectures, which enable powerful feature learning processes to replace manually designed features. These methods have achieved remarkable progress and pushed the performance of SOD to a new level. More details of recent deep learning based SOD methods can be found in the surveys/benchmarks .

The current success in building deep learning based salient object detectors is mainly due to the use of multi-scale/level feature aggregation, contextual modeling, top-down modeling, and edge-guided learning mechanisms. Specifically, models with a multi-scale/level feature aggregation mechanism enhance the features from different levels and scales of the network, and then fuse them to generate the final SOD results. These approaches can help discover salient objects of various sizes and highlight the salient regions under the guidance of both coarse semantics and fine details. For example, the network proposed by Zhang et al. first adaptively fuses multi-level features at five different scales, and then use them to generate predictions. Similarly, Luo et al. proposed to extract the global and local features at the low and high feature scales, respectively, and then fuse them to generate the final results.

Contextual modeling is another key mechanism in SOD. It helps infer the saliency of each local region by considering the surrounding contextual information. Current studies in the field of SOD usually design various attention modules to explore such information. Specifically, Zhao et al. proposed a pyramid feature attention network, where channel attention and spatial attention modules are introduced to process high- and low-level features, respectively, and consider the contextual information in different feature channels and spatial locations. Liu et al. proposed to learn a pixel-wise contextual attention for SOD. Deep models learned with such an attention module can infer the relevant importance between each pixel and its global/local context location, and thus achieve the selective aggregation of contextual information.

For top-down modeling, some SOD methods adopt carefully designed decoders to gradually infer salient regions under the guidance of high-level semantic cues. For example, Wang et al. built an iterative and cooperative inference network for SOD, where multiple top-down network streams work together with the bottom-up network streams in an iterative inference manner. Zhao et al. proposed a gated dual-branch decoding structure to achieve cooperation among different levels of features in the top-down flow, which improves the discriminability of the whole network. In , Liu et al. adopted a pyramid pooling to build global guiding features, to improve the top-down flow modeling.

In order to accurately predict salient object boundaries, another group of methods introduce additional network streams or learned objective functions to force the network to pay more attention to the contours that separate the salient objects from the surrounding background. For example, Wei et al. built a label decoupling framework for SOD, which explicitly decomposes the original saliency map into a body map and a details map. Specifically, the body map concentrates on the central areas of the salient objects, while the details map focuses on the regions around the object boundaries. To improve the prediction precision of the salient contours and reduce the local noise in salient edge predictions, Wu et al. proposed the mutual learning strategies to separately guide the foreground contour and edge detection tasks.

Although the aforementioned mechanisms can improve the SOD performance in several aspects, the detection results produced are still not optimal. In our opinion, this is likely due to the under-exploration of another helpful and important mechanism, i.e., the integrity learning mechanism (see Fig. 1 (a) & (b)). In this work, we define the integrity learning mechanism at two levels. At the micro level, the model should focus on part-whole relevance within a single salient object. At the macro level, the model needs to identify all salient objects within the given image scene. In Fig. 2, we present some examples of the integrity qualities at both the macro and micro levels. It is clear that there exists a strong correlation between integrity and prediction performance.

In order to pursue two-level integrity, we introduce three key components in our deep neural network design. The first is diverse feature aggregation (DFA). Unlike the existing models, which focus more on feature discriminability, DFA aggregates the features from various receptive fields (in terms of both the kernel shape and context) to increase their diversity. Such feature diversity provides the foundation for mining integral salient objects, since it considers richer contextual patterns to determine the activation of each neuron. The second component is called integrity channel enhancement (ICE), which aims at enhancing the feature channels that highlight the integral salient objects (at both the micro and macro levels), while suppressing the other distracting ones. As it is rare for the feature channels enhanced by ICE to perfectly match the real salient object regions, we further adopt a part-whole verification (PWV) component to judge whether the part features and whole features have a strong agreement to form the integral objects. This can help further improve integral learning at the micro level.

It is worth mentioning that some existing works have also tried to solve the macro-level integrity issue by introducing the auxiliary task for learning deep salient object detectors . However, these methods require additional supervision information on the number of salient objects within each image. In contrast, our newly proposed approach can tackle both macro- and micro-level integrity issues within a unified and entirely different learning framework, without requiring any additional supervision.

Our overall framework for integrity learning is called the Integrity Cognition Network (ICON), details of which are shown in Fig. 3. Specifically, our ICON first leverages five convolutional blocks for basic feature extraction. Then, it passes the deep features at each level to a diverse feature aggregation module to extract the different feature bases. Next, the diverse feature bases extracted from three adjacent feature levels are sent to an integrity channel enhancement module. Here, an integrity guiding map is generated and then used to guide the attention weighting of each feature channel. Finally, the integrity channel enhancing features produced from the three feature levels are combined and passed through the part-whole verification module, which is implemented by using the capsule routing layers . After further verifying the agreement between the object parts and whole regions, the missing parts will be reinforced. To sum up, this work includes three main contributions:

We investigate the integrity issue in SOD, which is essential yet under-studied in this field.

We introduce three key components for achieving integral SOD, namely diverse feature aggregation, integrity channel enhancement, and part-whole verification.

We design a novel network, i.e., ICON, that incorporates the three components and demonstrate its effectiveness on seven challenging datasets. In addition to its prominent performance, our approach also achieves real-time speed (∼\sim60fps).

The remainder of the paper is organized as follows. In Section 2, we discuss the related works. Then, we describe the proposed ICON in detail (see Section 3). Experimental results, including performance evaluations and comparison, are provided in Section 4. Finally, conclusions are drawn in Section 5.

Related Work

Over the past several decades, a number of SOD methods have been proposed and have achieved encouraging performance on various benchmark datasets. These existing SOD methods can be roughly categorized into scale learning based, boundary learning based, and integrity learning based approaches.

Scale variation is one of the major challenges for SOD. Many works have tried to handle this issue from different perspectives. Inspired by the HED model for edge detection, DSS introduced deep-to-shallow side-outputs with rich semantic features. This design enables shallow layers to distinguish real salient objects from the background, while retaining high resolution. In addition, Zhang et al. designed a multi-level feature aggregation framework and employed the hierarchical features as the saliency cues for final saliency prediction. Meanwhile, RADF integrated multi-level features and refines them within each layer with a recurrent pattern, which effectively suppresses the non-salient noise in lower layers and increases the salient details of features in higher layers. Further, Zhao et al. proposed to use the F-measure loss, which can generate precise contrastive maps to help segment multi-scale objects. To efficiently extract multi-scale features, Pang et al. embedded self-interaction modules into their decoder units to learn the integrated information. In the more recent work, GateNet adopted Fold-ASPP to gather multi-scale saliency cues. Finally, Liu et al. utilized a centralized information interaction strategy to simultaneously process multi-scale features.

2 Boundary Learning Approaches for SOD

Boundary learning plays another important role for improving SOD results. Early works used boundary learning via biologically inspired methods . However, these models exhibit undesirable blurring results and usually lose entire salient areas. The more recent CNN-based approaches, which operate at the patch level (instead of pixel level), also suffer from blurred edges, due to the stride and pooling operations. To address this issue, several works (e.g., ) use the pre-processing technology (e.g., superpixel ) to preserve the object boundaries, while other works, such as DSS , DCL , and PiCANet , employed post-processing (e.g., conditional random fields ) to enhance edge details. The main drawback of these approaches is their slow inference speed. To learn the intrinsic edge information, PoolNet employed an auxiliary module for edge detection. Besides, many other works have improved edge quality by introducing boundary-aware loss functions. For instance, the recent works used explicit boundary losses to guide the learning of boundary details. Considering that the cross entropy loss prefers to predict hard pixel samples (e.g., 0 or 1) as non-integer values, BASNet introduced a new prediction-refinement network and hybrid loss. Dealing with the inherent defect of blurry boundaries, HRSOD introduced the first high-resolution SOD dataset, which explores how high-resolution data can improve the performance of the salient object edges. F3Net demonstrated that assigning larger weights to boundary pixels in the loss functions is a simple way to handle boundary problems. In addition, the recent works such as SCRN , LDF , VST built two-stream architectures to model salient objects and boundaries simultaneously.

3 Integrity Learning Approaches for SOD

Integrity learning is an under-explored research topic in SOD. Among the limited existing models, DCL processed contrast information at both the pixel and patch levels in order to simultaneously integrate global and local structural information. CPD utilized an effective decoder to summarize the discriminative features, and segment the integral salient objects with the aid of holistic attention modules. TSPOANet modeled part-object relationships in SOD, and produced better wholeness and uniformity scores for segmented salient objects with the help of a capsule network. GCPANet made full use of global context to capture the relationships between multiple salient objects or regions, and alleviate the dilution effect of features. Wu et al. used a bi-stream network combining two feature backbones and gate control units to fuse complementary information. Recently, transformers have become a hot research area in the field of computer vision. Mao et al. proposed a transformer-based architecture for the context learning problem, which can also be considered as an integrity learning based approach.

Framework

As shown in Fig. 3, our method is based on an encoder-decoder architecture. The encoder uses ResNet-50 as the backbone to extract multi-level features. Meanwhile, the decoder integrates these multi-level features and generates the saliency map with multi-layer supervision. For simplicity, from then on we denote the features generated by the backbone as a set Fbkb={Fbkb(0),Fbkb(1),Fbkb(2),Fbkb(3),Fbkb(4)}\mathcal{F}_{bkb}=\{\mathbf{F}^{(0)}_{bkb},\mathbf{F}^{(1)}_{bkb},\mathbf{F}^{(2)}_{bkb},\mathbf{F}^{(3)}_{bkb},\mathbf{F}^{(4)}_{bkb}\}. To improve the computational efficiency, we do not use Fbkb(0)\mathbf{F}^{(0)}_{bkb} in the decoder due to its large spatial size.

Next, we enhance the backbone features by passing them through the diverse feature aggregation (DFA) module, which consists of various convolutional blocks. Thereafter, we further use the integrity channel enhancement (ICE) module to strengthen the responses of the integrity-related channels and coarsely highlight the integral salient parts. Finally, to further refine the saliency map, we utilize the part-whole verification (PWV) module to verify the agreement between object parts and the whole salient region.

2 Diverse Feature Aggregation

Recent works have demonstrated that enriching the receptive fields of the convolution kernel can help the network learn features that capture different object sizes. In this work, we go one step further and incorporate convolution kernels with different shapes to deal with the shape diversity of different objects. Specifically, as shown in Fig. 4-(A), we introduce the novel DFA module to enhance the diversity of the extracted multi-level features by using three kinds of convolutional blocks with different kernel sizes and shapes. Technically, we utilize a practical combination of the asymmetric convolution , atrous convolution , and original convolution operations to capture diverse spatial features. The overall procedure is summarized as follows:

where Fdfa(i)\mathbf{F}^{(i)}_{dfa} denotes the features produced by the above process, X∗\mathcal{X}_{*} denotes different types of blocks (i.e., asymmetric, atrous, original convolutions), and Concat⁡[⋅]\operatorname{Concat}[\cdot] is the concatenation operation.

Note that we use Xatrr\mathcal{X}_{atr}^{r} to denote atrous convolution operations with different dilation rates rr, for example, Xatr2\mathcal{X}_{atr}^{2} is the atrous convolution operation with the dilation rate of 2, and use the asymmetric convolution (Xasy\mathcal{X}_{asy}) operation with a crux-shape . Specially, Xasy\mathcal{X}_{asy} contains three layers, one with a normal 3×33\times 3 square kernel K3×3\textbf{K}_{3\times 3}, one with a horizontal 1×31\times 3 kernel K1×3\textbf{K}_{1\times 3}, and one with a vertical 3×13\times 1 kernel K3×1\textbf{K}_{3\times 1}, and all of them are shared in the same sliding window. It can be described as:

where ⋆\star is the 2D convolutional operator, ⊕\oplus is the element-wise addition, and I denotes the input feature.

In such a way, our DFA module can enrich the feature space by fusing the learned knowledge from the crux kernel, dilated kernel, and normal kernel in the first stage. As a result, DFA can cover different salient regions in various contexts, by enhancing integrity. We mark the features processed by DFA as \mathcal{F}_{dfa}=\{\mathbf{F}^{({\color[rgb]{0,0,0}1})}_{dfa},\mathbf{F}^{({\color[rgb]{0,0,0}2})}_{dfa},\mathbf{F}^{({\color[rgb]{0,0,0}3})}_{dfa},\mathbf{F}^{({\color[rgb]{0,0,0}4})}_{dfa}\}.

3 Integrity Channel Enhancement

Several recent studies have achieved promising visual categorization results by using the spatial or channel attention mechanism. Though these methods are driven by various motivations, they all essentially aim to build the correspondence between different features to highlight the most significant object parts. However, how to mine the integrity information hidden in different channel of features remains under studied. To address this issue, we propose a simple ICE module to further mine the relations within different channels, and enhance the channels that highlight the potential integral targets.

We consider multi-scale information from every three adjacent features. First, we re-scale the next and previous feature levels and use upsampling and downsampling operations to adjust them to the spatial resolution of H×WH\times W. Then, we generate the fusion maps Ffuse(i)\mathbf{F}^{(i)}_{fuse} by concatenating the three input features:

After that, we extract the integrity embedding Iemb(i)\mathbf{I}_{emb}^{(i)} by applying the l2l_{2} norm on Ffuse(i)\mathbf{F}_{fuse}^{(i)}. Next, to further integrate the integrity information, we use a parameter-efficient bottleneck design to learn Iemb(i)\mathbf{I}_{emb}^{(i)}. As the channel transform would slightly increase the difficulty of optimization, we add layer normalization inside two convolution layers (before ReLU) to facilitate optimization, similar to the design used in :

where ⊗\otimes is the element-wise multiplication operation and LN⁡\operatorname{LN} means layer normalization.

By using the proposed ICE module, the channels with better integrity can be effectively enhanced. As can be seen in Fig. 5, after feeding the features into our ICE, the foreground region is noticeably distinguished from the background, and the features produced by ICE tend to highlight the integral objects at both the micro and macro levels. In our implementation, if there are not enough multi-level features input in the first and last levels, we will fill the input with the features from the current level. Besides, we use two ICE modules with shared parameters to help our ICON model integrate cues at multiple levels.

4 Part-Whole Verification

The PWV module aims to enhance the learned integrity features by measuring the agreement between object parts and the whole salient region. To achieve this goal, we adopt a capsule network , which has been proved to be effective in modeling part-whole relationships. Motivated by the success of the prior work SegCaps , we embed the capsule network into our ICON. In PWV, one key issue is how to assign votes from the low-level to the high-level capsules. As the high-level capsules need to form the whole object representation by aggregating the object parts from the relevant low-level capsules, we use EM routing to model the association between the low-level and high-level capsules in a clustering-like manner. The inputs of PWV are three different ICE features (Fice\mathcal{F}_{ice}). Specifically, we first reduce the ICE features at each level to a united resolution, i.e., 22×2222\times 22, in order to reduce the computational costs.

Next, we build our primary capsules. To be specific, we use eight pose vectors to build a pose matrix M, and an activation ϕ∈[0,1]\phi\in\left[0,1\right] to represent each capsule. The pose matrix contains the instantiated parameters to reflect the properties of object parts or the whole object, while the activation represents the existence probability of the object. The capsules from the primary capsule layer pass information to those in the next PWV capsule layer through a routing-by-agreement mechanism. Specifically, when the capsules from a lower layer produce votes for the capsules in a higher level, the votes ωij\omega_{ij} are obtained by a matrix multiplication operation between the learned transformation matrices Tij\textbf{T}_{ij} and the lower-level pose matrix Mi\textbf{M}_{i}, where ii and jj are the indices of the lower- and higher level capsules, respectively. Once these votes are obtained, they are used in the EM routing algorithm to produce the higher-level capsule Cj\textbf{C}_{j} with the pose matrices Mj\textbf{M}_{j} and the activation ϕj\phi_{j}. After that, we obtain the part-whole verified features. Subsequently, element-wise addition and upsampling operations are introduced to fuse these part-whole verified features at adjacent levels in a bottom-up manner, which encourages cooperation among multi-scale features. After the PWV module, the model outputs Fpwv={Fpwv(1),Fpwv(2),Fpwv(3),Fpwv_cap(4)}\mathcal{F}_{pwv}=\{\mathbf{F}^{({1})}_{pwv},\mathbf{F}^{({2})}_{pwv},\mathbf{F}^{({3})}_{pwv},\mathbf{F}^{({4})}_{pwv\_cap}\}.

5 Supervision Strategy

In this work, in addition to the BCE loss, we also use the IoU loss . Specifically, the overall loss of the proposed ICON is formulated as LCPR(P,G)\mathcal{L}_{CPR}\left(\textit{P},\textit{G}\right), where P is the generated saliency prediction map, and G is the ground truth saliency map. LCPR\mathcal{L}_{CPR} incorporates the cooperative BCE loss and IoU loss, i.e., LCPR=LBCE+LIoU\mathcal{L}_{CPR}=\mathcal{L}_{BCE}+\mathcal{L}_{IoU}. Specifically, LBCE\mathcal{L}_{BCE} is formulated as follows:

where W and H are the width and height of the images, respectively. Meanwhile, LIoU\textit{L}_{IoU} is defined as:

where G(x,y)\textit{G}(\textit{x},\textit{y}) and P(x,y)\textit{P}(\textit{x},\textit{y}) are the ground truth label and the predicted saliency label at the location (x,y)(\textit{x},\textit{y}), respectively. During training, we use the multi-level supervision strategy widely used in this field . Apart from using four features from Fpwv\mathcal{F}_{pwv}, we fuse Fpwv(1)\mathbf{F}^{({1})}_{pwv} and Fice(1)\mathbf{F}^{({1})}_{ice} by dot-product, which is used as an extra feature for supervision. This feature is also used to generate the final prediction results during the inference period. To match the ground-truth maps in both training and inference stages, the features’ channel will be reduced to 1-dimension, and the spatial size will be recovered as the same as the input.

Experiments

We train our ICON on the DUTS-TR , which is commonly used for the SOD task and contains 10,553 images. Then, we evaluate all the models on seven popular datasets: ECSSD , HKU-IS , OMRON , PASCAL-S , DUTS-TE , SOD and attribute-based SOC , which are all annotated with pixel-level labels. Specifically, ECSSD is made up of 1,000 images with meaningful semantics. HKU-IS includes 4,447 images, containing multiple foreground objects. OMRON consists of 5,168 images with at least one object. These objects are usually structurally complex. PASCAL-S was built from a dataset originally used for semantic segmentation, and it consists of 850 challenging images. DUTS is a relatively large dataset with two subsets. The 10,553 images in DUTS-TR are used for training, and the 5,019 images in DUTS-TE are employed for testing. SOD includes 300 very challenging images. SOC contains images from complicated scenes, which are more challenging than those in the other six datasets.

2 Implementation Details

We run all experiments on the publicly available Pytorch 1.5.0 platform. An eight-core PC with an Intel Core i7-9700K CPU (with 4.9GHz Turbo boost), 16GB 3000 MHz RAM and an RTX 2080Ti GPU card (with 11GB memory) is used for both training and testing. During network training, each image is first resized to 352×\times352 (for the VGG /ResNet /PVT backbones) or 384×\times384 (for Swin /CycleMLP ), and data augmentation methods such as normalizing, cropping and flipping, are used. Some encoder parameters are initialized from VGG-16, ResNet-50, PVTv2, Swin-B and CycleMLP-B4. We initialize some layers of PWV by zeros or ones, while other convolutional layers are initialized based on . We use the SGD optimizer to train our network, and set its hyperparameters as: initial learning rate lr = 0.05, momen = 0.9, eps = 1e-8, weight_decay = 5e-4. The warm-up and linear decay strategies are used to adjust the learning rate. The batch size is set to 32 (ResNet), 10 (PVTv2/CycleMLP) or 8 (VGG/Swin), and the maximum number of epochs is set to 60 (ResNet-based training takes ∼\sim2.5 hours). In addition, we use apex https://github.com/NVIDIA/apex and fp16 to accelerate the training process. Gradient clipping is also used to prevent gradient explosion. The inference process of the ResNet-based architecture for a 352×\times352 image only takes 0.0164s, including the IO time.

3 Evaluation Metrics

We use five metrics to evaluate our model and the existing state-of-the-art algorithms:

(1) MAE (MM) evaluates the average pixel-wise difference between the predicted saliency map (P) and the ground-truth map (G). We normalize P and G to $,sotheMAEscorecanbecomputedas, so the MAE score can be computed asM=\frac{1}{\textit{W}\times\textit{H}}\sum_{\textit{x}=1}^{\textit{W}}\sum_{\textit{y}=1}^{\textit{H}}|P(\textit{x},\textit{y})-G(\textit{x},\textit{y})|.$

(2) Weighted F-measure (FβωF_{\beta}^{\omega}) offers an intuitive generalization of FβF_{\beta}, and is defined as Fβω=(1+β2)Precisionω⋅Recallωβ2⋅Precisionω+RecallωF_{\beta}^{\omega}=\frac{\left(1+\beta^{2}\right)\text{Precision}^{\omega}\cdot\text{Recall}^{\omega}}{\beta^{2}\cdot\text{Precision}^{\omega}+\text{Recall}^{\omega}}. As a widely adopted metric , FβωF_{\beta}^{\omega} can handle the interpolation, dependency and equal-importance issues, which might cause inaccurate evaluation by MAE and F-measure . As suggested in , we set β2\beta^{2} to 1.0 to emphasize the precision over recall. By assigning different weights (ω\omega) to different errors based on the specific location and neighborhood information, FβωF_{\beta}^{\omega} extends the F-measure to non-binary evaluation.

(4) E-measure (EξE_{\xi}) combines the local pixel values with the image-level mean value in one term and can be computed as: Eξ=1W×H∑x=1W∑y=1Hθ(ξ)E_{\xi}=\frac{1}{W\times H}\sum_{x=1}^{W}\sum_{y=1}^{H}\theta\left(\xi\right), where ξ\xi is the alignment matrix and θ(ξ)\theta\left(\xi\right) indicates the enhanced alignment matrix. We adopt mean E-measure (EξmE_{\xi}^{m}) as our final evaluation metric.

(5) FNR is the false negative ratio, which is computed by:

where FNFN is the pixel-level indicator that determines whether a pixel is a false negative. We show several examples of FNR in Fig. 6. It clearly and accurately reflects the integrity of prediction results and is sensitive at the macro and micro level.

4 Comparison with the SOTAs

We compare the proposed approach with 14 very recent state-of-the-art methods, including Condinst , PointRend , PiCANet , RAS , AFNet , BASNet , CPD , EGNet , SCRN , F3Net , MINet , ITSD , GateNet and VST .

Table I reports the quantitative results on six traditional benchmark datasets, in which we compare our method with the 14 state-of-the-art algorithms in terms of S-measure, E-measure, weighted F-measure, and MAE. Our model is clearly better than the other baseline methods. Besides, we also show the FNR results of ours and the baseline methods in Fig. 7. As can be seen, our approach achieves the lowest FNR scores across all datasets. Visual comparison (see Fig. 6) also demonstrates the efficiency of our method for capturing integral objects. In fact, our ICON performs favorably against the existing methods across all datasets in terms of nearly all evaluation metrics. This demonstrates its strong capability in dealing with challenging inputs. In addition, we present the precision-recall and F-measure curves in Fig. 8. The solid red lines belonging to the proposed method are obviously higher than the other curves, which further demonstrates the effectiveness of the proposed method based on integrity learning.

4.2 Visual Comparison

Fig. 9 provides visual comparison between our approach and the baseline methods. As can be observed, our ICON generates more accurate saliency maps for various challenging cases, e.g., small objects (1st), large objects (2nd row), delicate structures (3rd row), low-contrast (4th row), and multiple objects (5th row). Besides, our framework can detect salient targets integrally and noiselessly. The above results demonstrate the accuracy and robustness of the proposed method.

4.3 Attribute-Based Analysis

In addition to the most frequently used saliency detection datasets, we also evaluate our method on another challenging SOC dataset . When compared with the previous six SOD datasets, this dataset contains many more complicated scenes. In addition, the SOC dataset categorizes images according to nine different attributes, including AC (appearance change), BO (big object), CL (clutter), HO (heterogeneous object), MB (motion blur), OC (occlusion), OV (out-of-view), SC (shape complexity), and SO (small object).

In Table II, we compare our ICON with 16 state-of-the-art methods, including Amulet , DSS , NLDF , C2SNet , SRM , R3Net , BMPM , DGRL , PiCANet-R (PiCA-R) , RANet , AFNet , CPD , PoolNet , EGNet , BANet and SCRN in terms of attribute-based performance. As seen, our ICON achieves clear performance improvement over the existing methods.

5 Failure Cases

Although the proposed ICON method outperforms other SOD algorithms and rarely generates completely incorrect prediction results, there are still some failure cases, as shown in Fig. 10. Specifically, in the first row, which shows a tidy room, our method is confused by whether the pillow or the bed and wall is the salient object. Meanwhile, in the second image, the three lamp lights are the salient regions, but our method cannot detect them. Similarly, other SOTA methods also fail for these samples. We believe there are several reasons for these failure cases: (1) strong color contrast influencing the model’s judgment (e.g., 1st row); (2) lack of sufficient training samples (see the 2nd row) and (3) controversial annotations (i.e., 1st row).

6 Ablation Study

To demonstrate the effectiveness of different components in our ICON, we report the quantitative results of several simplified versions of our method. We start from the encoder-decoder baseline (a UNet-like network with skip connections) and progressively extend it with different modules, including DFA, ICE, and PWV. As shown in Table III, we first add the DFA (i.e., ID: 2) component upon the Baseline (i.e., ID: 1), which demonstrates an obvious performance improvements. This is reasonable because DFA has the ability to search for objects with diverse cues. Then we add the ICE module (i.e., ID: 3), which again shows a substantial performance improvement. Finally, as anticipated, adding all components (ID: 4) to the proposed model achieves the best performance.

6.2 DFA vs. Other Feature Enhancement Methods

DFA, ASPP , Inception , and PSP are four feature enhancement methods (FEMs), which share some common ideas to stimulate representative feature learning. Differently, our DFA is designed to enhance feature sub-spaces without enlarging the receptive field, which yields more diverse representations. In Table IV, our DFA with fewer convolutional block clearly outperforms or at least is on par with other FEMs in terms of SmS_{m}, EξmE_{\xi}^{m}, and FβwF_{\beta}^{w}. However, DFA also brings some drawbacks. For instance, it generates higher MAE scores when compared with other FEMs. We argue that one possible illustration is that DFA not only brings feature diversity but also some noise. Besides, our experiments (i.e., ID: 2 vs. ID: 8∼\sim10) reveal that our method after combining three different types of convolution operations can achieve the best score. Meanwhile, our method using only 3xAsyConv generally yields better results than that only using 3xOriConv or 3xAtrConv.

7 ICE vs. Attention Methods

In Table V, we make an additional control group (i.e., ID: 3 vs. ID: 11∼\sim13) to verify the improvement brought by the ICE mechanism. Following the same setting (ID: 3), we conduct the experiments to compare ICE with SE , CBAM and GCT . We observe that the alternative method using CBAM achieves an acceptable performance and ranks second among these four methods. However, the other two methods using SE and GCT would lead to a noticeable performance drop. One possible explanation is that our ICE can strengthen the integrity of features and highlight potential salient candidates through our designed attention mechanisms.

To evaluate the performance of EM routing (ID: 4), we also conduct additional experiments (see Table VI) by replacing it with dynamic routing (DR) and self-routing (SR) . We observe that the first alternative method (ID: 14) also achieves reasonable performance, but the second alternative method (ID: 15) yields worse performance, when compared to our method using EM routing. One possible illustration is that SR does not have the routing-by-agreement mechanism, making it incompatible with our PWV scheme.

7.2 Evaluation of Loss Function

Conclusion

We present a novel Integrity Cognition Network, called ICON, to detect salient objects from given image scenes. It is based on the observation that mining integral features (at both the micro and macro level) can substantially benefit the salient object detection process. Specifically, in this work, three novel network modules are designed: the diverse feature aggregation module, the integrity channel enhancement module, and the part-whole verification module. By integrating these modules, our ICON is able to capture diverse features at each feature level and enhance feature channels that highlight the potential integral salient objects, as well as further verify the part-whole agreement between the mined salient object regions. Comprehensive experiments on seven benchmark datasets are conducted. The experimental results demonstrate the contribution of each newly proposed component, as well as the state-of-the-art performance of our ICON.

Acknowledgments

The authors would like to thank the anonymous reviewers and editor for their helpful comments on this manuscript. And we thank Jing Zhang for sharing codes of their work. This work is partially funded by the National Key R&D Program of China under Grant 2021B0101200001; the National Natural Science Foundation of China under Grant U21B2048 and 61929104, and Zhejiang Lab (No.2019KD0AD01/010).

References