CIR-Net: Cross-modality Interaction and Refinement for RGB-D Salient Object Detection

Runmin Cong, Qinwei Lin, Chen Zhang, Chongyi Li, Xiaochun Cao, Qingming Huang, Yao Zhao

I Introduction

WHEN viewing an image, humans are involuntarily attracted by some objects or regions in the image (e.g., the Smurfs in the second image of Fig. 1), which is mainly caused by the human visual attention mechanism, and these objects are called salient objects . Simulating this scheme, in the field of computer vision, salient object detection (SOD) is the task of automatically locating the most visually attractive objects or regions in a scene, which has been successfully applied to numerous tasks, such as segmentation , retrieval , enhancement , and quality assessment .

With the development of SOD task research, there are many subtasks, such as co-salient object detection (CoSOD) , remote sensing SOD , video SOD , light field SOD , have also been developed. In fact, the natural binocular structure of humans can also perceive the depth of field of the scene, and then generate stereo perception. Expressing this depth relationship in the form of an image is a depth/disparity map. In recent years, the development and popularization of depth sensors, especially the rise of affordable and portable consumer depth cameras, has further promoted the applications of RGB-D data, such as depth map super-resolution , depth estimation , superpixel segmentation , and saliency detection . For the RGB-D images, the RGB image contains abundant details and appearance information (e.g., color, texture, structure, etc) while the depth map provides some valuable supplementary information (e.g., shape, surface normals, internal consistency, etc). Recently, more and more studies focus on the introduction of depth cue for the SOD task to effectively suppress the background interference in complex scenes and further completely highlight foreground salient regions. For example, in Fig. 1, the first two images have complex and mussy backgrounds, and the color contrast between the salient object and the background in the fourth image is low. Thus, for the RGB SOD method (i.e., the GCPANet ) shown in the last row of Fig. 1, it is difficult to accurately locate the salient regions with a clean background and a complete structure. In comparison, the RGB-D SOD methods (e.g., the fourth and fifth rows of Fig. 1) can alleviate this problem with the introduction of depth information. Notably, our method has better object positioning ability, completeness preserving ability, and background suppression ability.

The effectiveness of depth map for SOD task has been validated in previous work ; however, how to effectively utilize and integrate the RGB information and depth cue is still an open issue. This is because RGB image and depth map belong to different modalities that have different attributes. To achieve this, we design the three-stream structure network to fully capture and utilize cross-modality information. Considering the strengths and complementarities of different modalities, through the three-stream structure with independent RGB and depth streams, we can sufficiently preserve the rich information and explore the complementary relations of different modalities, which is beneficial to jointly integrate cross-modality information in the encoder and decoder stage with a more comprehensive and in-depth manner than the two-stream structure. It is manifested in the following two aspects:

1) Cross-Modality Interaction. In terms of the cross-modality information, the primary problem we face is how to interact them. Specifically, the purpose is to learn the strengths and complementarities of different modalities, then obtain more comprehensive and discriminative feature representations. Different from the existing cross-modality interaction methods that operated only in the encoder stage or decoder stage , we dedicate to integrating cross-modality information into both encoder and decoder stages jointly in a more comprehensive and in-depth manner, which sufficiently explores the complementary relations of different modalities. Concretely, in the feature encoder stage, we design a progressive attention guided integration (PAI) unit to fuse cross-modality and cross-level features, thereby attaining the RGB-D encoder representations. In the feature decoder stage, we design an aggregation structure to allow RGB and depth decoder features to flow into the RGB-D mainstream branch and generate more comprehensive saliency-related features. In this structure, the decoder features of the previous layer, the RGB and depth decoder features of the corresponding layer are integrated into confluence decoder features through an important gate fusion (IGF) unit in a dynamic weighting manner. The gradually refined decoder features of the last layer are then used to predict the final saliency map.

2) Cross-Modality Refinement. In addition to cross-modality interaction, refining the most valuable information from different modalities is also crucial for RGB-D SOD task. To this end, we insert a refinement middleware between the encoder and the decoder, including the self-modality refinement and cross-modality refinement. For the self-modality refinement, in order to reduce the feature redundancy of the channel dimension and emphasize the important location of the spatial dimension, we propose a simple but effective self-modality attention refinement (smAR) unit, which replaces the commonly used progressive interaction or feature fusion method with our proposed channel-spatial attention generation. We directly integrate spatial attention and channel attention in the feature map space to generate a 3D attention tensor that is used to refine the single modality features, which not only reduces the computational cost, but also better highlights the important features. Further, we design a cross-modality weighting refinement (cmWR) unit to refine the multi-modality features by considering cross-modality complementary information and cross-modality global contextual dependencies. Inspired by the non-local model , the RGB features, depth features, and RGB-D features are integrated to capture the long-range dependencies among different modalities. Then, we use the integrated features to weight and refine different modality features, thereby obtaining the refined features embedded with cross-modality global context cue, which is important for the perception of global information.

In summary, our method is unique in that the cross-modality interaction and refinement are closely coupled in a comprehensive and in-depth manner. In terms of the cross-modality interaction, for learning the strengths and complementarities of different modalities, we propose the PAI unit in the encoder stage and the IGF unit in the decoder stage to jointly explore the complementary relations of different modalities. In terms of cross-modality refinement, considering the information redundancy of the encoder features and the significance of global context cues for the SOD, we design the pluggable refinement middleware structure to refine the encoder features from the self-modality and cross-modality perspectives. The main contributions are summarized as follows:

We propose an end-to-end cross-modality interaction and refinement network (CIR-Net) for RGB-D SOD by fully capturing and utilizing the cross-modality information in an interaction and refinement manner.

The progressive attention guided integration unit and the importance gated fusion unit are proposed to achieve comprehensive cross-modality interaction in the encoder and decoder stages respectively.

The refinement middleware structure including the self-modality attention refinement unit and cross-modality weighting refinement unit is designed to refine the multi-modality encoder features by encoding the self-modality 3D attention tensor and the cross-modality contextual dependencies.

Without any pre-processing (e.g., HHA ) or post-processing (e.g., CRF ) techniques, our network achieves competitive performance against the state-of-the-art methods on six RGB-D SOD datasets.

The rest of this paper is organized as follows. In Section II, we briefly review the related works of RGB-D SOD. In Section III, we introduce the technical details of the proposed CIR-Net. Then, the experiments including the comparisons with state-of-the-art methods and ablation studies are conducted in Section IV. Finally, the conclusion is drawn in Section V.

II Related Work

Different from RGB SOD models , depth modality together with RGB appearance are introduced into RGB-D SOD models. In the past ten years, a mass of methods have been proposed, which can be roughly divided into traditional methods and deep learning-based methods . Especially in recent years, the deep learning-based methods have achieved great breakthroughs in the performance of RGB-D SOD. For the RGB-D SOD task, how to make full use of the cross-modality information and generate more discriminate saliency-related representation is a challenging issue to be addressed . In terms of the model structure, the existing works can be roughly divided into single-stream, two-stream and three-stream structures, as shown in Fig. 2(a)-(c).

For the single-stream models , the early feature fusion strategy is commonly adopted, where RGB image and depth map are concatenated into four channels as the input of a network. For examples, Zhao et al. adopted a single-stream encoder to make full use of the representation ability of the pre-trained network, and proposed a real-time and robust salient detection model. Zhang et al. proposed the first uncertainty-inspired RGB-D SOD model based on conditional variational auto-encoder. Ji et al. proposed a novel collaborative learning framework that integrated the edge, depth, and saliency collaborators, which is a more lightweight and versatile network due to the free of depth inputs during testing. However, such models ignore the difference between RGB and depth modalities and lack the comprehensive cross-modality interaction.

The two-stream models are currently the most widely used structure in RGB-D SOD task, mainly including two independent branches to respectively process RGB and depth modality information and generate cross-modality features in the encoder or decoder stage. For example, Li et al. proposed an attention steered interweave fusion network, which progressively and interactively captures cross-modality complementarity via the interweave fusion and weighs the saliency regions by the steering of the deeply supervised attention mechanism. Li et al. adopted the late feature fusion strategy to generate cross-modality representation which combines high-level RGB and depth features of two independent branches in the decoder stage. Zhai et al. leveraged the multi-modal and multi-level features to devise a novel cascaded refinement network, and the RGB and depth modalities can be fused in a complementary way. Zhang et al. focused on the roles of RGB and depth modalities in the cross-modality interaction, and presented a discrepant interaction mode, i.e., the RGB modality and the depth modality guide each other interactively. Some studies are taking an interest in the negative impact of low-quality depth maps by controlling, updating, or abandoning the depth information in the two-stream structure . Chen et al. introduced depth quality perception to control the impact of low-quality depth maps while performing cross-modality interaction in the two-stream structure. Chen et al. estimated an additional high-quality depth map as a complement to the original depth map, and all these depth maps are fed into a selective fusion network to achieve RGB-D SOD. Chen et al. introduced a depth-quality-aware subnet into the two-stream RGB-D SOD structure to locate the most valuable depth regions.

In addition, some studies adopted the three-stream network structure for comprehensive cross-modality feature interaction, where RGB, depth, and RGB-D are embedded in three sub-networks for learning and interaction, respectively. For example, Fan et al. designed a gate mechanism to filter out the low-quality depth maps using the decoder results of RGB, depth, and RGB-D branches.

Compared with the existing works, our work differs conceptually from theirs in that: Our proposed network architecture (as shown in Fig. 2(d)) is a form between two-stream and three-stream networks, and the RGB-D stream is formed by interacting with the high-level features learned by the single-branch network. In this way, the parameters of the network can be reduced, and the RGB-D features can be better established by our designed PAI unit. On balance, we classify our network as a three-stream network architecture. This is also the first point that makes our network different from other networks. Second, in addition to the cross-modality feature integration through the PAI unit in the encoder stage, we also perform cross-modality information interaction in the decoder stage to obtain the discriminative saliency prediction features. Considering that the decoder features of RGB and depth streams can further provide effective guidance information (e.g., sharp edge, internal consistency) for RGB-D stream, we design a convergence aggregation structure in the entire decoder stage. In this way, we are dedicated to jointly integrating cross-modality information into the encoder and decoder stages in a more comprehensive manner. Third, to better establish the relationship between encoder features and decoder features, we introduce a refinement middleware structure to further highlight the effective information before decoding from the perspective of self-modality and cross-modality. It is worth mentioning that such a middleware structure is pluggable for three-stream networks.

III Proposed Method

Fig. 3 shows the overview of the proposed CIR-Net that is an encoder-decoder three-stream architecture equipped with a refinement middleware between the encoder and the decoder. In what follows, we detail the proposed method.

The feature encoder aims to learn the multi-level three-stream features, i.e., RGB, depth, and RGB-D encoder features. First, the backbone network (e.g., ResNet50) is used to extract the multi-level features from the input RGB image and depth map, denoted as frif_{r}^{i} and fdif_{d}^{i}, respectively, where i∈{1,2,3,4,5}i\in\{1,2,3,4,5\} indexes the feature level. Then, the RGB and depth features at high levels are fed into the proposed progressive attention-guided integration (PAI) unit to generate the cross-modality RGB-D encoder features frgbdi (i∈{3,4,5})f_{rgbd}^{i}~{}(i\in\{3,4,5\}). At this point, the three-stream encoder structure is formed, as shown on the left side of Fig. 3.

Considering the information redundancy in the self-modality and the content complementarity in the cross-modality, we introduce a refinement middleware structure to further highlight the effective information before decoding. Specifically, a two-stage refinement mechanism composing of a self-modality attention refinement (smAR) unit and a cross-modality weighing refinement (cmWR) unit is designed to progressively refine the multi-modality top-level encoder features in a self- and cross-modality manner.

In the decoder stage, we devise a novel convergence aggregation structure, in which the corresponding decoder features of the RGB and depth streams flow into the corresponding RGB-D stream to achieve cross-modality interaction. During aggregation, an importance gated fusion (IGF) unit is proposed to integrate the corresponding decoder features of RGB and depth streams and the previous IGF outputs in a dynamic weighting manner. Finally, the output features of the last IGF unit are used to infer the final saliency map.

III-B Progressive Attention Guided Integration Unit

Taking the complementarity and diversity of different modalities into account, effective cross-modality information interaction plays a critical role in the RGB-D SOD task. For an encoder-decoder network architecture, the existing interaction strategies are mainly designed separately in the encoder stage or decoder stage . In comparison, we design specialized modules in both encoder and decoder stages according to the different interaction purposes. To achieve that, two key issues need to be addressed: (1) how to effectively integrate and generate the RGB-D representations based on the multi-level RGB and depth features in the encoder stage, and (2) how the single-modality stream can better collaborate with the RGB-D stream to learn more discriminate saliency-related features and predict more accurate saliency map in the decoder stage. To this end, a PAI unit in the encoder stage and an IGF unit in the decoder stage are proposed in our method. The IGF unit will be introduced in Section III-D.

Specifically, to effectively integrate the RGB-D representations in the encoder stage, we consider two aspects when designing the PAI unit: (1) sufficient multi-level information fusion, and (2) effective feature selection and highlighting. For the former, in the encoder stage, considering the fact that the features of different levels contain different information with varying scales, receptive fields, and contents. Thus, the progressive cross-level fusion strategy is designed to obtain more comprehensive RGB-D representations in a coarse-to-fine manner. For the latter, although the encoder features contain rich multi-level information, the commonly used fusion strategy (e.g., concat-conv) may introduce information redundancy and easily confuse the feature representations. Therefore, for feature selection and enhancement, we introduce the spatial attention scheme to guide the cross-level and cross-modality feature fusion by highlighting the complementary information and suppressing irrelevant redundancy.

First, motivated by the fact that the shallower depth features usually contain too much background noise and the high-level features contain clear information of salient objects but lack details, we choose to generate the initial cross-modality features by combining high-level RGB and depth features and start the combination of features and forward propagation from the third layer, which can be described as:

where frif_{r}^{i} and fdif_{d}^{i} respectively denote the RGB and depth features at ithi^{th} encoder level, [⋅,⋅][\cdot,\cdot] is the channel-wise concatenation operation, and convconv represents a convolutional layer followed by a batch normalization (BN) layer and a ReLU activation function.

Then, in order to highlight the complementary information and suppress the irrelevant redundancy in the cross-level and cross-modality fusion, we employ the spatial attention map generated by the previous RGB-D level to guide the current-level feature integration in a progressive manner. Thus, the final RGB-D features at 4th4^{th} and 5th5^{th} levels are updated as:

where ⊙\odot is the element-wise multiplication, Ai−1A^{i-1} denotes the attention map of (i−1)th(i-1)^{th} level, SASA is the spatial attention operation , and ↓\downarrow denotes the down-sampling operation. Note that, considering the inaccurate attention that may happen in some challenging cases, we adopt the residual connection in Eq. (2) to learn the optimal relationship between the learned features and the original features for effective feature learning. In Section IV-C, we provide the ablation study to demonstrate the effectiveness of this operation. Our PAI unit can not only integrate the different modality information, but also encode the different levels’ features in a progressive attention weighting manner, thereby generating the RGB-D encoder features.

III-C Refinement Middleware

To transfer more effective encoder features into the decoder stage, we insert a refinement middleware structure as a connecting link between the encoder and decoder to refine the encoder features from the perspectives of self-modality and cross-modality. For the design of refinement middleware, we consider two aspects: 1) the encoder features of each modality contain abundant spatial and channel information while indiscriminate information transmission may increase the difficulty of learning effective feature representations. Therefore, we design a smAR unit to suppress the background noises and highlight the important cues from a single modality perspective; and 2) considering the strong correlation and complementarity between different modalities where the RGB modality contains figure-ground color contrast and the depth modality contains internal consistency, we design a cmWR unit to capture the long-range dependencies of multiple modalities and refine the modality features from a global perspective.

After the feature encoding, the obtained RGB, depth, and RGB-D encoder features contain abundant spatial and channel information representing the salient objects. However, there will be redundancy in the single-modality information. Moreover, indiscriminate information transmission may increase the difficulty of feature learning, and even contaminate the inference of subsequent decoding process. Therefore, we design a smAR unit in the refinement middleware to suppress the background noises and highlight the important cues from the perspective of single modality in a new spatial-channel 3D attention manner.

The spatial attention (SA) and channel attention (CA) have been widely used in the existing RGB-D SOD tasks , which can be summarized into three forms: (a) Separate utilization. In , SA and CA are applied to low-level features and high-level features, respectively. (b) Serial utilization. In , the CA is first used to generate the CA-enhanced features, and then the SA is subsequently applied to obtain the final enhanced features. (c) Parallel utilization via feature fusion. In , the CA and SA are respectively used to enhance the same input features, and then the obtained enhanced features are fused to generate the final features. The CA and SA in separate utilization are used for different level features, which are not necessarily suitable for all vision tasks. However, the serial utilization is sensitive to the order of SA-CA combination, while the way of feature fusion in parallel utilization has some information redundancy in the structural design, and can only enhance the features in one dimension (i.e., spatial or channel) at a time, which increases computational complexity. To address this issue, we integrate SA and CA into a spatial-channel 3D attention tensor for: 1) enhancing the robustness via parallel utilization and reducing the computational complexity in a 3D attention manner; and 2) refining the single modality features in both spatial and channel dimension simultaneously.

As shown in the left side of Fig. 4, the output features of the three encoder branches (i.e., fr5f^{5}_{r}, fd5f^{5}_{d}, and frgbd5f^{5}_{rgbd}) are embedded into the smAR unit. We first calculate the CA and SA of the input features in a parallel structure respectively, thereby obtaining the corresponding spatial attention map and channel attention map. Then, we directly fuse them on the attention map space via matrix multiplication to generate the 3D attention tensor. This process can be described as:

where fmod5f_{mod}^{5} denotes the each modality features of the top encoder layer, mod∈{r,d,rgbd}mod\in\{r,d,rgbd\}, SASA and CACA represent the spatial attention and channel attention operations, respectively, and ⊗\otimes denotes the matrix multiplication. With the 3D attention tensor, we refine each modality features through a residual connection:

where ⊙\odot is the element-wise multiplication. In Section IV-C, we provide ablation studies with different attention combinations to demonstrate the effectiveness of our design.

III-C2 Cross-modality Weighting Refinement Unit

The smAR unit refines the encoder features in each modality, but does not make full use of the strong correlation and complementarity between different modalities. For example, the RGB modality contains figure-ground color contrast and object texture, and the depth modality provides the internal consistency and spatial relations of the salient objects. Therefore, inspired by the non-local model , we design a novel cmWR unit in the second stage of the refinement middleware to further capture long-range dependencies of multiple modalities.

where WθW_{\theta}, WξW_{\xi}, WφW_{\varphi}, and WψW_{\psi} denote the learnable embedding weights through the bottleneck convolutional layers.

Then, similar to the scaled dot-product attention, the correlation between the RGB features and depth features, and the self-correlation of the RGB-D features are calculated in a pixel-wise manner:

Finally, these two correlation information mapped to the RGB-D modality jointly generate cross-modality global dependency weights to refine the original input features:

III-D Importance Gated Fusion Unit

As we emphasized before, the cross-modality information interaction is essential for RGB-D SOD task. Existing methods usually only interact in a separate encoder or decoder stage, but this is insufficient. In fact, the encoder and decoder play different roles in feature learning, where the encoder focuses more on general feature extraction while the decoder places extra emphasis on the learning of saliency-related features. Thus, in addition to the cross-modality feature integration through the PAI unit in the encoder stage, we also perform cross-modality information interaction in the decoder stage to obtain the discriminative saliency prediction features. Considering that the decoder features of RGB and depth streams can further provide effective guidance information (e.g., sharp edge, internal consistency) for RGB-D stream, we design a convergence aggregation structure in the entire decoder stage. In detail, the single modality features (i.e., RGB and depth decoder features) at the same level will flow to the corresponding RGB-D stream to learn more comprehensive cross-modality decoder features. For the convergence aggregation structure, we face a challenging problem, i.e., how to effectively select the most valuable information from the afflux streams, because the direct and equal combination of different modality information may be uncontrollable and miscellaneous. To solve this issue, we design an IGF unit to learn an importance map PiP^{i}, which is used to selectively control the influence of different modalities in a dynamic weighting manner, as shown in Fig. 5. In this way, the IGF unit can determine the contribution of supplementary information of different modalities during cross-modality information interaction. Furthermore, with such learnable important weights, our network is somewhat resistant to situations where certain modal features are invalid, such as low-quality depth maps.

First, the RGB decoder features and the depth decoder features are fused with the corresponding skip-connection encoder features via two convolutional layers, thus attaining the fused decoder features. Then, the fused decoder features of the RGB and depth streams are concatenated to obtain the RGB-D decoder features HiH^{i}. Finally, the previous IGF features fIGFi+1f_{IGF}^{i+1} and the RGB-D decoder features HiH^{i} are combined into the current IGF outputs through the learnable importance weight:

where CACA denotes the channel-wise attention , and σ\sigma is the sigmoid activation function. The importance map determines the contribution of supplementary information of different modalities at the ithi^{th} decoder level.

III-E Loss Function

In our CIR-Net, the last layer of decoder features of the three streams are used to separately predict the corresponding saliency maps, which are denoted as SrS^{r}, SdS^{d}, and SrgbdS^{rgbd}. For network training, we employ the binary cross-entropy (BCE) loss function to optimize the RGB, depth, and RGB-D streams simultaneously. The final loss function is defined as:

IV Experiments

We first describe the six RGB-D SOD benchmark datasets and three commonly used evaluation metrics, then introduce the implementation details of the proposed model. After that, the comparisons with 15 state-of-the-art CNN-based methods are conducted. Finally, we conduct a series of ablation studies to validate the effectiveness of our proposed modules.

We conduct experiments on six popular RGB-D SOD benchmark datasets, including STEREO797 , NLPR , NJUD , DUT , LFSD and SIP . NJUD contains 1985 RGB-D images and corresponding manually labeled ground truth. The images are collected from the Internet and stereo movies with diverse objects and complex scenarios, and the depth maps are estimated from the stereo images. NLPR consists of 1000 multiple salient objects RGB-D images, where the depth maps are captured by the Kinect with a resolution of 640×480640\times 480. STEREO797 includes 797 stereoscopic images collected from the Internet, and the depth maps are estimated from the stereo images. DUT contains 1200 paired RGB-D images captured by a Lytro camera with a resolution of 600×400600\times 400. LFSD is a small-scale dataset including 100 small-resolution RGB-D images, where the depth maps are captured via a Lytro light field camera. SIP includes 929 RGB-D images with a high-resolution of 744×992744\times 992.

IV-A2 Evaluation Metrics

To quantitatively evaluate the performance of the proposed method, precision-recall (P-R) curves, F-measure (FβF_{\beta}) , Mean Absolute Error (MAE) score , and S-measure (SmS_{m}) are employed. By thresholding the saliency map from 0 to 255, the precision and recall scores can be calculated by comparing the binary mask with the corresponding ground truth, and the variation tendency of different precision and recall scores can be drawn in a precision-recall curve.

F-measure is a widely used comprehensive evaluation metrics by considering both precision and recall scores, which is defined as:

where PrecisionPrecision and RecallRecall respectively represent the precision score and recall score, and β2\beta^{2} is set to 0.3 for emphasizing the precision as suggested in .

The MAE score calculates the average pixel-wise absolute difference between the predicted saliency map SS and the corresponding ground truth GG, which is denoted as:

where HH and WW represent the height and width of the image, respectively.

S-measure denotes the structural similarity between the predicted saliency map and the corresponding ground truth:

where α\alpha{} is set to 0.5 to balance the region similarity SrS_{r} and object similarity SoS_{o} as suggested in .

IV-A3 Implementation Details

Following , we adopt 14851485 samples from NJUD dataset, 700700 samples from NLPR dataset, and 800800 samples from DUT dataset as the training data. The remaining samples in these three datasets and the rest three datasets are used as testing datasets. During training, the random flipping, rotating and multi-scale input are adopted for data augmentation. During the training phase, the training samples are randomly resized to 128×128128\times 128, 256×256256\times 256, and 352×352352\times 352. In the interference stage, the images are resized to 352×352352\times 352 and then fed into the network to obtain saliency prediction without any other post-processing or pre-processing techniques. We report the experimental results using ResNet50 and VGG16 as backbone networks, initialized by the pre-trained parameters on ImageNet .Unless otherwise stated, the results in this paper are obtained with ResNet as the backbone network. The Adam algorithm is used to optimize our network with a batch size of 1616, and the initial learning rate 1e1e-44 is divided by 55 every 4040 epochs. Our network is implemented in PyTorch and accelerated by two NVIDIA 20802080Ti GPUs. We also implement our network by using the MindSpore Lite toolhttps://www.mindspore.cn/. In order to show the training process of our model more clearly, we report the learning curve of our network in Fig. 6. It takes around 4 hours to optimize our network. The inference time of our method is 0.070.07 second for an image with the size of 352×352352\times 352.

IV-B Comparison with the State-of-the-art Methods

We compared the proposed model with 1515 state-of-the-art CNN-based RGB-D SOD methods, including DMRA , FRDT , SSF , S2MA , A2dele , JL-DCF , PGAR , DANet , cmMS , BiANet , D3Net , UCNet , ASIF-Net , BBSNet , and UCNet* (the extension version of UCNet). For fair comparisons, all the saliency maps are generated by the released code under the default settings or are provided by the authors directly.

To further illustrate the outperformance of our proposed method, we provide some visualization comparison results of different methods in Fig. 7. From it, we can clearly see that our proposed model achieves superior performance, which achieves accurate location and complete structure of the salient objects. For quantitative evaluations, we report the P-R curves of different methods on six benchmark datasets, which is shown in Fig. 8. The closer the P-R curve is to (1,1)(1,1), the better the algorithm performance. As visible, our model (i.e., the red solid line) achieves both higher precision and recall scores against other compared methods over all six benchmark datasets. Moreover, as shown in Table I, our method achieves the best performance except for the MAE metric on the SIP dataset, which also demonstrates the effectiveness and superiority of the proposed method. For example, compared with the second best method on the large-scale popular NLPR-test, DUT-test, STRERO797 datasets, the minimum percentage gain reaches 3.0%3.0\%, 15.3%15.3\%, 10.7%10.7\% for MAE scores, respectively. On the small-scale LFSD dataset, compared with the second best model, the percentage gain reaches 1.6%1.6\% in terms of F-measure and 5.0%5.0\% in terms of MAE score.

In order to better illustrate the advantages of our method, we analyze and summarize the qualitative and quantitative results from the following aspects:

For some common scenes, such as scenes with obvious foreground-background color contrast, large-size object, single object, simple structure, etc., although most of the existing methods can also achieve good results, our method is more stable and robust. As shown in the first two images of Fig. 7, the salient object is simple in structure and its color contrasts sharply with the background. In this case, while most of the works can effectively locate the salient object, our work is able to obtain more accurate results, such as sharp object boundaries (e.g., the pointy tip of the leaf in the first image), clean background suppression (e.g., the leaves in the second image).

In addition, to verify the robustness and performance of our method on the challenging scenes, we conduct several sensitive studies on the testing subsets. The quantitative comparison results are shown in Table II.

(1) Our method has certain advantages when dealing with unreliable depth maps. As shown in Table II (No.1), we conduct a sensitive experiment to evaluate the performance of our method on unreliable depth map samples. Specifically, we select depth maps with the depth confidence λd\lambda_{d} score less than 0.1 from the six testing datasets as unreliable depth maps, denoted as unreliable-depth subset. As reported in Table II (No.1), compared with the second best method (i.e., DANet), the percentage gain reaches 20.0%, 4.7%, and 3.8% for MAE score, F-measure, and S-measure, respectively. Moreover, as shown in the third and fourth images of Fig. 7, the depth values of the salient objects are similar to the background, which greatly interferes with the detection of salient objects. Due to the interference of unreliable depth information, most works (e.g., S2MA, A2dele, D3Net) fail to suppress the background noise, leading to the inaccurate results. Benefiting for the overall network architecture and effective cross-modality interactions, our model can obtain robust results in the face of these unreliable factors.

(2) Our method has certain advantages when dealing with multi-object scenes. To be specific, we collect all samples with multiple salient objects from the six testing datasets based on the ground truth, denoted as multi-object subset. As shown in Table II (No.2), the percentage gain in both F-measure and S-measure reaches 1.6% compared with the second best method. In addition, as can be seen from the fifth and sixth images of Fig. 7, benefiting from the cross-modality feature refinement in a global perspective, our method can not only correctly locate all salient objects, but also obtain a complete and consistent structure, such as the inner area of the person on the right in the sixth image.

(3) Our method has certain advantages when dealing with low-contrast scenes. Similarly, we select all low-contrast samples with the average color similarity between the salient objects and backgrounds exceeding 80% from the six testing datasets (denoted as low-contrast subset) to verify the superiority of our method in this case. As shown in Table II (No.3), compared with the second best method (i.e., SSF), the percentage gain reaches 25.9%, 3.0%, and 3.7% for MAE score, F-measure, and S-measure, respectively. As shown in the seventh and eighth images of Fig. 7, most of the existing works disturbed by the low color contrast interference, failing to obtain a complete result. In contrast, our method handles such a challenging scene by better exploiting complementary information across modalities, resulting in more complete and accurate results, such as hand regions of the person.

(4) Our method has certain advantages when dealing with small-object scenes. Experimentally, we select all samples with the salient object occupying less than 10% of the image from the six testing datasets (denoted as small-object subset) to measure the performance of the proposed model in small object scenes. In Table II (No.4), compared with the second best method (i.e., SSF),the percentage gain reaches 18.7%, 5.5%, and 3.3% for MAE score, F-measure, and S-measure, respectively. As can be seen from the ninth and tenth images in Fig 7, our method can effectively locate the small salient object, obtaining results with accurate locations, clean backgrounds, and sharp boundaries.

IV-C Ablation Study

To evaluate the effectiveness of each module in the proposed model, we conduct the ablation studies on the NJUD-test, STEREO797 and LFSD datasets. The quantitative evaluations and visual examples are shown in Table III and Fig. 9, respectively. We construct our baseline model by simplifying our full model as follows:

replacing the PAI unit with the feature concatenation of the fifth layer in the RGB and depth streams;

removing the refinement middleware structure including its smAR unit and cmWR unit;

replacing the IGF unit with the simple deconvolutional layers.

We use the method of progressively adding designed modules for ablation experiments. We first introduce the PAI unit into the baseline model (denoted as ‘+PAI’), then we progressively add the IGF unit, cmWR unit, and smAR unit into the model. In other words, ‘+IGF’ denotes the ‘baseline+PAI+IGF’, and the like. Moreover, all the ablation models are trained by using the same training configurations as our CIR-Net.

In Fig. 9, it shows that the baseline model roughly locates the salient objects but lacks complete structure and sharp boundary, and many background regions are not effectively suppressed. Compared with the baseline model, the introduction of the PAI module obtains more complete and consistent structural information (e.g., the flower in the first image), but still include many wrongly detected background regions. From the quantitative result, the F-measure is improved from 0.88800.8880 to 0.89520.8952 on the NJUD-test dataset, and the F-measure is increased from 0.87690.8769 to 0.88530.8853 on the STEREO797 dataset. Then, after adding the IGF unit for cross-modality feature integration in the decoder stage, the clearer boundaries of the salient objects (e.g., the flower in the first image) can be obtained and the quantitative performance is obviously improved. Specifically, the F-measure score is increased to 0.91350.9135 on the NJUD-test dataset, and the percentage gain of the F-measure score reaches 2.0%2.0\% compared with the ‘+PAI’ model. Furthermore, by introducing the cmWR unit to refine the different modalities from a global perspective, it is observed that the background suppression and object structure are improved to a certain extent. Finally, after adding the smAR unit to highlight the important cues from the single modality perspective, the full model (i.e., ‘+smAR’ in Fig. 9 and Table III) yields the best performance with the percentage gain of 4.5%4.5\% and 4.2%4.2\% in terms of F-measure on the NJUD-test dataset and STEREO797 dataset compared with the baseline model. In summary, the ablation studies further demonstrate the effectiveness of the proposed modules.

IV-C2 Analysis of the converged three-stream architecture

To demonstrate the effectiveness of the converged three-stream architecture, we conduct several experiments in Table IV and Fig. 10.

First, we remove the RGB-D branch in the decoder from the full model and fuse the output features of RGB and depth branches via concatenation to obtain the final saliency map, thereby constructing the two-stream architecture network (denoted as ‘Two-stream’). From the quantitative results, we can see that, with the help of comprehensive feature interaction in the three-stream structure, the CIR-Net is more effective than the two-stream architecture. For example, on the LFSD dataset, the F-measure of the three-stream network is 0.0209 higher than that of the two-stream network, and the S-measure is 0.0185 higher. Similarly, from the visualization results shown in Fig. 10, we can see the advantages of the three-stream structure in detection accuracy and completeness. Of course, the performance gain comes at a price. Compared with the two-stream structure, the three-stream design needs more computational resources and parameters due to the use of more branches. To be specific, due to the additional parameters, the inference speed for an image of the three-stream architecture is 14 fps, while that of the two-stream architecture is 18 fps.

In addition, we also quantify the saliency performance of the three branches separately. As can be seen from Fig. 10, the RGB branch and the Depth branch have their own advantages and disadvantages in different regions, but our final RGB-D branch can concentrate on the advantages of both and suppress the disadvantages, so as to achieve better results with sharper edges and complete structure. For the quantitative comparison, it can be found that, with the help of effective cross-modality feature interaction, compared with the RGB branch performance, the final RGB-D saliency performance is significantly improved. For example, compared with the RGB branch on the LFSD dataset, the F-measure is improved from 0.8324 to 0.8828 with a percentage gain of 6.0%, and the S-measure is improved from 0.8339 to 0.8753 with a percentage gain of 5.0%. These experiments demonstrate the robustness and effectiveness of the proposed model architecture.

IV-C3 Analysis of Refinement Middleware

We conduct various ablation experiments in Table V to validate the effectiveness of the refinement middleware.

In terms of the proposed smAR unit, we replace the 3D attention tensor with a single channel attention weight (denoted as ‘w/ CA, w/o SA’), a single spatial attention weight (denoted as ‘w/o CA, w/ SA’), and a serial utilization of the SA-CA combination (denoted as ‘SA-CA’). From Table V, we can see that the proposed smAR unit is more effective than other commonly used attention variants. For example, compared with the serial utilization of SA-CA combination module on the STEREO797 dataset, the F-measure of our full model with smAR unit reaches 0.9139 with a percentage gain of 1.6%, and the percentage gain of S-measure is 1.5%.

In terms of the proposed cmWR unit, we replace the final weight map (i.e., M1×M2M_{1}\times M_{2}) with the M1M_{1}-only (denoted as ‘w/ M1, w/o M2’) and M2M_{2}-only (denoted as ‘w/o M1, w/ M2’) cases to demonstrate the advantages of the cmWR unit. From the quantitative comparison reported in Table V, we can see that the way of M1×M2M_{1}\times M_{2} is more effective than the single weight map M1M_{1} or M2M_{2}. For example, on the LFSD dataset, compared with the case of only M2M_{2}, the F-measure score is improved from 0.8582 to 0.8828 with a percentage gain of 2.9% and the S-measure score is improved from 0.8462 to 0.8753 with a percentage gain of 3.4%.

IV-C4 Analysis of different feature interaction strategy in PAI and IGF units

To verify the effectiveness of our design of PAI and IGF units, we conduct various experiments, as shown in Table VI.

In terms of the PAI unit, we add two ablation experiments. One is to validate the combination and propagation of different layers, and the other is to replace the spatial attention maps with the fused cross-modality RGB-D features in the corresponding layer (denoted as ‘Trans Fusion’). From Table VI, we can see that starting combination from the third layer (the Full model) achieves the best performance, and the PAI unit is more effective than the commonly used feature fusion strategy. For example, on the LFSD dataset, the F-measure score is improved from 0.8555 to 0.8828 with a percentage gain of 3.2% compared with the forward propagation from the first layer, and the S-measure score is improved from 0.8480 to 0.8753 with a percentage gain of 3.2% compared with the feature fusion strategy.

In terms of IGF unit, we replace the dynamic fusion strategy with the addition (denoted as ‘w/ add’) or concatenation (denoted as ‘w/ cat’) to suggest the effectiveness of the IGF unit. From Table VI, we can see that with help of the proposed IGF unit, the performance is improved when compared with the commonly used fusions strategy (add or concatenation). For example, on the LFSD dataset, compared with the concatenation operation (i.e., w/ cat), the F-measure score is improved from 0.8605 to 0.8828 with a percentage gain of 2.6% and the S-measure score is improved from 0.8518 to 0.8753 with a percentage gain of 2.8%.

IV-C5 Analysis of residual connection

To demonstrate the effectiveness of residual features, we conduct the ablation studies that remove the addition operation in Eqs. (2, 5, 8). The quantitative results are shown in Table VII. Compared with the only direct multiplication operation, our method with residual connection achieves better quantitative performance. For example, in Table VII, compared with the result of removing the addition in Eq. (2), on the LFSD dataset, the F-measure is improved from 0.8500 to 0.8828 with a percentage gain of 3.9% and the percentage gain of S-measure reaches 3.4%. Similarly, removing the addition operations in Eq. (5, 8) also degrades the performance.

IV-C6 Analysis of effectiveness to different scenes

When the depth map is unreliable, through the cross-modality interaction of PAI unit in the encoder stage, the features of RGB-D branch can exploit the correlation between RGB and depth modalities to highlight the salient regions. Moreover, in the decoder stage, the IGF unit is able to selectively determine the contribution of depth modality, thus suppressing the interference of unreliable information in depth modality. To validate the effectiveness of the PAI and IGF units on unreliable depth maps, we add two ablation studies on the unreliable-depth subset that replace the PAI unit with the feature concatenation of the fifth layer in the RGB and depth branches (denoted as ‘w/o PAI’), and replace the IGF unit with the direct feature concatenation (denoted as ‘w/o IGF’), respectively. As shown in the left side of Table VIII, it can be found that after removing the PAI unit and the IGF unit, the detection effect of the model on unreliable depth maps decreases. For example, with the PAI unit, the F-measure is improved from 0.8886 to 0.9022 with a percentage gain of 1.5%. Similarly, with the IGF unit, the F-measure is improved from 0.8978 to 0.9022 with a percentage gain of 0.5%.

In addition, concerning the challenging scene containing multiple salient objects, the cmWR unit can extract the cross-modality global context information by calculating the long-range dependency, thus refining features from a global perspective and improving the completeness of saliency results. Similarly, we add an ablation experiment to demonstrate the effectiveness of the cmWR unit in the multi-object scene. As can be reported in the right side of Table VIII, compared with the model without cmWR unit (denoted as ‘w/o cmWR’), the F-measure is improved from 0.8626 to 0.8715 with a percentage gain of 1.0%.

IV-D Discussion

Several representative failure cases are shown in Fig. 11. We can see that it is difficult to perfectly locate salient objects in the following aspects: 1) Multiple and small salient objects. In the first scene, although the multiple salient objects contain the same characteristics in the scene, the salient objects far from the lens are too small, so that the corresponding depth map fails to provide effective depth information of these objects. Hence, it is difficult to completely detect all salient objects in such a scene. 2) High contrast but not salient objects. In the second scene, it is obvious that the bike seat contrasts sharply with the background in the depth map. However, the red logo, the real salient object, is also in sharp contrast to the bike seat in the RGB image. Therefore, the ambiguity introduced by this conflict prevents our model from accurately detecting the red logo as the salient object. 3) Complex background noise. In the third scene, due to the small contrast between the salient object and the background of the RGB image and the misleading depth information in the depth map, our algorithm fails to suppress the background effectively. It is worthy to note that, for the above challenging scenes, the recent state-of-the-art methods (S2MA and DANet ) also fail to detect the salient objects correctly.

IV-D2 Future Work

In the future, work in three areas can be further studied. First, our paper mainly focuses on how to achieve cross-modality interaction more fully and effectively and does not specifically consider the solution when the quality of the depth map is unreliable, but only uses some control mechanisms (e.g., cmWR and IGF modules) to reduce the negative impact of low-quality depth maps. Under the existing depth imaging equipment, how to stably and explicitly achieve salient object detection in the case of poor depth map quality is a problem worthy of study. Second, as we all know, deep learning-based methods are data-driven. Thus, more training data would improve the generalization capability of deep models in most cases. As presented in Table IX, when discarding the training data of the NLPR dataset (i.e., ‘w/o NLPR’), the final performance is all degraded, but to varying degrees. For example, on the STEREO797 dataset, the F-measure drops by only 0.5%, but on the LFSD dataset it drops by 4.4%. Put like that, constructing larger datasets or reducing the dependence on data volume under the premise of ensuring performance can be worked as future research directions. Furthermore, weakly supervised RGB SOD task has received a lot of attention, but very little in RGB-D SOD. Exploring RGB-D SOD models with less supervisory information can reduce the dependence on data annotation and is a very valuable and promising research direction. Last but not least, two- and three-stream based RGB-D SOD models have achieved satisfactory performance, but as a fundamental pre-processing task, how to pursue real-time efficiency while maintaining performance is also a valuable research point.

V Conclusion

In this work, we proposed an end-to-end network, named CIR-Net, for the task of RGB-D SOD. The strength of our algorithm comes from the synergy of the model architecture and technical modules.

From the perspective of model architecture, we design a new three-stream-like model architecture to more comprehensively realize cross-modality information interaction. As we all know, the two-stream models are currently the most widely used structure in RGB-D SOD task, mainly including an RGB branch and a depth branch, which can achieve the cross-modality interaction in the feature encoder or decoder stage. However, the two-stream models can only complete the interaction of RGB and depth modalities, while ignoring the role of RGB-D modality. In contrast, the three-stream structure has the opportunity to model the correlation and interaction among the RGB, depth, and RGB-D modalities. Moreover, our proposed model architecture is also different from the existing three-stream structures. On the one hand, the generation of our RGB-D stream is not learned from scratch, but obtained through the fusion of the high-level features from the RGB branch and depth branch through the PAI module, which can make the learned RGB-D features more discriminative and reduce the amount of calculation. On the other hand, we adopt a clear convergence structure at the decoder stage to realize the information interaction centered on the RGB-D modality, which can further capture the complementarity of the three modalities (i.e., RGB, depth, and RGB-D), thereby obtaining more discriminative and saliency-related features.

From the technical design level, as our title says, we do two things in this paper: cross-modality interaction and cross-modality refinement. For the cross-modality interaction, different from the existing cross-modality interaction methods that operated only in the encoder or decoder stage, we dedicate to integrating cross-modality information into both encoder and decoder stages jointly in a more comprehensive and in-depth manner. Concretely, in the feature encoder stage, a PAI unit is designed to fuse the cross-modality and cross-level features, thereby attaining the RGB-D encoder representations. In the feature decoder stage, we design a convergence structure equipped with the IGF unit to make the RGB and depth decoder features flow into the RGB-D mainstream branch, and effectively select the most valuable supplementary information from RGB and depth modalities to obtain more discriminative cross-modality saliency prediction features. For the cross-modality refinement, we insert a refinement middleware between the encoder and decoder to further highlight the effective information before decoding from the perspective of self-modality and cross-modality. Specifically, we propose a simple but effective smAR unit in a 3D-tensor manner to reduce the feature redundancy of the channel dimension and emphasize the important location of the spatial dimension, as well as propose a cmWR unit to refine the multi-modality features by considering cross-modality complementary information and cross-modality global contextual dependencies. It is worth mentioning that such a middleware structure is pluggable for three-stream networks.

The mutual cooperation and facilitation of model structure and technical modules enable our method to achieve competitive performance on six datasets both qualitatively and quantitatively.

References