RGB-D Salient Object Detection: A Survey
Tao Zhou, Deng-Ping Fan, Ming-Ming Cheng, Jianbing Shen, Ling Shao
I Introduction
Salient object detection (SOD) aims to locate the most visually prominent object(s) in a given scene . SOD plays a key role in a range of real-world applications, such as stereo matching , image understanding , co-saliency detection , action recognition , video detection and segmentation , semantic segmentation , medical image segmentation , object tracking , person re-identification , camouflaged object detection , image retrieval , etc. Although significant progress has been made in the SOD field over the past several years , there is still room for improvement when faced with challenging factors, such as complicated background or different lighting conditions in the scenes. One way to overcome these challenges is to employ depth maps, which provide complementary spatial information for RGB images and have become easier to capture due to the large availability of depth sensors (e.g., Microsoft Kinect).
Recently, RGB-D based SOD has gained increasing attention and various methods have been developed . Early RGB-D based SOD models tended to extract handcrafted features and then fuse RGB image and depth maps. For example, Lang et al. , the first work on RGB-D based SOD, utilized Gaussian mixture models to model the distribution of depth-induced saliency. Ciptadi et al. extracted 3D layout and shape features from depth measurements. Besides, several methods measure depth contrast using the depth difference between different regions. In , a multi-contextual contrast model including local, global, and background contrast was developed to detect salient objects using depth maps. More importantly, however, this work also provided the first large-scale RGB-D dataset for SOD. Despite the effectiveness achieved by traditional methods using handcrafted features, they tend to suffer from a limited generalization ability for low-level features and lack the high-level reasoning required for complex scenes. To address these limitations, several deep learning-based RGB-D SOD methods have been developed, showing improved performance. DF was the first model to introduce deep learning technology into the RGB-D based SOD task. More recently, various deep learning-based models have focused on exploiting effective multi-modal correlations and multi-scale/level information to boost SOD performance. To more clearly describe the progress in the RGB-D based SOD field, we provide a brief chronology in Fig. 2.
In this paper, we provide a comprehensive survey on RGB-D based SOD, aiming to thoroughly cover various aspects of the models for this task and provide insightful discussions on the challenges and open directions for future work. We also review another related topic, i.e., light field SOD, in which the light field can provide more information (including focal stack, all-focus images, and depth maps) to boost the performance of salient object detection. Further, we provide a comprehensive comparison to evaluate existing RGB-D based SOD models and discuss their main advantages.
There are several surveys that are closely related to salient object detection. For example, Borji et al. provided a quantitative evaluation of 35 state-of-the-art non-deep saliency detection methods. Cong et al. reviewed several different saliency detection models, including RGB-D based SOD, co-saliency detection, and video SOD. Zhang et al. provided an overview of co-saliency detection and reviewed its history, and summarized several benchmark algorithms in this field. Han et al. reviewed the recent progress in SOD, including models, benchmark datasets, and evaluation metrics, as well as discussed the underlying connection among general object detection, SOD, and category-specific object detection. Nguyen et al. reviewed various works related to saliency applications and provided insightful discussions on the role of saliency in each. Borji et al. provided a comprehensive review of recent progress in SOD and discussed some related works, including generic scene segmentation, saliency for fixation prediction, and object proposal generation. Fan et al. provided a comprehensive evaluation of several state-of-the-art CNNs-based SOD models, and proposed a high quality SOD dataset, termed SOC (details can be found at: http://dpfan.net/socbenchmark/). Zhao et al. reviewed various deep learning-based object detection models and algorithms in detail, as well as various specific tasks, including SOD works. Wang et al. focused on reviewing deep learning-based SOD models. Different from previous SOD surveys, in this paper, we focus on reviewing the existing RGB-D based SOD models and benchmark datasets.
I-B Contributions
Our main contributions are summarized as follows:
We provide the first systematic review of RGB-D based SOD models from different perspectives. We summarize existing RGB-D SOD models into traditional or deep methods, fusion-wise methods, single-stream/multi-stream methods, and attention-aware methods.
We review nine RGB-D datasets that are commonly used in this field, and provide details for each dataset. Moreover, we provide a comprehensive as well as an attribute-based evaluation of several representative RGB-D based SOD models.
We supply the first collection and review of the related light field SOD models and benchmark datasets.
We thoroughly investigate several challenges for RGB-D based SOD, and the relation between SOD and other topics, shedding light on potential directions for future research.
I-C Organization
In II, we review existing RGB-D based models in terms of different aspects. In III, we summarize and provide details for current benchmark datasets for RGB-D salient object detection. In IV, we conduct a comprehensive review of light field SOD models and benchmark datasets. In V, we provide a comprehensive and attribute-based evaluation of several representative RGB-D based models. We then discuss challenges and open directions of this field in VI. Finally, we conclude this paper in VII.
II RGB-D based SOD Models
Over the past few years, several RGB-D based SOD methods have been developed and obtained promising performance. These models are summarized in Tables I, II, III and IV. The complete benchmark can be found at http://dpfan.net/d3netbenchmark/. To review these RGB-D based SOD models in detail,we introduce them from different perspectives as follows. (1) Traditional/deep models: they are viewed from the perspective of feature extraction, that is using the manual features or deep features. It is convenient for follow-up researchers to grasp the historical development trends of RGB-D SOD models. (2) Fusion-wise models: it is critical to effectively fuse RGB and depth images in this task, thus we review different fusion strategies to understand their effectiveness. (3) Single-stream/multi-stream models: we consider this problem from the perspective of model parameters. Single stream can save parameters, but the final result may not be optimal, and multiple streams may require more parameters. Thus, it is helpful to understand the balance between the amount of calculation and accuracy of different models. (4) Attention-aware models: attention mechanisms have widely been applied in various visual tasks including SOD. We review related works on RGB-D SOD to analyze how do different models use attention. Thus, it is an alternative to design attention modules for future works.
Traditional Models. With depth cues, several useful attributes, such as boundary cues, shape attributes, surface normals, etc., can be explored to boost the identification of salient objects in complex scenes. Over the past several years, many traditional RGB-D models based on handcrafted features have been developed . For example, the early work focused on modeling the interaction between layout and shape features generated from the RGB image and depth map. Besides, the representative work developed a novel multi-stage RGB-D model, and constructed the first large-scale RGB-D benchmark dataset, termed NLPR.
Deep Models. However, the above-mentioned methods suffer from unsatisfactory SOD performance due to the limited expression ability of handcrafted features. To address this, several studies have turned to deep neural networks (DNNs) to fuse RGB-D data . These models can learn high-level representations to explore complex correlations across RGB images and depth cues for improving SOD performance. We review some representative works in detail as follows.
DF develops a novel convolutional neural network (CNN) to integrate different low-level saliency cues into hierarchical features, for effectively locating salient regions in RGB-D images. This was the first CNN-based model for the RGB-D SOD task. However, it utilizes a shallow architecture to learn the saliency map.
PCF presents a complementarity-aware fusion module to integrate cross-modal and cross-level feature representations. It can effectively exploit complementary information by explicitly using cross-modal/level connections and modal/level-wise supervision to decrease fusion ambiguity.
CTMF employs a computational model to identify salient objects from RGB-D scenes, utilizing CNNs to learn high-level representations for RGB images and depth cues, while simultaneously exploiting the complementary relationships and joint representation. Besides, this model transfers the structure of the model from the source domain (i.e., RGB images) to be applicable to the target domain (i.e., depth maps).
CPFP proposes a contrast-enhanced network to produce an enhanced map, and presents a fluid pyramid integration module to effectively fuse cross-modal information in a hierarchical manner. Besides, considering the fact that depth cues tend to suffer from noise, a feature-enhanced module is proposed to learn an enhanced depth cue for boosting the SOD performance. It is worth noting that this is an effective solution.
UC-Net proposes a probabilistic RGB-D based SOD network via conditional variational autoencoders (VAEs) to model human annotation uncertainty. It generates multiple saliency maps for each input image by sampling in the learned latent space. This was the first work to investigate uncertainty in RGB-D based SOD, and was inspired by the data labeling process. This method leverages the diverse saliency maps to improve the final SOD performance.
II-B Fusion-wise Models
For RGB-D based SOD models, it is important to effectively fuse RGB images and depth maps. The existing fusion strategies can be grouped into three categories, including 1) early fusion, 2) multi-scale fusion, and 3) late fusion. We provide details for each fusion strategy as follows.
Early Fusion. Early fusion-based methods can follow one of two veins: 1) RGB images and depth maps are directly integrated to form a four-channel input . This is denoted as “input fusion” (shown in Fig. 3); 2) RGB and depth images are first fed into each independent network and their low-level representations are combined as joint representations, which are then fed into a subsequent network for further saliency map prediction . This is denoted as “early feature fusion” (shown in Fig. 3).
Late Fusion. Late fusion-based methods can also be further divided into two families: 1) Two parallel network streams are adopted to learn high-level features for RGB and depth data, respectively, which are concatenated and then used for generating the final saliency prediction . This is denoted as “later feature fusion” (shown in Fig. 3). 2) Two parallel network streams are used to obtain the independent saliency maps for RGB images and depth cues, and then the two saliency maps are concatenated to obtain a final prediction map . This is denoted as “late result fusion” (shown in Fig. 3).
Multi-scale Fusion. To effectively explore the correlations between RGB images and depth maps, several methods propose a multi-scale fusion strategy . These models can be divided into two categories. The first category learn the cross-modal interactions and then fuse them into a feature learning network. For example, Chen et al. developed a multi-scale multi-path fusion network to integrate RGB images and depth maps, with a cross-modal interaction (termed MMCI) module. This method introduces cross-modal interactions into multiple layers, which can empower additional gradients for enhancing the learning of the depth stream, as well as enable complementarity across low-level and high-level representations to be explored. The second category fuse the features from RGB images and depth maps in different layers and then integrate them into a decoder network (e.g., skip connection) to produce the final saliency detection map (as shown in Fig. 3). Some representative works are briefly discussed as follows.
ICNet proposes an information conversion module to convert high-level features in an interactive manner. In this model, a cross-modal depth-weighted combination (CDC) block is introduced to enhance RGB features with depth features at different levels.
DPANet uses a gated multi-modality attention (GMA) module to exploit long-range dependencies. The GMA module can extract the most discriminative features by utilizing a spatial attention mechanism. Besides, this model controls the fusion rate of the cross-modal information using a gate function, which can reduce some effects brought by the unreliable depth cues.
BiANet employs a multi-scale bilateral attention module (MBAM) to capture better global information in multiple layers.
JL-DCF treats a depth image as a special case of a color image and employs a shared CNN for both RGB and depth feature extraction. It also proposes a densely-cooperative fusion strategy to effectively combine the learned features from different modalities.
BBS-Net uses a bifurcated backbone strategy (BBS) to split the multi-level feature representations into teacher and student features, and develops a depth-enhanced module (DEM) to explore informative parts in depth maps from the spatial and channel views.
II-C Single-stream/Multi-stream Models
Single-stream Models. Several RGB-D based SOD works focus on a single-stream architecture to achieve saliency prediction. These models often fuse RGB images and depth information in the input channel or feature learning part. For example, MDSF employs a multi-scale discriminative saliency fusion framework as the SOD model, in which four types of features in three levels are computed and then fused to obtain the final saliency map. BED utilizes a CNN architecture to integrate bottom-up and top-down information for SOD, which also incorporates multiple features, including background enclosure distribution (BED) and low level depth maps (e.g., depth histogram distance and depth contrast) to boost the SOD performance. PDNet extracts depth-based features using a subsidiary network, which makes full use of depth information to assist the main-stream network.
Multi-stream Models. Two-stream models consist of two independent branches that process RGB images and depth cues, respectively, and often generate different high-level features or saliency maps and then incorporate them in the middle stage or end of the two streams. It is worth noting that most recent deep learning-based models utilize this two-stream architecture with several models capturing the correlations between RGB images and depth cues across multiple layers. Moreover, some models utilize a multi-stream structure and then design different fusion modules to effectively fuse RGB and depth information in order to exploit their correlations.
II-D Attention-aware Models
Existing RGB-D based SOD methods often treat all regions equally using the extracted features equally, while ignoring the fact that different regions can have different contributions to the final prediction map. These methods are easily affected by cluttered backgrounds. In addition, some methods either regard the RGB images and depth maps as having the same status or overly rely on depth information. This prevents them from considering the importance of different domains (RGB images or depth cues). To overcome this, several methods introduce attention mechanisms to weight the importance of different regions or domains.
ASIF-Net captures complementary information from RGB images and depth cues using an interweaved fusion, and weights the saliency regions through a deeply supervised attention mechanism.
AttNet introduces attention maps for differentiating between salient objects and background regions to reduce the negative influence of some low-quality depth cues.
TANet formulates a multi-modal fusion framework using RGB images and depth maps from the bottom-up and top-down views. It then introduces a channel-wise attention module to effectively fuse the complementary information from different modalities and levels.
II-E Open-source Implementations
We summarize the open-source implementations of RGB-D based SOD models reviewed in this survey. The implementations and hyperlinks of the source codes of these models are provided in Tab V. More source codes will be updated at: https://github.com/taozh2017/RGBD-SODsurvey.
III RGB-D Datasets
With the rapid development of RGB-D based SOD, various datasets have been constructed over the past several years. Tab VI summarizes nine popular RGB-D datasets, and Fig. 4 shows examples of images (including RGB images, depth maps, and annotations) from these datasets. Moreover, we provide the details for each dataset as follows.
STERE . The authors first collected 1,250 stereoscopic images from Flickr http://www.flickr.com/, NVIDIA 3D Vision Live http://photos.3dvisionlive.com/, and Stereoscopic Image Gallery http://www.stereophotography.com/. The most salient objects in each image were annotated by three users. All annotated images were then sorted based on the overlaping salient regions and the top 1,000 images were selected to construct the final dataset. This is the first collection of stereoscopic images in this field.
GIT consists of 80 color and depth images, which were collected using a mobile-manipulator robot in a real-world home environment. Moreover, each image is annotated based on the pixel-level segmentation of the objects.
DES consists of 135 indoor RGB-D images, which were taken by Kinect with a resolution of . When collecting this dataset, three users were asked to label the salient object in each image, and then the overlapping areas of the labeled object were regarded as the ground truth.
NLPR consists of 1,000 RGB images and their corresponding depth maps, which were obtained by a standard Microsoft Kinect. This dataset includes a series of outdoor and indoor locations, e.g., offices, supermarkets, campuses, streets, and so on.
LFSD includes 100 light fields collected using a Lytro light field camera, and consists of 60 indoor and 40 outdoor scenes. To label this dataset, three individuals were asked to manually segment salient regions, and then the segmented results were deemed ground truth when the overlap of the three results was over .
NJUD consists of 1,985 stereo image pairs, and these images were collected from the internet, 3D movies, and photographs that are taken by a Fuji W3 stereo camera.
SSD was constructed using three stereo movies and includes indoor and outdoor scenes. This dataset includes 80 samples, and each image has the size of .
DUT-RGBD consists of 800 indoor and 400 outdoor scenes with their corresponding depth images. This dataset includes several challenging factors, i.e., multiple or transparent objects, complex backgrounds, similar foregrounds and backgrounds, and low-intensity environments.
SIP consists of 929 annotated high-resolution images, with multiple salient persons in each image. In this dataset, depth maps were captured using a real smartphone (i.e., Huawei Mate10). Besides, it is worth noting that this dataset covers diverse scenes, and various challenging factors, and is annotated with pixel-level ground truths.
Note that a detailed dataset statistics analysis (including center bias, size of objects, background objects, object boundary conditions, and number of salient objects) can be found in .
IV Saliency Detection on Light Field
Existing works for SOD can be grouped into three categories according to the input data type, including RGB SOD, RGB-D SOD, and light field SOD . We have already reviewed RGB-D based SOD models, in which depth maps provide layout information to improve SOD performance to some extent. However, inaccurate or low-quality depth maps often decrease the performance. To overcome this issue, light field SOD methods have been proposed to make use of rich information captured by the light field. Specifically, light field data contains an all-focus image, a focal stack, and a rough depth map . A summary of related light field SOD works is provided in Tab VII. Further, to provide an in-depth understanding of these models, we also review them in more detail as follows.
Traditional/Deep Models. The classic models for light field SOD often use superpixel-level handcrafted features . Early work showed that the unique refocusing capability of light fields can provide useful focusness, depth, and objectness cues. Thus, several SOD models using light field data were further proposed. For example, Zhang et al. utilized a set of focal slices to compute the background prior, and then combined it with the location prior for SOD. Wang et al. proposed a two-stage Bayesian fusion model to integrate multiple contrasts for boosting SOD performance. Recently, several deep learning-based light field SOD models have also been developed, obtaining remarkable performance. Besides, in , an attentive recurrent CNN was developed to fuse all focal slices, while the data diversity was increased using adversarial examples to enhance model robustness. Zhang et al. developed a memory-oriented decoder for light field SOD, which fuses multi-level features in a top-down manner using high-level information to guide low-level feature selection. LFNet employs a new integration module to fuse features from light field data according to their contributions and captures the spatial structure of a scene to improve SOD performance.
Refinement based Models. Several refinement strategies have been used to enforce neighboring constraints or reduce the homogeneity of multiple modalities for SOD. For example, in , the saliency dictionary was refined using the estimated saliency map. The MA method employs a two-stage saliency refinement strategy to produce the final prediction map, which enables adjacent superpixels to obtain similar saliency values. Besides, LFNet presents an effective refinement module to reduce the homogeneity among different modalities as well refine their dissimilarities
IV-B Light Field Data for SOD
There are five representative datasets widely used in existing light field SOD models. We describe the details of each dataset as follows.
LFSD https://sites.duke.edu/nianyi/publication/saliency-detection-on-light-field/ consists of 100 light fields of different scenes with a spatial resolution, captured using a Lytro light field camera. This dataset contains 60 indoor and 40 outdoor scenes, and most scenes consist of only one salient object. Besides, three individuals were asked to manually segment salient regions in each image, and then the ground truth was determined when all three segmentation results had an overlap of over 90%.
HFUT https://github.com/pencilzhang/HFUT-Lytro-dataset consists of 255 light fields captured using a Lytro camera. In this dataset, most scenes contain multiple objects that appear within different locations and scales under complex background clutter.
DUTLF-FS https://github.com/OIPLab-DUT/ICCV2019_Deeplightfield_Saliency consists of 1,465 samples, 1,000 of which are used as the training set, while the remaining 465 images make up the test set. The resolution of each image is . This dataset contains several challenges, including lower contrast between salient objects and cluttered background, multiple disconnected salient objects, and dark or strong light conditions.
DUTLF-MV https://github.com/OIPLab-DUT/IJCAI2019-Deep-Light-Field-Driven-Saliency-Detection-from-A-Single-View consists of 1,580 samples, 1,100 of which are for training and the remaining is for testing. Images were captured by a Lytro Illum camera, and each light field consists of multi-view images and a corresponding ground truth.
Lytro Illum https://github.com/pencilzhang/MAC-light-field-saliency-net consists of 640 light fields and the corresponding per-pixel ground-truth saliency maps. It includes several challenging factors, e.g., inconsistent illumination conditions, and small salient objects existing in a similar or cluttered background.
V Model Evaluation and Analysis
We briefly review several popular metrics for SOD evaluation, i.e., precision-recall (PR), F-measure , mean absolute error (MAE) , structural measure (S-measure) , and enhanced-alignment measure (E-measure) .
PR. Given a saliency map , we can convert it to a binary mask , and then compute the precision and recall by comparing with ground-truth :
A popular strategy is to partition the saliency map using a set of thresholds (i.e., it changes from 0 to 255). For each threshold, we first calculate a pair of recall and precision scores, and then combine them to obtain a PR curve that describes the performance of the model at the different thresholds.
F-measure (). To comprehensively consider both precision and recall, the F-measure is proposed by calculating the weighted harmonic mean:
where is set to 0.3 to emphasize the precision . We use different fixed $FFF_{\beta}$.
MAE. This measures the average pixel-wise absolute error between a predicted saliency map and a ground truth for all pixels, which can be defined by
where and denote the width and height of the map, respectively. MAE values are normalized to .
S-measure (). To capture the importance of the structural information in an image, is used to assess the structural similarity between the regional perception () and object perception (). Thus, can be defined by
where is a trade-off parameter. Here, we set = 0.5 as the default setting, as suggested by Fan et al. .
E-measure (). was proposed based on cognitive vision studies to capture image-level statistics and their local pixel matching information. Thus, can be defined by
where denotes the enhanced-alignment matrix .
V-B Performance Comparison and Analysis
To quantify the performance of different models, we conduct a comprehensive evaluation of 24 representative RGB-D based SOD models, including 1) nine traditional methods: LHM , ACSD , DESM , GP , LBE , DCMC , SE , CDCP , CDB ; and 2) fifteen deep learning-based methods: DF , PCF , CTMF , CPFP , TANet , AFNet , MMCI , DMRA , D3Net , SSF , A2dele , S2MA , ICNet , JL-DCF , and UC-Net . We report the mean values of and MAE across the five datasets (STERE , NLPR , LFSD , DES , and SIP ) for each model in Fig. 5. It is worth noting that better models are shown in the upper left corner (i.e., with a larger and smaller MAE). From Fig. 5, we have following observations:
Traditional vs. Deep Models. Compared with traditional RGB-D based SOD models, deep learning methods obtain significantly better performance. This confirms the powerful feature learning ability of deep networks.
Comparison of Deep Models. Among the deep learning-based models, D3Net , JL-DCF , UC-Net , SSF , ICNet , and S2MA obtain the best performance.
Moreover, Fig. 6 and Fig. 7 show the PR and F-measure curves for the 24 representative RGB-D based SOD models on eight datasets (i.e., STERE , NLPR , LFSD , DES , SIP , GIT , SSD , and NJUD ). Note that there are 1000, 300, 100, 135, 929, 80, and 80 test samples for NLPR, LFSD, DES, SIP, GIT, and SSD, respectively. For the NJUD dataset, there are 485 test images for CPFP , S2MA , ICNet , JL-DCF , and UC-Net , while 498 testing images for all other models.
To understand the top six models in depth, we discuss their main advantages for the six models below.
D3Net consists of two key components, i.e., a three-stream feature learning module and a depth depurator unit. In the three-stream feature learning module, there are three subnetworks, i.e., RgbNet, RgbdNet, and DepthNet. The RgbNet and DepthNet are used to learn high-level feature representations for RGB and depth images, respectively, while the RgbdNet is used to learn their fused representations. It is worth noting that this three-stream feature learning module can capture modality-specific information as well as the correlation between modalities. Thus, balancing the two aspects is very important for multi-modal learning and it has helped to improve the SOD performance. Besides, the depth depurator unit acts as a gate to explicitly filter out low-quality depth maps, which several existing methods do not consider the effects. Because low-quality depth maps can inhibit the fusion between RGB images and depth maps, thus the depth depurator unit can ensure effective multi-modal fusion to achieve robust SOD performance.
In JL-DCF , there are two key components, i.e., a joint learning (JL) and a densely-cooperative fusion (DCF). Specifically, the JL module is used to learn robust saliency features, while the DCF module is used for complementary feature discovery. It is worth noting that this method uses a middle-fusion strategy to extract deep hierarchical features from RGB images and depth maps, in which the cross-modal complementarity can be effectively exploited to achieve accurate prediction.
In UC-Net , instead of producing a single saliency prediction, this model produces multiple predictions by modeling the distribution of the feature output space as a generative model conditioned on RGB-D images. Because each person has some specific preferences in labeling a saliency map, it could fail to capture the stochastic characteristic of saliency while only a single saliency map is produced for an image pair using a deterministic learning pipeline. Thus, the strategy in this model can take into account human uncertainty in saliency annotations. Moreover, considering the fact that depth maps could suffer from noise, directly fusing RGB images and depth maps could cause the network to fit to this noise. Therefore, a depth correction network, designed as an auxiliary component, is proposed to refine depth information with a semantic guided loss. Thus, the above key components are all helpful for improving SOD performance.
In SSF , a complementary interaction module (CIM) is developed to explore discriminative cross-modal complementarities and fuse cross-modal features, where a region-wise attention is introduced to supplement rich boundary information for each modality. Besides, a compensation-aware loss is proposed to improve the network’s confidence for hard samples in unreliable depth maps. Thus, these key components enable the proposed model to effectively explore and establish the complementarity of cross-modal feature representations, while at the same time reducing the negative effects introduced by low-quality depth maps, boosting SOD performance.
In ICNet , an information conversion module is proposed to interactively and adaptively explore the correlations between high-level RGB and depth features. Besides, a cross-modal depth-weighted combination block is introduced to enhance the difference between the RGB and depth features in each level, which ensures that the features are treated differently. It is also worth noting that ICNet exploits the complementarity of cross-modal features, as well as explores the continuity of cross-level features, both of which are helpful for achieving accurate predictions.
In S2MA , a self-mutual attention module (SAM) is proposed to fuse RGB and depth images, integrating self-attention and each other’s attention to propagate context more accurately. The SAM can provide additional complementary information from multi-modal data to improve SOD performance, overcoming the limitations of the original self-attention, which only uses a single modality. Besides, to reduce the low-quality (e.g., noise) effects of depth cues, a selection mechanism is proposed to reweight the mutual attention. This mechanism can filter out unreliable information, resulting in more accurate saliency prediction.
V-B2 Attribute-based Evaluation
To investigate the influence of different factors, such as object scale, background clutter, number of salient objects, indoor or outdoor scene, background objects, and lighting conditions, we carry out diverse attribute-based evaluations on several representative RGB-D based SOD models.
Object Scale. To characterize the scale of a salient object area, we compute the ratio between the size of the salient area and the whole image. We define three types of object scales: 1) when the ratio is less than 0.1, it is denoted as “small”; 2) when the ratio is larger than 0.4, it is denoted as “large”; and 3) when the ratio is in the range of , it is denoted as “medium”. In this evaluation, we build a hybrid dataset with 2,464 images collected from STERE , NLPR , LFSD , DES , and SIP , where 24%, 69.2% and 6.8% of images have small, medium, and large salient object areas, respectively. The constructed hybrid dataset can be found at https://github.com/taozh2017/RGBD-SODsurvey. Some sample images with different object scales are shown in Fig. 8. The comparison results of the attribute-based study w.r.t. object scale are shown in Tab. VIII. From the results, it can be observed that all comparison methods obtain better performance in detecting small salient objects while they obtain worse performance in detecting large salient objects. Besides, the three most recent models, i.e., JL-DCF , UC-Net , and S2MA , obtain the best performance. D3Net , SSF , A2dele , and ICNet also obtain promising performance.
Background Clutter. It is difficult to directly characterize background clutter. Since classic SOD methods tend to use prior information or color contrast to locate salient objects, they often fail under complex backgrounds. Thus, in this evaluation, we utilize five traditional SOD methods, i.e., BSCA , CLC , MDC , MIL , and WFD , to first detect salient objects in various images and then group these images into different categories (e.g., simple or complex background) according to the results. Specifically, we first construct a hybrid dataset with 1,400 images collected from three datasets (STERE , NLPR , and LFSD ). Then, we apply the five models to this dataset and obtain the values for each, which we use to characterize images as follows: 1) If all values are higher than , the image is denoted as having a “simple” background; 2) If all values are lower than , the image is said to have a “complex” background; 3) The remaining images are denoted as “uncertain”. Some example images with the three types of background clutter are shown in Fig. 9. The constructed hybrid dataset can be found at https://github.com/taozh2017/RGBD-SODsurvey. The comparison results of the attribute-based study w.r.t. background clutter are shown in Tab. IX. As can be seen, all models obtain worse SOD performance on images containing complex backgrounds than simple ones. Among the representative models, JL-DCF , UC-Net and SSF achieve the top-three best results. Besides, the four most recent models, i.e., D3Net , S2MA , A2dele , and ICNet , also obtain better performance than the other models.
Single vs. Multiple Objects. In this evaluation, we construct a hybrid dataset with 1,229 images collected from the NLPR and SIP datasets. Some example images with single or multiple salient objects are shown in Fig. 10. The comparison results are shown in Fig. 11. From the results, we can see that it is easier to detect single salient object than multiple ones.
Indoor vs. Outdoor. We evaluate the performance of different RGB-D based SOD models on indoor and outdoor scenes. In this evaluation, we construct a hybrid dataset collected from the DES , NLPR , and LFSD datasets. The comparison results are shown in Fig. 12. From the results, it can be seen that most models struggle more to detect salient objects in indoor scene than outdoor scenes. This is possibly because indoor environments often have varying light conditions.
Background Objects. We evaluate the performance of the RGB-D based SOD models when different background objects are present. We use the SIP dataset , and split it into nine categories, i.e., car, barrier, flower, grass, road, sign, tree, and other. The comparison results are shown in Tab. X. As can be seen, all methods obtain diverse performances under different background objects. Among the 24 representative RGB-D based models, JL-DCF , UC-Net and SSF achieve the top-three best results. In addition, the four most recent models, i.e., D3Net , S2MA , A2dele , and ICNet obtain better performance than the others.
Lighting Conditions. The performance of SOD can be affected by different lighting conditions. To determine the performance of different RGB-D based SOD models under different lighting conditions, we conduct an evaluation on the SIP dataset , which we split it into two categories, i.e., sunny and low-light. The comparison results are shown in Tab. XI. As can be seen, low-light negatively impacts SOD performance. Among comparison models, UC-Net obtains the best performance under sunny conditions while JL-DCF achieves the best result under low-light condition.
In addition, we report the saliency maps generated for various challenging scenes to visualize the performance of different RGB-D based SOD models. Fig. 13 and Fig. 14 show some representative examples using two classic non-deep methods (DCMC and SE ) and eight state-of-the-art CNN-based models (DMRA , D3Net , SSF , A2dele , S2MA , ICNet , JL-DCF , and UC-Net ). The row shows a small object, while the row is an example of a large one. The and rows contain complex backgrounds and boundaries, respectively. The and rows contain multiple salient objects. In the row, there are low-light condition. In the row, the depth map is coarse with very inaccurate object boundaries, which could inhibit the SOD performance. From the results in Fig. 13 and Fig. 14, it can be observed that deep models perform better than non-deep models on these challenging scenes, confirming the powerful expression ability of deep features over handcrafted ones. In addition, D3Net , S2MA , JL-DCF , and UC-Net perform better than other deep models.
VI Challenges and Open Directions
Effects of Low-quality Depth Maps. Depth maps with affluent spatial information have been proven beneficial in detecting salient objects from cluttered backgrounds, while the depth quality also directly affects the subsequent SOD performance. The quality of depth maps varies tremendously across different scenarios due to the limitations of depth sensors, posing a challenge when trying to reduce the effects of low-quality depth maps. However, most existing methods directly fuse RGB images and original raw data from depth maps, without considering the effects of low-quality depth maps. There are a few notable exceptions. For example, in , a contrast-enhanced network was proposed to learn enhanced depth maps, which have much higher contrasts compared with the original depths. In , a compensation-aware loss was designed to pay more attention to hard samples containing unreliable depth information. Moreover, D3Net uses a depth depurator unit (DDU) to classify depth maps into two classes (i.e., reasonable and low-quality). The DDU also acts as a gate that can filter out the low-quality depth maps. However, the above methods often employ a two-step strategy to achieve depth enhancement and multi-modal fusion or an independent gate operation for filtering out poor depths, which could bring a suboptimal problem. There is thus a need to develop an end-to-end framework that can achieve depth enhancement or adaptively weight the depth maps (e.g., assign low weights to poor depth maps) during multi-modal fusion, which would be more helpful for reducing the effects of low-quality depth maps and boosting SOD performance.
Incomplete Depth Maps. In RGB-D datasets, it is inevitable for there to be some low-quality depth maps due to the limitations of the acquisition devices. As previously discussed, several depth enhancement algorithms have been used to improve the quality of depth maps. However, depth maps that suffer from severe noise or blurred edges, are often discarded. In this case, we have complete RGB images but some samples do not have depth maps, which is similar to the incomplete multi-view/modal learning problem . Thus, we call it “incomplete RGB-D based SOD”. As current models only focus on the SOD task using complete RGB images and depth maps, we believe this could be a new direction for RGB-D SOD.
Depth Estimation. Depth estimation provides an effective solution to recover high-quality depths and overcome the effects of low-quality depth maps. Various depth estimation approaches have been developed, which could be introduced into the RGB-D based SOD task to improve performance.
VI-B Effective Fusion Strategies
Adversarial Learning-based Fusion. It is important to effectively fuse RGB images and depth maps for RGB-D based SOD. Existing models often employ different fusion strategies (e.g., early fusion, middle fusion, or late fusion) to exploit the correlations between RGB images and depth maps. Recently, generative adversarial networks (GANs) have gained widespread attention for the saliency detection task . In common GAN-based SOD models, a generator takes RGB images as inputs and generates the corresponding saliency maps, while a discriminator is adopted to determine whether a given image is synthetic or ground-truth. GAN-based models could easily be extended to RGB-D SOD, which could be helpful for boosting performance due to their superior feature learning ability. Moreover, GANs could also be used to learn common feature representations for RGB images and depth maps , which could help with feature or saliency map fusion and further boost the SOD performance.
Attention-induced Fusion. Attention mechanisms have been widely applied to various deep learning-based tasks , allowing networks to selectively pay attention to a subset of regions for extracting discriminative and powerful features. Besides, co-attention mechanisms have been developed to explore the underlying correlations across multiple modalities, and are widely studied in visual question answering and video object segmentation . Thus, for the RGB-D based SOD task, we could also develop attention-based fusion algorithms to exploit correlations between RGB images and depth cues to improve the performance.
VI-C Different Supervision Strategies
Existing RGB-D models often use a fully supervised strategy to learn saliency prediction models. However, annotating pixel-level saliency maps is a tedious and time-consuming procedure. To alleviate this issue, there has been increased interest in weakly and semi-supervised learning, which have been applied to salient object detection . Semi-/weak supervision could also be introduced into RGB-D SOD, by leveraging image-level tags and pseudo pixel-wise annotations , for improving the detection performance. Besides, several studies have suggested that models pretrained using self-supervision can effectively be used to achieve better performance. Therefore, we could train saliency prediction models on large amounts of annotated RGB images in a self-supervised manner and then transfer the pretrained models to the RGB-D SOD task.
VI-D Dataset Collection
Dataset size. Although there are nine public RGB-D datasets for SOD, their size is quite limited, e.g., the maximum size is about 2,000 samples for NJUD . When compared with other RGB-D datasets for generic object detection or action recognition , the size of RGB-D datasets for SOD is also very small. Thus, it is essential to develop new large-scale RGB-D datasets that can serve as baselines for future research.
Complex Background & Task-driven Datasets. Most existing RGB-D datasets collect images that contain one salient object or multiple objects but with a relatively clean background. However, real-world applications often suffer from much more complicated situations (e.g., occlusion, appearance change, low illumination, etc), which could decrease the SOD performance. Thus, collecting images with complex background is critical to improve the generalization ability of RGB-D SOD models. Moreover, for some tasks, images with specific salient object(s) must be collected. For example, one important technology is road sign recognition in driver assistance systems, which requires images with road signs to be collected. Thus, it is essential to construct task-driven RGB-D datasets like SIP .
VI-E Model Design for Real-world Scenarios
Some smartphones can capture depth maps (e.g., images in the SIP dataset were captured using Huawei Mate 10). Thus it would be feasible to conduct the SOD task in real-world applications, e.g., on smart devices. However, most existing methods include complicated and deep DNNs to increase the model capacity and achieve better performance, preventing them from being directly applied on real-work platforms. To overcome this, model compression techniques could be used to learn compact RGB-D based SOD models with promising detection accuracy. Moreover, JL-DCF utilizes a shared network to locate salient objects using RGB and depth views, which largely reduces the model parameters and makes real-world applications feasible.
VI-F Extension to RGB-T SOD
In addition to RGB-D SOD, there are several other methods that fuse different modalities for better detection, such as RGB-T SOD, which integrates RGB and thermal infrared data. Thermal infrared cameras can capture the radiation emitted from any object with a temperature above absolute zero, making thermal infrared images insensitive to illumination conditions . Therefore, thermal images can provide supplementary information to improve SOD performance when salient objects suffer from varying light, reflective light, or shadows. Some RGB-T models and datasets (VT821 , VT1000 and VT5000 ) have already been proposed over the past few years. Similar to RGB-D SOD, the key aim of RGB-T SOD is to fuse RGB and thermal infrared images and exploit the correlations between the two modalities. Thus, several advanced multi-modal fusion technologies in RGB-D SOD could be extended to the RGB-T SOD task.
VII Conclusion
In this paper we present, to the best of our knowledge, the first comprehensive review of RGB-D based SOD models. We first review the models from different perspectives, and then summarize popular RGB-D SOD datasets as well as provide details for each. Considering the fact that light fields also provide depth information, we also review popular light field SOD models and the related benchmark datasets. Next, we provide a comprehensive evaluation of 24 representative RGB-D based SOD models as well as an attribute-based evaluation. Specifically, we perform attribute-based performance analysis by constructing new datasets for the 24 representative RGB-D based SOD models. Moreover, we discuss several challenges and highlight open directions for future research. In addition, we briefly discuss the extension work to RGB-T SOD to improve performance when salient objects suffer from varying light, reflective light, or shadows. Although RGB-D based SOD has made notable progress over the past several decades, there is still significant room for improvement. We hope this survey will generate more interest in this field.