Recurrent Attentional Networks for Saliency Detection

Jason Kuen, Zhenhua Wang, Gang Wang

Introduction

Saliency detection refers to the challenging computer vision task of identifying salient objects in imagery and segmenting their object boundaries. Despite that it has been studied for years, saliency detection still remains an unsolved research problem due to its tough goal to model high-level subjective human perceptions. Recently, saliency detection methods have received considerable amount of attention, as there is a wide and growing range of applications facilitated by it. Some of the notable applications of saliency detection are object recognition ren2014region, visual tracking borji2012adaptive, and image retrieval chen2009sketch2photo.

Traditionally, methods in saliency detection leverage low-level saliency priors such as contrast prior and center prior to model and approximate human saliency. However, such low-level priors can hardly capture high-level information about the objects and its surroundings: the traditional methods are still very far away from how saliency works in the context of human perceptions. To incorporate high-level visual concepts into a saliency detection framework, it is natural to consider convolutional neural networks (CNN). For a lot of computer vision tasks gu2015recent, CNNs have shown to be remarkably effective. It is also the first learning algorithm to achieve human-competitive performances he2015delving in large-scale image classification task, which is a high-level vision task like saliency detection. Although there have been works on developing CNNs for visual saliency modeling, they either focus on predicting eye fixations nian2015predicting, or applying CNNs to predict just the saliency value of visual sub-units (e.g. superpixels) independently zhao2015saliency. Besides, conventional CNNs downsize feature maps over multiple convolutional and pooling layers and lose detailed information for our problem of densely segmenting salient objects.

Inspired by the success of convolutional-deconvolutional network (CNN-DecNN) in semantic segmentation noh2015learning, in this paper, we adapt the network to detect salient objects in an end-to-end fashion. For this framework, the input is an image, and the output is its corresponding saliency map. A deconvolutional network (DecNN) is a variant of CNN that performs convolution and unpooling to produce dense pixel-precise outputs. However, CNN-DecNN works poorly for objects of multiple scales long2015fully; noh2015learning due to the fixed-size receptive fields. To overcome this limitation, we propose a recurrent attentional convolutional-deconvolutional network (RACDNN) to refine the saliency maps generated by CNN-DeCNN. RACDNN uses spatial transformer and recurrent network units to iteratively attend to flexibly-sized image sub-regions, and refines the saliency predictions on those sub-regions. As shown in Figure 1, RACDNN can perform saliency detection at finer scales due to its ability to attend to smaller sub-regions. Another advantage of RACDNN is that the attended sub-regions in the previous iterations can provide contextual information for the saliency refinement of the sub-region in the current iteration. For example, in Figure 1, RACDNN can make use of the more visible front legs of the deers to help at refining the saliency values of the less-visible back legs.

We perform experiments on several challenging saliency detection benchmark datasets, and compare the proposed method with state-of-the-art saliency detection methods. Experimental results show the effectiveness of our proposed method.

Related work

Saliency detection methods can be coarsely categorized into bottom-up and top-down methods. Bottom-up methods itti1998model; harel2006graph; hou2007saliency; achanta2009frequency; liu2011learning; cheng2015global; margolin2013what make use of level local visual cues like color, contrast, orientation and texture. Top-down methods zhang2008sun; yang2012top; judd2009learning are based on high-level task-specific prior knowledge. Recently, deep learning-based saliency detection methods wang2015deep; zhang2015co; zhao2015saliency; li2015visual; vig2014large have been very successful. Instead of manually defining and tuning saliency-specific features, these methods can learn both low-level features and high-level semantics useful for saliency detection straight from minimally processed images. However, these works employ neither attention mechanism nor RNN to improve saliency detection. To the best of our knowledge, ours is the first work to exploit recurrent attention along with deep learning for saliency detection.

Attention models are a new variant of neural networks aiming to model visual attention. They are often used with recurrent neural networks to achieve sequential attention. mnih2014recurrent formulates a recurrent attention model that surpasses CNN on some image classification tasks. ba2015multiple extends the work of mnih2014recurrent by making the model deeper and apply it for multi-object classification task. To overcome the training difficulty of recurrent attention model, gregor2015draw propose a differentiable attention mechanism and apply it for generative image generation and image classification. jaderberg2015spatial propose a differentiable and efficient sampling-based spatial attention mechanism, in which any spatial transformation can be used. Unlike the above works mnih2014recurrent; ba2015multiple; gregor2015draw which mostly use small attention networks for low-resolution digit classification task, the attention mechanism used in our work is much more complex, as it is tied with a large CNN-DecNN for dense pixelwise saliency refinement.

Proposed Method

In this section, we describe our proposed saliency detection method in detail. In our method, initial saliency maps are first generated by a convolutional-deconvolutional network (CNN-DecNN) which takes entire images as input, and outputs saliency maps. The saliency maps are then refined iteratively via another CNN-DecNN operated under a recurrent attentional framework. Unlike the initial saliency map prediction which is done through single feedforward passes on the entire images, the saliency refinement is done locally on selected image sub-regions in a progressive way. At every processing iteration, the recurrent CNN-DecNN attends to an image sub-region, through the use of a spatial transformer-based attention mechanism. The attentional saliency refinement helps to alleviate the inability of CNN-DecNN to deal with multiscale saliency detection. In addition, the sequential nature of the attention enables the network to exploit contextual patterns from past iterations to enhance the representation of the attended sub-region, hence to improve the saliency detection performance.

Conventionally, CNNs downsize feature maps over multiple convolutional and pooling layers, to construct spatially compact image representations. Although these spatially compact feature maps are well-suited for whole-image classification tasks, they tend to produce very coarse outputs when being applied for dense pixelwise prediction tasks (e.g., semantic segmentation). To tackle dense prediction tasks in the multi-layered convolutional learning setting, one can append a deconvolutional network (DecNN) to a CNN as shown in noh2015learning. In such a convolutional-deconvolutional (CNN-DecNN) framework, the CNN learns globally meaningful representations, while the DecNN upsizes feature maps and learns increasingly localized representations. Unlike the work of noh2015learning, we preserve the spatial information of CNN’s output (the input to DecNN) by using only convolutional layers. In practice, we find that preserving such spatial information works better than without preserving it. This is because the preserved spatial information provides a good head start for DecNN to gradually introduce more spatial information to the feature maps. A generic network architecture of CNN-DecNN is shown in Figure 2.

A DecNN is almost identical to conventional CNNs except for a few minor differences. Firstly, in deconvolutional networks, convolution operations are often carried out in such a way that the resulting feature maps retain the same spatial sizes as those of the input feature maps. This is done by adding appropriate zero paddings beforehand. Secondly, the pooling operators adopted by CNNs are substituted with unpooling operators in DecNNs. Given input feature maps, unpooling operators work by upsizing the feature maps, contrary to what pooling operators achieve. A few variants of unpooling methods dosovitskiy2015learning; noh2015learning have been proposed previously to tackle several computer vision tasks involving spatially large and dense outputs. In this paper, we employ the simple unpooling method demonstrated in dosovitskiy2015learning, whereby each block (with spatial size 1×11\times 1) in the input feature maps is mapped to the top left corner of a blank output block with spatial size k×kk\times k. This effectively increases the spatial size of the whole feature maps by a factor of kk.

In the processing pipeline of CNN-DecNN for saliency detection, the CNN first transforms the input image xx to a spatially compact hidden representation zz, as z=\mboxCNN(x)z=\mbox{{CNN}}(x). Then, zz is transformed to a raw saliency map rr through the DecNN, as r=\mboxDecNN(z)r=\mbox{{DecNN}}(z). To obtain the final saliency map Sˉ\bar{S} that lies within the probability range of $,weperform, we perform\bar{S}=\sigma(r),passingtherawsaliencymap, passing the raw saliency maprintoelement−wisesigmoidactivationfunctioninto element-wise sigmoid activation function\sigma(\cdot).Giventhegroundtruthsaliencymap. Given the groundtruth saliency map\bar{G},thelossfunctionofCNN−DecNNforsaliencydetectionisthebinarycross−entropybetween, the loss function of CNN-DecNN for saliency detection is the binary cross-entropy between\bar{G}andand\bar{S}$. The resulting network can be trained in end-to-end fashion to perform saliency detection. Although CNN-DecNN can achieve pixelwise labeling, it works poorly for objects of multiple scales long2015fully; noh2015learning due to the fixed-size receptive fields used. Furthermore, long-distance contextual information which is important for saliency detection, cannot be well captured by the locally applied convolution filters in DecNN. To address these issues, we propose an recurrent attentional network that iteratively attends to image sub-regions (of unconstrained scale and location) for saliency refinement, which is described in the next two subsections.

2 Attentional Inputs and Outputs with Spatial Transformer

To realize the attention mechanism for saliency refinement, we adopt the spatial transformer network proposed in jaderberg2015spatial. Spatial transformer is a sub-differentiable sampling-based neural network which spatially transform its input feature maps (may also be images), resulting in an output feature maps that is an attended region of the input feature maps. Due to its differentiability, spatial transformer is relatively easier to train compared to some non-differentiable neural network-based attention mechanisms mnih2014recurrent; ba2015multiple proposed recently.

where asa_{s}, atxa_{tx}, and atya_{ty} are the scaling, horizontal translation, and vertical translation parameters respectively. Aligning with the recent works mnih2014recurrent; ba2015multiple; gregor2015draw in recurrent visual attention modeling, the parameters deciding where the attention takes place (in our case, τ\tau) is produced by the localization network floc(⋅)f_{loc}(\cdot). More details on floc(⋅)f_{loc}(\cdot) will be introduced in Equation 9 in Section 3.3. Subsequently, the transformation matrix τ\tau is applied to the regular coordinates of VV to obtain sampling coordinates. Based on the sampling coordinates, VV is formed by sampling feature map points from UU using bilinear interpolation.

Generally, attention mechanisms are applied only to input images. However, our saliency refinement method (see Section 3.3) via DecNN demands that the input and output ends point to the same image sub-region. To this end, we propose an inverse spatial transformer which can map refined saliency output back to the same sub-region attended at input end. Assuming that τ\tau is the transformation matrix for the input end, the inverse spatial transformer takes the inverse of τ\tau as the output transformation matrix τ−1\tau^{-1}:

3 Recurrent Attentional Networks for Saliency Refinement

Recurrent neural networks (RNN) elman1990finding are a class of neural networks developed for modeling the sequential dependencies between sub-instances of sequential data. In RNN, the hidden state hih_{i} at time step or iteration ii is computed as a non-linear function of the input and the previous iteration’s hidden state hi−1h_{i-1}. Given an input xix_{i} at iteration ii, the hidden state hih_{i} of a RNN is formulated as:

where WIW_{I} and WRW_{R} are the learnable weights for input-to-hidden and hidden-to-hidden connections respectively, while bb is a bias term, and ϕ(⋅)\phi(\cdot) is a nonlinear activation function. By explicitly making the current hidden state hih_{i} dependable on the previous hidden state hi−1h_{i-1} , RNN is able to encode contextual information gained from past iterations for use in future iterations. As a result, a more powerful representation hih_{i} can be learned.

In this work, we combine the recurrent computational structure of RNN with CNN-DecNN as well as the spatial transformer attention mechanism, to establish the recurrent attentional convolutional-deconvolutional networks (RACDNN). As illustrated in Figure 4, given an intiail saliency map produced by the initial CNN-DeCNN, RACDNN iteratively uses spatial transformer to attend to a sub-region, and applies its CNN-DecNN to perform saliency refinement for the attended sub-region, by learning powerful context-aware features using RNN.

At every computational iteration ii, RACDNN first receives an attended input xix_{i} from the full input image xx as follows:

where ST(⋅\cdot) is a spatial transformer function which produces an output image sampled from the input image, given the transformation matrix τi\tau_{i}. τi\tau_{i} is computed at the previous iteration i−1i-1 through the localization network floc(⋅)f_{loc}(\cdot). Then, RACDNN uses a recurrent-based CNN \mboxCNNr\mbox{{CNN}}_{r} to encode the attended input xix_{i} into a spatially-compact hidden representation ziz_{i}. \mboxCNNr\mbox{{CNN}}_{r} is similar to CNN except that \mboxCNNr\mbox{{CNN}}_{r} is used in the recurrent setting, and all recurrent instances of \mboxCNNr\mbox{{CNN}}_{r} share the same network parameters. To form the recurrent hidden state hi1h^{1}_{i} of iteration ii, the representation ziz_{i} is combined with the hidden state hi−11h^{1}_{i-1} of the previous iteration:

where WI1W^{1}_{I} is the convolution filters for input-to-hidden connections, WR1W^{1}_{R} is the convolution filters for hiddent-to-hidden connections between any two consecutive iterations, b1b^{1} is a bias term. As in RNN, the hidden-to-hidden connections allow contextual information gathered at previous iterations to be passed to the future iterations. Since RACDNN is attentional, the already attended sub-regions can help to guide saliency refinement for the upcoming sub-regions. This is beneficial for the task of saliency detection, as the saliency of an object is highly dependable on its surrounding regions. Different from conventional RNNs that use matrix product (fully-connected network layers) for both input-to-hidden and hidden-to-hidden connections, these connections in our method are convolution operations (convolutional layers) as in pinheiro2014recurrent. By using recurrent connections that are convolutional, we can preserve the spatial information of hidden representation hi1h^{1}_{i}. As mentioned in Section 3.1, preserving the spatial information of hidden representation between CNN and DecNN is favorable for DecNN’s upsizing-related operations.

After obtaining hi1h^{1}_{i}, we can then perform saliency refinement on initial saliency maps using \mboxDecNNr\mbox{{DecNN}}_{r}. The initial saliency maps are generated by the global CNN-DecNN in single forward passes. Instead of replacing the values of initial saliency map with the output of RACDNN at each iteration, the initial saliency map r0r_{0} is refined cumulatively for NN number of iterations. At iteration ii, the saliency map rir_{i} is refined as

Before being added to rir_{i}, the saliency output of \mboxDecNNr\mbox{{DecNN}}_{r} is spatially transformed back to the attended sub-region using inverse spatial transformer (STST). For the unattended regions, the saliency refinement values are set as zero and thus those regions do not affect rir_{i}. After NN number of iterations, as in Section 3.1, sigmoid activation function σ(⋅)\sigma(\cdot) is applied to rNr_{N}, resulting in the final saliency map Sˉr\bar{S}_{r}.

Besides saliency refinement outputs, at every iteration, RACDNN should generate τ\tau to determine which sub-region to attend to in the next iteration. A simple way to achieve that is by simply treating hi1h^{1}_{i} as input to a fully-connected network-based regressor. However, to model the sequential dependencies between attended locations, such a simplistic approach is insufficient. This is because hi1h^{1}_{i} should focus mainly on modeling contextual dependencies for saliency refinement, not multiple kinds of dependency. To better model locational dependencies, we propose to add another recurrent layer to RACDNN. The hidden state of the second recurrent layer at iteration ii is denoted by hi2h^{2}_{i} and it is formulated as

where the weights WI2,WR2W^{2}_{I},W^{2}_{R} and bias b2b^{2} are semantically the same as their counterparts in the first recurrent layer in Equation (5). The input of the second recurrent layer is the output of the first recurrent layer, making the RACDNN a stacked recurrent network. Considering the nature of the regression task, we use only fully-connected layers for both recurrent input and hidden connections in the second recurrent layer. Finally, given hi2h^{2}_{i}, a floc(⋅)f_{loc}(\cdot) can be used to regress the transformation matrix for the next iteration i+1i+1:

Wloc1W_{loc^{1}} and Wloc2W_{loc^{2}} are respectively the weight matrices of the first and second layers of the two-layered fully-connected network floc(⋅)f_{loc}(\cdot) used in our work.

In RACDNN, the hidden representations (h01,h02)(h^{1}_{0},h^{2}_{0}) at the 00-th iteration are provided by a CNN (sharing the same architectural properties as \mboxCNNr\mbox{{CNN}}_{r}) which accepts the whole image region as input. Observing the full image region at the 00-th iteration helps RACDNN to better decide which sub-regions to attend subsequently.

Similar to the CNN-DecNN used for saliency detection, the loss function of RADCNN is the binary cross-entropy between the final saliency output Sˉr\bar{S}_{r} and the groundtruth saliency map Gˉ\bar{G}. Since every component in RADCNN is differentiable, errors can be backpropagated to all network layers and parameters of RADCNN, making it trainable with any gradient-based optimization methods (e.g., gradient descent). WI1W^{1}_{I}, WR1W^{1}_{R}, b1b^{1}, WI2W^{2}_{I}, WR2W^{2}_{R}, b2b^{2}, Wloc1W_{loc^{1}}, Wloc2W_{loc^{2}}, and the network weights in \mboxCNNr\mbox{{CNN}}_{r} and \mboxDecNNr\mbox{{DecNN}}_{r} are learnable parameters in RADCNN.

Implementation Details

For initial saliency detection, we use a CNN-DecNN independent from the CNN-DecNN used in the saliency refinement stage. The CNN part is initialized from the weights of VGG-CNN-S chatfield2014return, a relatively powerful CNN model pre-trained on ImageNet dataset. VGG-CNN-S consists of 5 convolutional layers and 3 fully-connected layers. We discard the fully-connected layers of VGG-CNN-S and retain only its convolutional and pooling layers for network initialization. The CNN accepts 224×224224\times 224 RGB images as inputs, and it outputs a 7×77\times 7 feature maps with 256 feature channels. The DecNN part of the initial CNN-DecNN is a network with 3 convolutional layers (5×55\times 5 kernel size, 1×11\times 1 stride, 2×22\times 2 zero paddings), and there is an unpooling layer before each convolutional layer. To increase the representational capability of the DecNN without adding too many weight parameters, we append a layer convolution layer with 1×11\times 1 convolution kernel, to each DecNN convolutional layer. At the end of the initial CNN-DecNN, the DecNN outputs a 56×5656\times 56 saliency map. The output size of 56×5656\times 56 achieves a good balance between computational complexity and saliency pixels details. For performance evaluation, the 56×5656\times 56 saliency map is resized to the input image’s original size. The initial CNN-DecNN is trained with Adam kingma2014adam in default learning settings.

As mentioned previously, the \mboxCNNr\mbox{{CNN}}_{r} and \mboxDecNNr\mbox{{DecNN}}_{r} used in RACDNN are trained and executed independently of those in the initial CNN-DecNN. On the other hand, \mboxDecNNr\mbox{{DecNN}}_{r} is initialized using the pre-trained weights of DecNN of the initial CNN-DecNN. In the recurrent layers of RACDNN, rectified linear unit (ReLU) is employed as the non-linear activation ϕ(⋅)\phi(\cdot). The feature maps of the hidden state hi1h^{1}_{i} (the first recurrent layer of RACDNN) is of size 7×77\times 7 and has 256256 feature channels. For the second recurrent layer’s hidden state hi2h^{2}_{i}, the feature representation is a 512512-dimensional vector. The weight parameters Wloc1W_{loc^{1}} and Wloc2W_{loc^{2}} of floc(⋅)f_{loc}(\cdot) are 512×256512\times 256 and 256×3256\times 3 matrices respectively. The number of recurrent iterations of RACDNN (inclusive of the 00-th iteration) is set to 99 for all saliency detection experiments. RACDNN is trained using RMSProp tieleman2012 with an initial learning rate of 0.00010.0001. The learning rate is reduced by an order of magnitude whenever validation performance stops improving. During training, gradients are hard-clipped to be within the range of $$ as a way to mitigate the gradient explosion problem which occurs when training recurrent-based networks. To speed up training and improve training convergence, we apply Batch Normalization ioffe2015batch to all weight layers (except for recurrent hidden-to-hidden connections) in both the initial CNN-DeCNN and RADCNN.

Most of the saliency detection methods employ object segmentation techniques which can output image segments with consistent saliency values within each segment. Furthermore, the edges of the output segments are sharp. To achieve similar effects, we apply a mean shift-based segmentation method frintrop2015traditional; garcia2015saliency to the outputs of RACDNN as a post-processing step.

Saliency Training Datasets

Learning-based methods require a big amount of training samples to generalize to new examples well. However, most of the saliency detection datasets are too small. It is not possible to train the deep models well if the experimental evaluations are done in such a way that each dataset is split into training, testing and validation sets in proportions. Here, we follow the dataset procedure in one recent deep learning-based saliency detection work zhao2015saliency. We train the deep models (initial CNN-DecNN and RADCNN) in our proposed method on saliency datasets different from the datasets used for experimental evaluations. The training datasets we use are: DUT-OMRON yang2013saliency, NJU2000 ju2015depth, RGBD Salient Object Detection dataset peng2014rgbd, and ImageNet segmentation dataset guillaumin2014imagenet. The data samples in these datasets reach a total number of 12,430, which is roughly the size of the dataset (with 10,000 samples) used in zhao2015saliency. We randomly split the combined datasets into 10,565 training samples and 1865 validation samples. Although the training set is considered large in saliency detection context, it is still small for deep learning methods, and may cause overfitting. Thus, we apply data augmentation in the form of cropping, translation, and color jittering on the training samples.

Experiments

We evaluate our proposed on a number of challenging saliency detection datasets: MSRA10K cheng2015global is by far the largest publicly available saliency detection dataset, containing 10,000 annonated saliency images. THUR15K cheng2014salientshape has 6,232 images which belong to five object classes of “butterfly”, “coffee mug”, “dog jump”, “giraffe”, and “plane”. It is challenging because some of its images do not contain any salient object. HKUIS li2015visual is a recently released saliency detection dataset with 4,447 annonated images. ECSSD shi2015hierarchical is a challenging saliency detection dataset with many semantically meaningful but structurally complex images. It contains 1,000 images. SED2 alpert2007image is a small saliency dataset having only 100 images. For each image, there are two salient objects.

Even though F-measure is the most commonly used evaluation metric for saliency detection, it is not comprehensive enough as it does not consider true negative saliency labeling. To have a more comprehensive experimental evaluation, we consider another evaluation metric known as Mean Absolute Error (MAE) adopted by borji2015salient. MAE is given by: 1W×H∑n=1W∑m=1H∣Sˉ(n,m)−Gˉ(n,m)∣\frac{1}{\mathcal{W}\times\mathcal{H}}\sum\limits_{\mathit{n}=1}^{\mathcal{W}}\sum\limits_{\mathit{m}=1}^{\mathcal{H}}|\bar{S}(n,m)-\bar{G}(n,m)|, where W\mathcal{W} and H\mathcal{H} are width and height of saliency map; Sˉ\bar{S} is the real-valued saliency map output normalized to the range of $,and, and\bar{G}$ is the saliency groundtruth. Saliency map binarization is not needed in MAE as it measures the mean of absolute differences between groundtruth saliency pixels and given saliency pixels.

2 Comparison with Baseline Methods

To highlight the advantages of recurrent attention mechanism in the proposed network RACDNN, we use CNN-DecNN as one of the baseline methods in our experiments. Compared to the proposed method, the baseline CNN-DecNN has no recurrent attention mechanism to perform iterative saliency refinement. The other baseline method is a CNN-DecNN paired with a non-recurrent attentional convolutional-deconvolutional network (NACDNN) in place of RACDNN. NACDNN is a RACDNN variant whose layers h1h^{1} and h2h^{2} are made non-recurrent. By removing the recurrent connections, NACDNN cannot learn context-aware features useful for saliency refinement despite having attention mechanism. At each computational iteration, NACDNN works almost like a CNN-DeCNN except that it has a localization network floc(⋅)f_{loc}(\cdot) that accepts CNN’s output as input and outputs spatial transformation matrix.

To compare the proposed method with baseline methods, we use F-measure and MAE as evaluation metrics. The F-measure scores and Mean Square Errors (MAEs) for comparisons with the baselines are shown in Table 1. On all of the five datasets and two evaluation metrics, the proposed method achieves better results than both the baseline methods. This shows that the RACDNN can help to improve the saliency map outputs of CNN-DecNN, using a recurrent attention mechanism to alleviate the scale issues of CNN-DecNN, and to learn region-based contextual dependencies not easily modeled by mere convolutional and deconvolutional network operations. The second baseline method NRACDNN that has attention mechanism performs better than the non-attentional first baseline. However, due to the lack of recurrent connections, NRACDNN is inferior to RACDNN because it does not exploit contextual information from past iterations for saliency refinement.

3 Comparison with State-of-the-art Methods

In addition to the baseline methods, we compare the proposed method “CNN-DecNN + RACDNN” with several state-of-the-art saliency detection methods: RRWR li2015robust, BSCA qin2015saliency, DRFI jiang2013salient, RBD zhu2014saliency, DSR li2014saliency, MC jiang2013saliency, and HS shi2015hierarchical. DRFI, RBD, DSR, MC, and HS are the top-performing methods evaluated in borji2015salient, while RRWR and BSCA are two very recent saliency detection works. To obtain the results for these methods, we run the original codes provided by the authors with recommended parameter settings. The precision-recall curves are given in Figure 5. We compute the curves based on the saliency maps generated by the proposed method. In overall, the proposed method “CNN-DecNN + RACDNN” performs better than the evaluated state-of-the-art methods. Especially in datasets with complex scenes (ECSSD & HKUIS), the performance gains of the proposed method over the state-of-the-art methods are more noticeable.

We also compare the proposed method “CNN-DecNN + RACDNN” with the state-of-the-art methods in terms of F-measure scores and Mean Square Errors (MAEs) (Table 2). In these evaluation metrics, its performance gains over the other methods are very significant. For the HKUIS and ECSSD dataset, the F-measure improvements of the proposed method over the next top-performing method DRFI are more than 5%. The proposed method also pushes down the MAEs on these challenging datasets by a large margin.

Besides quantitative results, we show some qualitative results in Figure 6. The proposed method “CNN-DecNN + RACDNN” can better detect multiple intermingled salient objects, as shown in the second image with a dog and a rabbit. Our method is the only one that can detect both objects well. The success of our method on this image is attributed to the attention mechanism that allows it to attend to different object regions for local refinement, making it is less likely to be negatively affected by distant noises and other objects. However, the proposed method tends to fail to detect salient objects which are mostly made up of background-like colors and textures (e.g., sky: third image, soil: fourth image).

To further evaluate the proposed method “CNN-DecNN + RACDNN”, we compare it with two recent deep learning-based saliency detection methods (MCDL zhao2015saliency and MDF li2015visual) on HKUIS, ECSSD, and SED2 datasets. We use the trained models provided by the authors. The F-measure scores and MAEs are given in Table 3, showing that the proposed method is comparable to both MCDL and MDF in terms of F-measure, but outperforming them in terms of MAEs.

Conclusion

In this paper, we introduce a novel method of using recurrent attention and convolutional-deconvolutional network to tackle the saliency detection problem. The proposed method has shown to be very effective experimentally. Still, the performance of proposed method may be limited by the quality of the initial saliency maps. To overcome such limitation, the recurrent attentional network can be potentially revamped to detect saliency from scratch in end-to-end manner. Also, this work can be readily adapted for other vision tasks that require pixel-wise prediction dosovitskiy2015learning; long2015fully.

Acknowledgement: The research is supported by Singapore Ministry of Education (MOE) Tier 2 ARC28/14, and Singapore A*STAR Science and Engineering Research Council PSF1321202099.

References