Joint Learning of Saliency Detection and Weakly Supervised Semantic Segmentation

Yu Zeng, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang

Introduction

Semantic image segmentation is an important and challenging task of computer vision, of which the goal is to predict a category label for every image pixel. Recently, convolutional neural networks (CNNs) have achieved remarkable success in semantic image segmentation . Due to the expensive cost for annotating semantic segmentation labels to train CNNs, weakly supervised learning has attracted increasing interest, resulting in various weakly supervised semantic segmentation (WSSS) methods. Saliency detection (SD) aims at identifying the most distinct objects or regions in an image, which has helped many computer vision tasks such as scene classification , image retrieval , visual tracking , to name a few. With the success of deep CNNs, it has been made a lot of attempts to use deep CNNs or deep features for saliency detection .

The two tasks both require to generate accurate pixel-wise masks. Hence, they have close connections. On the one hand, given the saliency map of an image, the computation load of a segmentation model can be reduced because of avoiding processing background. On the other hand, given the segmentation result of an image, the saliency map can be readily derived by selecting the salient category. Therefore, many existing WSSS methods have greatly benefited from SD. It is a widespread practice to exploit class activation maps (CAMs) for locating the objects of each category and use SD methods for selecting background regions. For example, Wei et al. use CAMs of a classification network with different dilated convolutional rates to find object regions, and use saliency maps of to find background regions for training a segmentation model. Wang et al. use saliency maps of to refine object regions produced by classification networks.

However, those WSSS methods simply utilize the results of pre-trained saliency detection models, which is not the most efficient configuration. On the one hand, they use SD methods as a pre-processing step to generate annotations for training their segmentation models while ignoring the interactions between SD and WSSS, which blocks the WSSS models from fully exploiting the segmentation cues of the strong saliency annotations. On the other hand, heuristic rules are usually required for selecting background regions according to the results of SD models, thereby complicating the training process and leading to a not end-to-end manner.

In this paper, we propose a unified, end-to-end training framework to solve both SD and WSSS tasks jointly. Unlike most existing WSSS methods that used pre-trained saliency detection models, we directly take advantage of pixel-level saliency labels. The core motive is to utilize semantic information of the image-level category labels and the segmentation cues of the category-agnostic saliency labels. The image-level category labels can make a CNN recognize the semantic categories, but they do not contain any spatial information, which is essential for segmentation. Although it has been suggested that CNNs trained with image-level labels are also informative of object locations, only a coarse spatial distribution can be inferred, as shown in the first row of Figure 1. We solve this problem with the pixel-level saliency labels. Through explicitly modelling the connection between SD and WSSS, we derive the saliency maps from the segmentation results and minimize the loss between them and the saliency ground-truth. So that the CNN has to precisely cut the recognized objects so as to make the derived saliency maps match the ground-truth.

Specifically, we propose a saliency and segmentation network (SSNet), which includes a segmentation network (SN) and a saliency aggregation module (SAM). For an input image, SN generates the segmentation results, as shown in The second column of Figure 1. SAM predicts the saliency score of each category and then aggregates the segmentation masks of all categories into a saliency map according to their saliency scores, which bridges the gap between semantic segmentation and saliency detection. As shown in the third column of Figure 1, given the segmentation map and saliency score of each category, saliency detection result can be generated by highlighting the masks of salient objects (e.g., the mask of persons in the third column) and suppressing the masks of the objects of low salience (e.g., the mask of bottles in the third column). When training, the loss is computed between the segmentation results and the image-level category labels as well as the saliency maps and the saliency ground-truth.

Our approach has several advantages. First, compared with existing WSSS methods that exploit pre-trained SD models for pre-processing, our method explicitly models the relationships between saliency and segmentation, which can transfer the learned segmentation knowledge from class-agnostic image-specific saliency categories with pixel-level annotations to unseen semantic categories with only image-level annotations. Second, as a low-level vision task, annotating pixel-level ground truth for saliency detection is less expensive than semantic segmentation. Therefore, compared with fully supervised segmentation methods, our method is trained with image-level category labels and saliency annotations, requiring less labeling cost. Third, compared with existing segmentation or saliency methods, our method can simultaneously predict the segmentation results and saliency results using a single model, with most parameters shared between the two tasks.

In summary, our main contributions are three folds:

We propose a unified end-to-end framework for both SD and WSSS tasks, in which segmentation is split into two learning tasks respectively based on image-level category labels and pixel-level saliency annotations.

We design a saliency aggregation module to explicitly bridge the two tasks, through which WSSS can directly benefit from saliency inference and vice versa.

The experiments on the PASCAL VOC 2012 segmentation benchmark and four saliency benchmarks demonstrate the effectiveness of the proposed method. It achieves favorable performance against weakly supervised semantic segmentation methods and fully supervised saliency detection methods. We make our code and models available for further researcheshttps://github.com/zengxianyu/jswshttp://ice.dlut.edu.cn/lu/.

Related work

Earlier saliency detection methods used low-level features and heuristic priors to detect salient objects, which were not robust to complex scenes. Recently, deep learning based methods have achieved remarkable performance improvements. Incipient deep learning based methods usually used regions as computation units, such as superpixels, image patches, and region proposals. Wang et al. trained two neural networks that estimate saliency of image patches and regional proposals respectively. Li and Yu used CNNs to extract multi-scale features and predict the saliency of each superpixel. Inspired by the success of fully convolutional network (FCN) on semantic segmentation, some methods have been proposed to exploit fully convolutional structure for pixel-wise saliency prediction. Liu and Han proposed a deep hierarchical network to learn a coarse global saliency map and then progressively refine it. Wang et al. proposed a recurrent FCN incorporates saliency priors. Zhang et al. propose to make CNNs learn deep uncertain convolutional features (UCF) to encourage the robustness and accuracy of saliency detection. Zhang et al. proposed an attention guided network which selectively integrates multi-level contextual information in a progressive manner. Chen et al. proposed reverse attention to guide residual feature learning in a top-down manner for saliency detection. All of the above saliency detection methods trained fully supervised models for a single task. Although our method slightly increases labeling cost, it achieves state-of-the-art performance in both saliency detection and semantic segmentation.

2 Segmentation with weak supervision

In recent years, a lot of weakly supervised semantic segmentation methods have been proposed to alleviate the cost of labeling. Various supervision has been exploited, such as the image-level labels, bounding boxes, scribbles, etc. Among all kinds of weak supervision, the weakest one, i.e., image-level supervision, has attracted the most attention. In image-level weakly supervised segmentation, some methods exploited results of the pre-trained saliency detection models. A simple-to-complex method was presented in , in which an initial segmentation model is trained with simple images using saliency maps for supervision. Then the ability of the segmentation model is enhanced by progressively including samples of increasing complexity. Wei et al. iteratively used CAM to discover object regions and used saliency detection results of to find background regions to train the segmentation model. Oh et al. used an image classifier to find the high confidence points over the objects classes, i.e. object seeds, and exploit a CNN-based saliency detection model to find the masks corresponding to some of the detected object seeds. Then these class-specific masks were used to train a segmentation model. Wei et al. used a classification network with convolutional blocks of different dilated rates to find object regions and used saliency detection results of to find background regions to train a segmentation model. Wang et al. started from the object regions produced by classification networks. The object regions were expanded using the mined features and refined using saliency maps produced by . Then the refined object regions were used as supervision to train a segmentation network. The above weakly supervised segmentation methods all exploited results of pre-trained saliency detection models, either using the existing models or separately training their saliency models and segmentation models. The proposed method has two main differences from these methods. First, these methods used pre-trained saliency detection models, while we directly exploit strong saliency annotations and work in an end-to-end manner. Second, in these methods, saliency detection was used as a pre-processing step to generate training data for segmentation. In contrast, we simultaneously solve saliency detection and semantic segmentation using a single model, of which most parameters are shared between the two tasks.

3 Multi-task learning

Multi-task learning has been used in a wide range of computer vision problems. Teichman et al. proposed an approach to joint classification, detection, and segmentation using a unified architecture where the encoder is shared among the three tasks. Kokkinos proposed an UberNet that jointly handles low-, mid-, high-level tasks including boundary detection, normal estimation, saliency estimation, semantic segmentation, human part segmentation, semantic boundary detection, region proposal generation, and object detection. Eigen and Fergus used a multiscale CNN to address three different computer vision tasks: depth prediction, surface normal estimation, and semantic labeling. Xu et al. proposed a PAD-Net that first solves several auxiliary tasks ranging from low level to high level, and then used the predictions as multi-modal input for the final task. The models above all worked in full supervision setting. In contrast, we jointly learn to solve a task in weak supervision setting and another task in full supervision setting.

The proposed approach

In this section, we detail the joint learning framework for simultaneous saliency detection and semantic segmentation. We first give an overview of the proposed saliency and segmentation network (SSNet). Then we describe the details of the segmentation network (SN) and the saliency aggregation module (SAM) in Section 3.2 and 3.3. Finally, we present the joint learning strategy in Section 3.4. Figure 2 illustrates the overall architecture of the proposed method.

We design two variants of SSNet, i.e. SSNet-1, and SSNet-2, for two training stages, respectively. In the first training stage, the SSNet-1 is trained with pixel-level saliency annotations and image-level semantic category labels. In the second stage, the SSNet-2 is trained with saliency annotations and image-level semantic category labels as well as semantic segmentation results predicted by SSNet-1. Both the SSNet-1 and SSNet-2 consist of a segmentation network (SN) and a saliency aggregation module (SAM). Given an input image, SN predicts a segmentation result. SAM predicts a saliency score for each semantic class and aggregates the segmentation map into a single channel saliency map according to the saliency score of each class. Both the SSNet-1 and SSNet-2 are trained end-to-end.

2 Segmentation networks

The segmentation network consists of a feature extractor to extract features from the input image and several convolution layers to predict segmentation results given the features. Feature extractors of our networks are designed based on state-of-the-art CNN architectures for image recognition, e.g., VGG and DenseNet , which typically contain five convolutional blocks for feature extraction and a fully connected classifier. We remove the fully connected classifier and use the convolutional blocks as our feature extractor. To obtain larger feature maps, we remove the downsampling operator from the last two convolution blocks and use dilated convolution to retain the original receptive field. The feature extractor generates feature maps of 1/81/8 the input image size. We resize the input images to 256×256256\times 256, so the resulted feature maps are 32×3232\times 32 in spatial scale.

In the first training stage, the only available semantic supervision cue is the image-level labels. Trained with image-level labels, a coarse spatial distribution of each class can be inferred but it is difficult to train a sophisticated model. Therefore, we use a relatively simple structure in SSNet-1 for generating segmentation results, i.e. a 1×11\times 1 convolution layer. The predicted CC-channel segmentation map and one-channel background map are of 1/81/8 the input image size, in which CC is the number of semantic classes. Each element of the segmentation map and background map is a value in $.Thevaluesofallclassessumto1foreachpixel.Thenthesegmentationresultsareupsampledbyadeconvolutionlayertotheinputimagesize.Inthesecondtrainingstage,thesegmentationresultsofSSNet−1canbeusedfortraining,whichisastrongersupervisioncue.Therefore,wecanuseamorecomplexsegmentationnetworktogeneratefinersegmentationresults.InspiredbyDeeplab,weusefour. The values of all classes sum to 1 for each pixel. Then the segmentation results are upsampled by a deconvolution layer to the input image size. In the second training stage, the segmentation results of SSNet-1 can be used for training, which is a stronger supervision cue. Therefore, we can use a more complex segmentation network to generate finer segmentation results. Inspired by Deeplab , we use four3\times 3convolutionlayerswithdilationrateconvolution layers with dilation rate6,12,18,24inSSNet−2andtakethesummationoftheiroutputsasthesegmentationresults.SimilartoSSNet−1,thesesegmentationresultsareofin SSNet-2 and take the summation of their outputs as the segmentation results. Similar to SSNet-1, these segmentation results are of1/8$ the input image size, and are upsampled by a deconvolution layer to the input size.

3 Saliency aggregation

We design a saliency aggregation module (SAM) as a bridge between the two tasks so that the segmentation network can make use of the class-agnostic pixel-level saliency labels and generate more accurate segmentation results. This module takes the 32×3232\times 32 outputs FF of the feature extractor, and generates a CC-dimensional vector v\bm{v} with a 32×3232\times 32 convolution layer and a sigmoid function, of which each element viv_{i} is the saliency score of the ii-th category. Then the saliency map SS is given by a weighted sum of the segmentation masks of all classes:

where HiH_{i} denotes the ii-th channel of the segmentation results encoding the spatial distribution of the ii-th category, which are the output of the segmentation network.

4 Jointly learning of saliency and segmentation

We use two training sets to train the proposed SSNet: the saliency dataset with pixel-level saliency annotations, and the classification dataset with image-level semantic category labels. Let Ds={(Xn,Yn)}n=1Ns\mathcal{D}_{s}=\{(X^{n},Y^{n})\}_{n=1}^{N_{s}} denote the saliency dataset, in which XnX^{n} is the image, and YnY^{n} is the ground truth. Each element of YnY^{n} is either 1 or 0, representing the corresponding pixel belongs to salient objects or background, respectively. The classification dataset is denoted as Dc={(Xn,tn)}n=1Nc\mathcal{D}_{c}=\{(X^{n},\bm{t}^{n})\}_{n=1}^{N_{c}}, in which XnX^{n} is the image, and tn\bm{t}^{n} is the one-hot encoding of the categories of the image.

For an input image, the segmentation network generates its segmentation result, from which the probability of each category can be derived by averaging the segmentation results over spatial locations. We compute the loss between these values and the ground-truth category labels and backward propagate it to make the segmentation results semantically correct, i.e., the semantic categories appearing in the input image are correctly recognized. This loss, denoted as Lc\mathcal{L}_{c}, is defined as follow,

in which tint_{i}^{n} is the ii-th element of tn\bm{t}^{n}. tin=1t_{i}^{n}=1 represents the image XnX^{n} contains objects of the ii-th category, and tin=0t_{i}^{n}=0 otherwise. t^n\bm{\hat{t}}^{n} is the average over spatial positions of the segmentation maps HnH^{n} of image XnX^{n}, of which each element t^in∈\hat{t}^{n}_{i}\in represents the predicted probability of the ii-th class objects presenting in the image.

The image-level category labels can make the segmentation network recognize the semantic categories, but they do not contain any spatial information, which is essential for segmentation. We solve this problem with the pixel-level saliency labels. As stated in Section 3.3, the SAM generates the saliency score of each category and aggregates the segmentation result into a saliency map. We minimize a loss Ls1\mathcal{L}_{s1} between the derived saliency maps and the ground-truth so that the segmentation network has to precisely cut the recognized objects to make the derived saliency maps match the ground-truth. The loss Ls1\mathcal{L}_{s1} between the saliency maps and the saliency ground-truth is defined as follow,

where ymn∈{0,1}y_{m}^{n}\in\{0,1\} is the value of the mm-th pixel of the saliency ground truth YnY^{n}. smn∈s_{m}^{n}\in is the value of the mm-th pixel in the saliency map of the image XnX^{n}, encoding the predicted probability of the mm-th pixel being salient.

In the first training stage, we train SSNet-1 with the loss Lc+Ls1\mathcal{L}_{c}+\mathcal{L}_{s1}. After having trained SSNet-1, we run it on the classification dataset Dc\mathcal{D}_{c} and obtain the C+1C+1-channel segmentation results, of which the first CC channels correspond to the CC semantic categories and the last channel corresponds to the background. Then the first CC channels of the segmentation result are cross-channel multiplied with the one-hot class label tn\bm{t}^{n} to suppress wrong predictions and refined with CRF to enhance spatial smoothness. Finally, we obtain some pseudo labels by assigning each pixel mm of each training image Xn∈DcX^{n}\in\mathcal{D}_{c} a class label including the background label corresponding to the maximum value in the refined segmentation result. We define a loss Ls2\mathcal{L}_{s2} between the segmentation results and the pseudo labels as follow,

in which himn∈,i=1,...,Ch^{n}_{im}\in,i=1,...,C is the value of HnH^{n} at pixel mm and channel ii, representing the probability of pixel mm belonging to the ii-th class. SSNet-2 is trained with the loss Ls1+Ls2\mathcal{L}_{s1}+\mathcal{L}_{s2}.

Experiments

Segmentation For semantic segmentation task, we evaluate the proposed method on the PASCAL VOC 2012 segmentation benchmark . This dataset has 20 object categories and one background category. It is split into a training set of 1,464 images, a validation set of 1,449 images and a test set of 1,456 images. Following the common practice , we increase the number of training images to 10,582 by augmentation. We only use image-level labels for training. The performance of our method and other state-of-the-art methods are evaluated on the validation set and test set. The performance for semantic segmentation is evaluated in terms of inter-section-over-union averaged over 21 classes (mIOU) according to the PASCAL VOC evaluation criterion. We obtain the mIOU on the test set by submitting our results to the PASCAL VOC evaluation server.

Saliency For saliency detection task, we use the DUT-S training set for training, which has 10,553 images with pixel-level saliency annotations. The proposed method and other state-of-the-art methods are evaluated on four benchmark datasets: ECSSD , PASCAL-S , HKU-IS , SOD . ECSSD contains 1000 natural images with multiple objects of different sizes. PASCAL-S stems from the validation set of PASCAL VOC 2010 segmentation dataset and contains 850 natural images. HKU-IS has 4447 images chosen to include multiple disconnected objects or objects touching the image boundary. SOD has 300 challenging images, of which many images contain multiple objects either with low contrast or touching the image boundary. The performance for saliency detection is evaluated in terms of maximum F-measure and mean absolute error (MAE).

Training/Testing Settings We adopt DenseNet-169 pre-trained on ImageNet as the feature extractor of our segmentation network due to its ability to achieve comparable performance with a smaller number of parameters than other architectures. Our network is implemented based on Pytorch framework and trained on two NVIDIA GeForce GTX 1080 Ti GPU. We use Adam optimizer to train our network. We randomly crop a patch of 9/10 of the original image size and rescaled to 256×256256\times 256 when training. The batch size is set to 16. We train both the SSNet-1 and SSNet-2 for 10,000 iterations with initial learning rate 1e-4, and decrease the learning rate by 0.5 every 1000 iterations. When testing, the input image is resized to 256×256256\times 256. Then, the predicted segmentation results and saliency maps are resized to the input size by nearest interpolation. We do not use any post-processing on the segmentation results. We apply CRF to refine the saliency maps.

2 Comparison with saliency methods

We compare our method with the following state-of-the-art deep learning based fully supervised saliency detection methods: PAGR (CVPR’18) , RAS (ECCV’18) , UCF (ICCV’17) , Amulet (ICCV’17) , RFCN (ECCV’16) , DS (TIP’16) , ELD (CVPR’16) , DCL (CVPR’16) , DHS (CVPR’16) , MCDL (CVPR’15) , MDF (CVPR’15) . Figure 3 shows a visual comparison of our method against state-of-the-art fully supervised saliency detection methods. The comparison in terms of MAE and maximum F-measure is shown in Table 1 and Table 2 respectively. As shown in Table 1, the proposed method achieves the smallest MAE among across all datasets. Maximum F-measure in Table 2 also shows that our method achieves the second largest F-measure in one dataset, and achieves the third largest F-measure in the other three datasets. Together the two metrics, it can be seen that our method achieves state-of-the-art performance in saliency detection task.

3 Comparison with segmentation methods

In this section, we compare our method with previous state-of-the-art weakly supervised semantic segmentation methods, i.e. MIL (CVPR’15) , WSSL , RAWK , BFBP (ECCV’16) , SEC (ECCV’16) , AE (CVPR’17) , STC (PAMI’17) , CBTS (CVPR’17) , ESOS (CVPR’17) , MCOF (CVPR’18) , MDC (CVPR’18) . WSSL uses bounding boxes as supervision, RAWK uses scribbles as supervision, and other methods use image-level categories as supervision. Among the methods using image-level supervision, ESOS exploits the saliency detection results of a deep CNN trained with bounding box annotations. AE, STC, MCOF, MDC use the results of fully supervised saliency detection models and thus implicitly use pixel level saliency annotations. As some previous methods used VGG16 as its backbone network, we also report the performance of our method using VGG16. It can be seen from Table 3 and Table 4 that our method compares favourably against all the above methods, including methods using stronger supervision such as bounding boxes (WSSL) and scribbles (RAWK). Our method also outperforms the methods i.e., ESOS, AE, STC, MCOF, MDC, that implicitly use saliency annotations by using pre-trained saliency detection models. Compared with these methods, our method simultaneously solves semantic segmentation and saliency detection and can be trained in an end-to-end manner, which is more efficient and easier to train.

4 Ablation study

In this section, we analyze the effect of the proposed jointly learning framework. To validate the impact of multitasking, we show the performance of the networks trained in different single-task and multi-task settings.

Semantic segmentation The quantitative and the qualitative comparison of the models trained in the different settings for segmentation task is shown in Table 5 and Figure 5 respectively. For the first training stage, we firstly train SSNet-1 in the single-task setting, where only the image-level category labels and Lc\mathcal{L}_{c} are used. The resulted model is denoted as SSNet-S, of which the mIOU is shown in the first column of Table 5. Then we add the saliency task to train SSNet-1 in the multi-task setting. In this setting Lc+Ls1\mathcal{L}_{c}+\mathcal{L}_{s1} is used as loss function, with both the image-level category labels and the saliency dataset are used as training data. The resulted model is denoted as SSNet-M, of which mIOU is shown in the second column of Table 5. It can be seen that SSNet-M has a much larger mIOU than SSNet-S, demonstrating that jointly learning saliency detection is of great benefit to WSSS. In the second training stage, the training data for SSNet-2 consists of two splits: the one is the predictions of SSNet-1, and the other is the saliency dataset. In order to verify the contribution of each split, we train SSNet-2 in three settings: 1) train with only the predictions of SSNet-S using the Ls2\mathcal{L}_{s2} as loss function, 2) train with only the predictions of SSNet-M using the Ls2\mathcal{L}_{s2} as loss function, and 3) train with the predictions of SSNet-M and the saliency dataset using Ls1+Ls2\mathcal{L}_{s1}+\mathcal{L}_{s2} as loss function. The resulted models under the three settings is denoted as SSNet-SS, SSNet-MS and SSNet-MM, of which the mIOU scores are shown in the third to the fifth column of Table 5. From the comparison of SSNet-SS and SSNet-MS, it can be seen that the model trained with the multi-task setting in the first training stage can provide better training data for the second training stage. The comparison of SSNet-MS and SSNet-MM shows that when trained with the same pixel-level segmentation labels, the model trained in the multi-task setting is still better than the single-task setting.

Saliency detection To study the effect of jointly learning for saliency detection, we compare the performance of SSNet-2 trained in multi-task settings and single-task settings. We firstly train SSNet-2 only for saliency detection task, resulting in a model denoted as SSNet-2S, of which the maximum F-measure and MAE are shown in the first column of Table 6. Then we run the model SSNet-MM mentioned above on saliency dataset, and the resulted F-measure and MAE are shown in the second column of Table 6. As can be seen, the models trained in multi-task setting has a comparable performance to the model trained in single-task setting, where the former has a larger maximum F-measure and the later is better in terms of MAE. Therefore, it is safe to conclude that conduct jointly learning of semantic segmentation dose not harm the performance on saliency detection. This result validates the superiority of the proposed jointly learning framework considering its great benefit to semantic segmentation.

Conclusion

This paper presents a joint learning framework for saliency detection (SD) and weakly supervised semantic segmentation (WSSS) using a single model, i.e. the saliency and segmentation network (SSNet). Compared with WSSS methods exploiting pre-trained SD models, our method makes full use of segmentation cues from saliency annotations and is easier to train. Compared with existing fully supervised SD methods, our method can provide more informative results. Experiments shows that our method achieves state-of-the-art performance among both fully supervised SD methods and WSSS methods.

Acknowledgements

Supported by the National Natural Science Foundation of China #61725202, #61829102, #61751212, #61876202, Fundamental Research Funds for the Central Universities under Grant #DUT19GJ201 and Dalian Science and Technology Innovation Foundation #2019J12GX039.

References