Self-Erasing Network for Integral Object Attention

Qibin Hou, Peng-Tao Jiang, Yunchao Wei, Ming-Ming Cheng

Introduction

Semantic segmentation aims at assigning each pixel a label from a predefined label set given a scene. For fully-supervised semantic segmentation , the requirement of large-scale pixel-level annotations considerably limits its generality . Some weakly-supervised works attempt to leverage relatively weak supervisions, such as scribbles , bounding boxes , or points , but they still need large amount of hand labors. Therefore, semantic segmentation with image-level supervision is becoming a promising way to relief lots of human labors. In this paper, we are also interested in the problem of weakly-supervised semantic segmentation. As only image-level labels are available, most recent approaches , more or less, rely on different attention models due to their ability of covering small but discriminative semantic regions. Therefore, how to generate high-quality attention maps is essential for offering reliable initial heuristic cues for training segmentation networks. Earlier weakly-supervised semantic segmentation methods mostly adopt the original Class Activation Maps (CAM) model for object localization. For small objects, CAM does work well but when encountering large objects of large scales it can only localize small areas of discriminative regions, which is harmful for training segmentation networks in that the undetected semantic regions will be judged to background.

Interestingly, the adversarial erasing strategy (Fig. 1) has been proposed recently. Benefiting from the powerful localization ability of CNNs, this type of methods is able to further discover more object-related regions by erasing the detected regions. However, a key problem of this type of methods is that as more semantic regions are mined, the attentions may spread to the background and further the localization ability of the initial attention generator is downgraded. For example, trains often run on rails and hence as trains are erased rails may be classified as the train category, leading to negative influence on learning semantic segmentation networks.

In this paper, we propose a promising way to overcome the above mentioned drawback of the adversarial erasing strategy by introducing the concept of self-erasing. The background regions of common scenes often share some similarities, which motivates us to explicitly feed attention networks with a roughly accurate background prior to confine the observable regions in semantic fields. To do so, we present two self-erasing strategies by leveraging the background prior to purposefully suppress the spread of attentions to the background regions. Moreover, we design a new attention network that takes the above self-erasing strategies into account to discover more high-quality attentions from a potential zone instead of the whole image . We apply our attention maps to weakly-supervised semantic segmentation, evaluate the segmentation results on the PASCAL VOC 2012 benchmark, and show substantial improvements compared to existing methods.

Related Work

To date, a great number of attention networks have been developed, attempting to reveal the working mechanism of CNNs. At earlier stage, error back-propagation based methods were proposed for visualizing CNNs. CAM adopted a global average pooling layer followed by a fully connected layer as a classifier. Later, Selvaraju proposed the Grad-CAM, which can be embedded into a variety of off-the-shelf available networks for visualizing multiple tasks, such as image captioning and image classification. Zhang et al. , motivated by humans’ visual system, used the winner-take-all strategy to back-propagate discriminative signals in a top-down manner. A similar property shared by the above methods is that they only attempt to produce an attention map.

In , Wei et al. proposed the adversarial erasing strategy, which aims at discovering more unseen semantic objects. A CAM branch is used to determine an initial attention map and then a threshold is used to selectively erase the discovered regions from the images. The erased images are then sent into another CNN to further mine more discriminative regions. In , Zhang et al. and Li et al. extended the initial adversarial erasing strategy in an end-to-end training manner.

2 Weakly-Supervised Semantic Segmentation

Due to the fact that collecting pixel-level annotations is very expensive, more and more works are recently focusing on weakly-supervised semantic segmentation. Besides some works relying on relatively strong supervisions, such as scribble , points , and bounding boxes , most weakly-supervised methods are based on only image-level labels or even inaccurate keyword . Limited by keyword-level supervision, many works harnessed attention models for generating the initial seeds. Saliency cues are also adopted by some methods as the initial heuristic cues. Beyond that, there are also some works proposing different strategies to solve this problem, such as multiple instance learning and the EM algorithm .

Self-Erasing Network

In this section, we describe details of the proposed Self-Erasing Network (SeeNet). An overview of our SeeNet can be found in Fig. 3. Before the formal description, we first introduce the intuition of our proposed approach.

As stated in Sec. 1, with the increase of training iterations, adversarial erasing strategy tends to mine more areas not belonging to any semantic objects at all. Thus, it is difficult to determine when the training phase should be ended. An illustration of this phenomenon has been depicted in Fig. 1. In fact, we humans always ‘deliberately’ suppress the areas that we are not interested in so as to better focus on our attentions . When looking at a large object, we often seek the most distinctive parts of the object first and then move the eyes to the other parts. In this process, humans are always able to inadvertently and successfully neglect the distractions brought by the background. However, attention networks themselves do not possess such capability with only image-level labels given. Therefore, how to explicitly introduce background prior to attention networks is essential. Inspired by this cognitive process of humans, other than simply erasing the attention regions with higher confidence as done in existing works , we propose to explicitly tell CNNs where the background is so as to let attention networks better focus on discovering real semantic objects.

2 The Idea of Self-Erasing

To highlight the semantic regions and keep the detected attention areas from expanding to background areas, we propose the idea of self-erasing during training. Given an initial attention map (produced by SAS_{A} in Fig. 3), we functionally separate the images into three zones in spatial dimension, the internal “attention zone”, the external “background zone”, and the middle “potential zone” (Fig. 2c). By introducing the background prior, we aim to drive attention networks into a self-erasing state so that the observable regions can be restricted to non-background areas, avoiding the continuous spread of attention areas that are already near a state of perfection. To achieve this goal, we need to solve the following two problems: (I) Given only image-level labels, how to define and obtain the background zone. (II) How to introduce the self-erasing thought into attention networks.

Regarding the circumstance of weak supervision, it is quite difficult to obtain a precise background zone, so we have to seek what is less attractive than the above unreachable objective to obtain relatively accurate background priors. Given the initial attention map MAM_{A}, other than thresholding MAM_{A} with δ\delta for a binary mask BAB_{A} as in , we also consider using another constant which is less than δ\delta to get a ternary mask TAT_{A}. For notational convenience, we here use δh\delta_{h} and δl\delta_{l} (δh>δl\delta_{h}>\delta_{l}) to denote the two thresholds. Regions with values less than δl\delta_{l} in MAM_{A} will all be treated as the background zone. Thus, we define our ternary mask TAT_{A} as follows: TA,(i,j)=0T_{A,(i,j)}=0 if MA,(i,j)≥δhM_{A,(i,j)}\geq\delta_{h}, TA,(i,j)=−1T_{A,(i,j)}=-1 if MA,(i,j)<δlM_{A,(i,j)}<\delta_{l}, and TA,(i,j)=1T_{A,(i,j)}=1 otherwise. This means the background zone is associated with a value of -1 in TAT_{A}. We empirically found that the resulting background zone covers most of the real background areas for almost all input scenes. This is reasonable as SAS_{A} is already able to locate parts of the semantic objects.

With the background priors, we introduce the self-erasing strategies by reversing the signs of the feature maps corresponding to the background outputted by the backbone to make the potential zone stand out. To achieve so, we extend the ReLU layer to a more general case. Recall that the ReLU function, according to its definition, can be expressed as ReLU(x)=max⁡(0,x).\text{ReLU}(x)=\max(0,x). More generally, our C-ReLU function takes a binary mask into account and is defined as

where BB is a binary mask, taking values from {−1,1}\{-1,1\}. Unlike ReLUs outputting tensors with only non-negative values, our C-ReLUs conditionally flip the signs of some units according to a given mask. We expect that the attention networks can focus more on the regions with positive activations after C-ReLU and further discover more semantic objects from the potential zone because of the contrast between the potential zone and the background zone.

3 Self-Erasing Network

Our architecture is composed of three branches after a shared backbone, denoted by SA,SB,S_{A},S_{B}, and SCS_{C}, respectively. Fig. 3 illustrates the overview of our proposed approach. Similarly to , our SAS_{A} has a similar structure to , the goal of which is to determine an initial attention. SBS_{B} and SCS_{C} have similar structures to SAS_{A} but differently, the C-ReLU layer is inserted before each of them.

Self-erasing strategy I. By adding the second branch SBS_{B}, we introduce the first self-erasing strategy. Given the attention map MAM_{A} produced by SAS_{A}, we can obtain a ternary mask TAT_{A} according to Sec. 3.2. When sending TAT_{A} to the C-ReLU layer of SBS_{B}, we can easily adjust TAT_{A} to a binary mask by setting non-negative values to 1. When taking the erasing strategy into account, we can extend the binary mask in C-ReLU function to a ternary case. Thus, Eqn. (1) can be rewritten as

An visual illustration of Eqn. (2) has been depicted in Fig. 2c. The zone highlighted in yellow corresponds to attentions detected by SAS_{A}, which will be erased in the output of the backbone. Units with positive values in the background zone will be reversed to highlight the potential zone. During training, SBS_{B} will fall in a state of self-erasing, deterring the background stuffs from being discovered and meanwhile ensuring the potential zone to be distinctive.

Self-erasing strategy II. This strategy aims at further avoiding attentions appearing in the background zone by introducing another branch SCS_{C}. Specifically, we first transform TAT_{A} to a binary mask by setting regions corresponding to the background zone to 1 and the rest regions to 0. In this way, only the background zone of the output of the C-ReLU layer has non-zero activations. During the training phase, we let the probability of the background zone belonging to any semantic classes learn to be 0. Because of the background similarities among different images, this branch will help correct the wrongly predicted attentions in the background zone and indirectly avoid the wrong spread of attentions.

The overall loss function of our approach can be written as: L=LSA+LSB+LSC\mathcal{L}=\mathcal{L}_{S_{A}}+\mathcal{L}_{S_{B}}+\mathcal{L}_{S_{C}}. For all branches, we treat the multi-label classification problem as MM independent binary classification problems by using the cross-entropy loss, where MM is the number of semantic categories. Therefore, given an image II and its semantic labels y\mathbf{y}, the label vector for SAS_{A} and SBS_{B} is ln=1\mathbf{l}_{n}=1 if n∈yn\in\mathbf{y} and 0 otherwise, where ∣l∣|\mathbf{l}| = MM. The label vector of SCS_{C} is a zero vector, meaning that no semantic objects exist in the background zone.

To obtain the final attention maps, during the test phase, we discard the SCS_{C} branch. Let MBM_{B} be the attention map produced by SBS_{B}. We first normalize both MAM_{A} and MBM_{B} to the range $anddenotetheresultsasand denote the results as\hat{M}_{A}andand\hat{M}_{B}.Then,thefusedattentionmap. Then, the fused attention mapM_{F}iscalculatedbyis calculated byM_{F,i}=\max(\hat{M}_{M,i},\hat{M}_{B,i}).Toobtainthefinalattentionmap,duringthetestphase,wealsohorizontallyfliptheinputimagesandgetanotherfusedattentionmap. To obtain the final attention map, during the test phase, we also horizontally flip the input images and get another fused attention mapM_{H}.Therefore,ourfinalattentionmap. Therefore, our final attention mapM_{final}canbecomputedbycan be computed byM_{final,i}=\max(M_{F,i},M_{H,i})$.

Weakly-Supervised Semantic Segmentation

To test the quality of our proposed attention network, we applied the generated attention maps to the recently popular weakly-supervised semantic segmentation task. To compare with existing state-of-the-art approaches, we follow a recent work , which leverages both saliency maps and attention maps. Instead of applying an erasing strategy to mine more salient regions, we simply use a popular salient object detection model to extract the background prior by setting a hard threshold as in . Specifically, given an input image II, we first simply normalize its saliency map obtaining DD taking values from $.Let. Let\mathbf{y}betheimage−levellabelsetofbe the image-level label set ofItakingvaluesfromtaking values from\{1,2,\dots,M\},where, whereMisthenumberofsemanticclasses,andis the number of semantic classes, andA_{c}beoneofattentionmapsassociatedwithlabelbe one of attention maps associated with labelc\in\mathbf{y}.Wecancalculateour“proxyground−truth”accordingtoAlgorithm1.Following,hereweharnessthefollowingharmonicmeanfunctiontocomputetheprobabilityofpixel. We can calculate our “proxy ground-truth” according to Algorithm 1. Following , here we harness the following harmonic mean function to compute the probability of pixelI_{i}belongingtoclassbelonging to classc$:

Parameter ww here is used to control the importance of attention maps. In our experiments, we set ww to 1.

Experiments

To verify the effectiveness of our proposed self-erasing strategies, we apply our attention network to the weakly-supervised semantic segmentation task as an example application. We show that by embedding our attention results into a simple approach, our semantic segmentation results outperform the existing state-of-the-arts.

Datasets and evaluation metrics. We evaluate our approach on the PASCAL VOC 2012 image segmentation benchmark , which contains 20 semantic classes plus the background category. As done in most previous works, we train our model for both the attention and segmentation tasks on the training set, which consists of 10,582 images, including the augmented training set provided by . We compare our results with other works on both the validation and test sets, which have 1,449 and 1,456 images, respectively. Similarly to previous works, we use the mean intersection-over-union (mIoU) as our evaluation metric.

Network settings. For our attention network, we use VGGNet as our base model as done in . We discard the last three fully-connected layers and connect three convolutional layers with 512 channels and kernel size 3 to the backbone as in . Then, a 20 channel convolutional layer, followed by a global average pooling layer is used to predict the probability of each category as done in . We set the batch size to 16, weight decay 0.0002, and learning rate 0.001, divided by 10 after 15,000 iterations. We run our network for totally 25,000 iterations. For data augmentation, we follow the strategy used in . Thresholds δh\delta_{h} and δl\delta_{l} in SBS_{B} are set to 0.7 and 0.05 times of the maximum value of the attention map inputted to C-ReLU layer, respectively. For the threshold used in SCS_{C}, the factor is set to (δh+δl)/2(\delta_{h}+\delta_{l})/2. For segmentation task, to fairly compare with other works, we adopt the standard Deeplab-LargeFOV architecture as our segmentation network, which is based on the VGGNet pre-trained on the ImageNet dataset . Similarly to , we also try the ResNet version Deeplab-LargeFOV architecture and report the results of both versions. The network and conditional random fields (CRFs) hyper-parameters are the same to .

Inference. For our attention network, we resize the input images to a fixed size of 224×224224\times 224 and then resize the resulting attention map back to the original resolution. For segmentation task, following , we perform multi-scale test. For CRFs, we adopt the same code as in .

2 The Role of Self-Erasing

To show the importance of our self-erasing strategies, we perform several ablation experiments in this subsection. Besides showing the results of our standard SeeNet (Fig. 3), we also consider implementing another two network architectures and report the results. First, we re-implement the simple erasing network (ACoL) proposed in (setting 1). The hyper-parameters are all same to the default ones in . This architecture does not use our C-ReLU layer and does not have our SCS_{C} branch as well. Furthermore, to stress the importance of the conditionally sign-flipping operation, we also try to zero the feature units associated with the background regions and keep all other settings unchanged (setting 2).

The quality of attention maps. In Fig. 4, we sample some images from the PASCAL VOC 2012 dataset and show the results by different experiment settings. When localizing small objects as shown on the top two rows of Fig. 4, our attention network is able to better focus on the semantic objects compared to the other two settings. This is due to the fact that our SCS_{C} branch helps better recognize the background regions and hence improves the ability of our approach to keep the attentions from expanding to unexpected non-object regions. When encountering large objects as shown on the bottom two rows of Fig. 4, other than discovering where the semantic objects are, our approach is also capable of mining relatively integral objects compared to the other settings. The conditional reversion operations also protect the attention areas from spreading to the background areas. This phenomenon is specially clear in the monitor image of Fig. 4.

Quantitative results on PASCAL VOC 2012. Besides visual comparisons, we also consider reporting the results by applying our attention maps to the weakly-supervised semantic segmentation task. Given the attention maps, we first carry out a series of operations following the instructions described in Sec. 4, yielding the proxy ground truths of the training set. We utilize the resulting proxy ground truths as supervision to train the segmentation network. The quantitative results on the validation set are listed in Table 1. Note that the segmentation maps are all based on single-scale test and no post-processing tools are used, such as CRFs. According to Table 1, one can observe that with the same saliency maps as background priors, our approach achieves the best results. Compared to the approach proposed in , we have a performance gain of 1.2% in terms of mIoU score, which reflects the high quality of the attention maps produced by our approach.

3 Comparison with the State-of-the-Arts

In this subsection, we compare our proposed approach with existing weakly-supervised semantic segmentation methods that are based on image-level supervision. Detailed information for each method is shown in Table 2. We report the results of each method on both the validation and test sets.

From Table 2, we can observe that our approach greatly outperforms all other methods when the same base model, such as VGGNet , is used. Compared to DCSP , which leverages the same procedures to produce the proxy ground-truths for segmentation segmentation network, we achieves a performance gain of more than 2% on the validation set. This method uses the original CAM as their attention map generator while our approach utilizes the attention maps produced by our SeeNet, which indirectly proofs the better performance of our attention network compared to CAM. To further compare our attention network with adversarial erasing methods, such as AE-PSL and GAIN , our segmentation results are also much better than theirs. This also reflects the high quality of our attention maps.

4 Discussions

To better understand the proposed network, we show some visual results produced by our segmentation network in Fig. 6. As can be seen, our segmentation network works well because of the high-quality attention maps produced by our SeeNet. However, despite the good results, there are still a small number of failure cases, part of which has been shown on the bottom row of Fig. 6. These bad cases are often caused by the fact that semantic objects with different labels are frequently tied together, making the attention models difficult to precisely separate them. Specifically, as attention models are trained with only image-level labels, it is hard to capture perfectly integral objects. In Fig. 5, we show more visual results sampled from the Pascal VOC dataset. As can be seen, some scenes are with complex background or low contrast between the semantic objects and the background. Although our approach has involved background priors to help confine the attention regions, when processing these kinds of images it is hard to localize the whole objects and the quality of the initial attention maps are also essential. In addition, it is still difficult to deal with images with multiple semantic objects as shown in the first row of Fig. 5. The attention networks may easily predict which categories exist in the images but localizing all the semantic objects are not easy. A promising way to solve this problem might be incorporating a small number of pixel-level annotations for each category during the training phase to offer attention networks the information of boundaries. The pixel-level information will tell attention networks where the boundaries of semantic objects are and also accurate background regions that will help produce more integral results. This is also the future work that we will aim at.

Conclusion

In this paper, we introduce the thought of self-erasing into attention networks. We propose to extract background information based on the initial attention maps produced by the initial attention generator by thresholding the maps into three zones. Given the roughly accurate background priors, we design two self-erasing strategies, both of which aim at prohibiting the attention regions from spreading to unexpected regions. Based on the two self-erasing strategies, we build a self-erasing attention network to confine the observable regions in a potential zone which exists semantic objects with high probability. To evaluate the quality of the resulting attention maps, we apply them to the weakly-supervised semantic segmentation task by simply combining it with saliency maps. We show that the segmentation results based on our proxy ground-truths greatly outperform existing state-of-the-art results.

This research was supported by NSFC (NO. 61620106008, 61572264), the national youth talent support program, Tianjin Natural Science Foundation for Distinguished Young Scholars (NO. 17JCJQJC43700), and Huawei Innovation Research Program.

References