Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

Yude Wang, Jie Zhang, Meina Kan, Shiguang Shan, Xilin Chen

Introduction

Semantic segmentation is a fundamental computer vision task, which aims to predict pixel-wise classification results on images. Thanks to the booming of deep learning researches in recent years, the performance of semantic segmentation model has achieved great progress , promoting many practical applications, e.g., autopilot and medical image analysis. However, compared to other tasks such as classification and detection, semantic segmentation needs to collect pixel-level class labels which are time-consuming and expensive. Recently many efforts are devoted to weakly supervised semantic segmentation (WSSS) which utilizes weak supervisions, e.g., image-level classification labels, scribbles, and bounding boxes, attempting to achieve equivalent segmentation performance of fully supervised approaches. This paper focuses on semantic segmentation by image-level classification labels.

To the best of our knowledge, most of advanced WSSS methods are based on the class activation map (CAM) , which is an effective way to localize objects by image classification labels. However, the CAMs usually only cover the most discriminative part of the object and incorrectly activate in background regions, which can be summarized as under-activation and over-activation respectively. Moreover, the generated CAMs are not consistent when images are augmented by affine transformations. As shown in Fig. 1, applying different rescaling transformations on the same input images causes significant inconsistency on the generated CAMs. The essential causes of these phenomena come from the supervision gap between fully and weakly supervised semantic segmentation.

In this paper, we propose a self-supervised equivariant attention mechanism (SEAM) to narrow the supervision gap mentioned above. The SEAM applies consistency regularization on CAMs from various transformed images to provide self-supervision for network learning. To further improve the network prediction consistency, SEAM introduces the pixel correlation module (PCM), which captures context appearance information for each pixel and revises original CAMs by learned affinity attention maps. The SEAM is implemented by a siamese network with equivariant cross regularization (ECR) loss, which regularizes the original CAMs and the revised CAMs on different branches. Fig. 1 shows that our CAMs are consistent over various transformed input images, with fewer over-activated and under-activated regions than baseline. Extensive experiments give both quantitative and qualitative results, demonstrating the superiority of our approach.

We propose a self-supervised equivariant attention mechanism (SEAM), incorporating equivariant regularization with pixel correlation module (PCM), to narrow the supervision gap between fully and weakly supervised semantic segmentation.

The design of siamese network architecture with equivariant cross regularization (ECR) loss efficiently couples the PCM and self-supervision, producing CAMs with both fewer over-activated and under-activated regions.

Experiments on PASCAL VOC 2012 illustrate that our algorithm achieves state-of-the-art performance with only image-level annotations.

Related Work

The development of deep learning has led to a series of breakthroughs on fully supervised semantic segmentation in recent years. In this section, we introduce some works, including weakly supervised semantic segmentation and self-supervised learning.

Compared to fully supervised learning, WSSS uses weak labels to guide network training, e.g., bounding boxes , scribbles and image-level classification labels . A group of advanced researches utilizes image-level classification labels to train models. Most of them refine the class activation map (CAM) generated by the classification network to approximate the segmentation mask. SEC proposes three principles, i.e., seed, expand, and constrain, to refine CAMs, which are followed by many other works. Adversarial erasing is a popular CAM expansion method, which erases the most discriminative part of CAM, guides the network to learn classification features from other regions and expands activations. AffinityNet trains another network to learn the similarity between pixels, which generates a transition matrix and multiplies with CAM several times to adjust its activation coverage. IRNet generates a transition matrix from the boundary activation map and extends the method to weakly supervised instance segmentation. Here are also some researches endeavor to aggregate self-attention module in the WSSS framework, e.g., CIAN proposes cross-image attention module to learn activation maps from two different images containing the same class objects with the guidance of saliency maps.

2 Self-supervised Learning

Instead of using massive annotated labels to train network, self-supervised learning approaches aim at designing pretext tasks to generate labels without additional manual annotations. Here are many classical self-supervised pretext tasks, e.g., relative position prediction , spatial transformation prediction , image inpainting , and image colorization . To some extent, the generative adversarial network can also be regarded as a self-supervised learning approach that the authenticity labels for discriminator do not need to be annotated manually. Labels generated by pretext tasks provide self-supervision for the network to learn a more robust representation. The feature learned by self-supervision can replace the feature pretrained by ImageNet on some tasks, such as detection and part segmentation .

Considering there is a large supervision gap between fully and weakly supervised semantic segmentation, it is an intuition that we should seek additional supervision to narrow the gap. Since image-level classification labels are too weak for network to learn segmentation masks which should well fit object boundary, we design pretext task using the equivariance of ideal segmentation function to provide additional self-supervision for network learning with only image-level annotations.

Approach

This section details our SEAM method. Firstly, we illustrate the motivation of our work. Then we introduce the implementation of equivariant regularization by a shared-weight siamese network. The proposed pixel correlation module (PCM) is integrated into the network to further improve the consistency of prediction. Finally, the loss design of SEAM is discussed. Fig. 2 shows our SEAM network structure.

Self-attention is a widely accepted mechanism that can significantly improve the network approximation ability. It revises feature maps by capturing context feature dependency, which also meets the ideas of most WSSS methods using the similarity of pixels to refine the original activation map. Following the denotation of , the general self-attention mechanism can be defined as:

2 Equivariant Regularization

During the data augmentation period of fully supervised semantic segmentation, the pixel-level labels should be applied with the same affine transformation as input images. It introduces an implicit equivariant constraint for the network. However, considering that the WSSS can only access image-level classification labels, the implicit constraint is missing here. Therefore, we propose equivariant regularization as follows:

Here F(⋅)F(\cdot) denotes the network, and A(⋅)A(\cdot) denotes any spatial affine transformation, e.g., rescaling, rotation, flip. To integrate regularization on the original network, we expand the network into a shared-weight siamese structure. One branch applies the transformation on the network output, the other branch warps the images by the same transformation before the feedforward of the network. The output activation maps from two branches are regularized to guarantee the consistency of CAMs.

3 Pixel Correlation Module

Although equivariant regularization provides additional supervision for network learning, it is hard to achieve ideal equivariance with only classical convolution layers. Self-attention is an efficient module to capture context information and refine pixel-wise prediction results. To integrate the classical self-attention module given by Eq. (1) and Eq. (2) for CAM refinement, the formulation can be written as:

To further refine original CAMs by context information, we propose a pixel correlation module (PCM) at the end of the network to integrate the low-level feature of each pixel. The structure of PCM refers to the core part of the self-attention mechanism with some modifications and trained by the supervision from equivariant regularization. We use cosine distance to evaluate inter-pixel feature similarity:

Here we take the inner-product in normalized feature space to calculate the affinity between current pixel ii and others. The ff can be integrated into Eq. (1) with some modifications as:

The similarities are activated by ReLU to suppress negative values. The final CAM is the weighted sum of the original CAM with normalized similarities. Fig. 3 gives an illustration of the PCM structure.

Compared to classical self-attention, PCM removes the residual connection to keep the same activation intensity of the original CAM. Moreover, since the other network branch provides pixel-level supervision for PCM, which is not as accurate as ground truth, we reduce parameters by removing embedding function ϕ\phi and gg to avoid overfitting on inaccurate supervision. We use ReLU activation function with L1 normalization to mask out irrelevant pixels and generate an affinity attention map which is smoother in relevant regions.

4 Loss Design of SEAM

The classification loss provides learning supervision for object localization. And it is necessary to aggregate equivariant regularization on original CAM to preserve the consistency of output. The equivariant regularization (ER) loss on original CAM can be easily defined as:

The PCM outputs are regularized by the original CAMs on the other branch of the siamese network. This strategy can avoid CAM degeneration during PCM refinement.

Although the CAMs are learned by foreground object classification loss, there are many background pixels, which should not be ignored during PCM processing. The original foreground CAMs have zero vectors on these background positions, which cannot produce gradients to push feature representations closer between those background pixels. Therefore, we define the background score as:

In summary, the final loss of SEAM is defined as:

The classification loss is used to roughly localize objects and the ER loss is used to narrow the gaps between pixel- and image-level supervisions. The ECR loss is used to integrate PCM with the trunk of the network, in order to make consistent predictions over various affine transformations. The network architecture is illustrated in Fig. 2. We give the details of network training settings and carefully investigate the effectiveness of each module in the experiments section.

Experiments

We evaluate our approach on PASCAL VOC 2012 dataset with 21 class annotations, i.e., 20 foreground objects and the background. The official dataset separation has 1464 images for training, 1449 for validation and 1456 for testing. Following the common experimental protocol for semantic segmentation, we take additional annotations from SBD to build an augmented training set with 10582 images. Noting that only image-level classification labels are available during network training. Mean intersection over union (mIoU) is used as a metric to evaluate segmentation results.

In our experiments, ResNet38 is adopted as backbone network with output_stride=8output\_stride=8. We extract the feature maps from stage 3 and stage 4, reduce their channel numbers into 64 and 128 respectively by individual 1×11\times 1 convolution layers. In PCM, these features are concatenated with images and fed into function θ\theta in Eq. (5), which is implemented by another 1×11\times 1 convolution layer. The images are randomly rescaled in the range of by the longest edge and then cropped by 448×448448\times 448 as network inputs. The model is trained on 4 TITAN-Xp GPUs with batch size 8 for 8 epochs. The initial learning rate is set as 0.01, following the poly policy lritr=lrinit(1−itrmax_itr)γ\mathit{lr}_{\mathit{itr}}=\mathit{lr}_{\mathit{init}}(1-\frac{itr}{\mathit{max\_itr}})^{\gamma} with γ=0.9\gamma=0.9 for decay. Online hard example mining (OHEM) is employed on the ECR loss remaining the largest 20%20\% pixel losses.

During network training, we cut off gradients back-propagation at the intersection point between PCM stream and the trunk of the network to avoid the mutual interference. This setting simplifies the PCM into a pure context refinement module which still can be trained with the backbone of the network at the same time. And the learning of original CAMs will not be affected by PCM refinement process. During inference, since our SEAM is a shared-weight siamese network, only one branch needs to be restored. We adopt multi-scale and flip test during inference to generate pseudo segmentation labels.

2 Ablation Studies

To verify the effectiveness of our SEAM, we generate pixel-level pseudo labels from revised CAMs on PASCAL VOC 2012 train set. In our experiments, we traverse all background threshold options and give the best mIoU of pseudo labels, instead of comparing with the same background threshold. Because the highest pseudo label accuracy represents the best matching results between CAMs and ground truth segmentation masks. Specifically, the foreground activation coverage will expand with the increase of average activation intensity, while its matching degree with ground truth is not changed. And the highest pseudo label accuracy will not be improved when CAMs only increase average activation intensity rather than becoming more matchable with ground truth.

Tab. 1 gives an ablation study of each module in our approach. It shows that using the siamese network with equivariant regularization has a 2.47% improvement compared to baseline. Our PCM achieves significant performance elevation by 5.18%. After applying OHEM on equivariant cross regularization loss, the generated pseudo labels further achieve 55.41% mIoU on PASCAL VOC train set. We also test the baseline CAM with dense CRF to refine predictions. The results show that dense CRF improves the mIoU to 52.40%, which is lower than the SEAM result 55.41%. And our SEAM can further improve the performance up to 56.83% after aggregating dense CRF as post process. Fig. 4 shows that the CAMs generated by SEAM have fewer over-activations and more complete activation coverage, whose shape is closer to the ground truth segmentation masks than baseline. To further verify the effectiveness of our proposed SEAM, we visualize the affinity attention maps generated by PCM. As shown in Fig. 5, the selected foreground and background pixels are very close in spatial, while their affinity attention maps are greatly different. It proves that the PCM can learn boundary sensitive features from self-supervision.

Improved Localization Mechanism:

It is an intuition that improved weakly supervised localization mechanism will elevate mIoU of pseudo segmentation labels. To verify the idea, we simply evaluate GradCAM and GradCAM++ before aggregating our proposed SEAM. However, the evaluation results given by Tab. 2 illustrates that both GradCAM and GradCAM++ cannot narrow the supervision gap between fully and weakly supervised semantic segmentation tasks, since the best mIoU results do not have improvement. We believe the improved localization mechanisms are only designed to represent object correlated parts without any constraints by low-level information, which is not suitable for the segmentation task. The CAMs generated by these improved localization methods are not becoming more matchable with ground truth masks. The following experiments further illustrate that our proposed SEAM can substantially improve the quality of CAM to fit the shape of object masks.

Affine Transformation:

Ideally, the A(⋅)A(\cdot) in Eq. (3) can be any affine transformation. Several transformations are conducted in the siamese network to evaluate the effect of them on equivariant regularization. As shown in Tab. 3, there are four candidate affine transformations: rescaling with 0.3 down-sampling rate, random rotation in degrees, translation by 15 pixels and horizontal flip. Firstly, our proposed SEAM simply adopts rescaling during network training. Tab. 3 shows that the mIoU of pseudo labels has significant improvement from 47.43% to 55.41%. Tab. 3 also shows that simply incorporating different transformations is not much effective. When rescaling transformation integrates with flip, rotation, and translation respectively, only flip makes tiny improvement. In our view, it is because the activation maps between flip, rotation, and translation are too similar to produce sufficient supervision. Without additional instructions, we only preserve rescaling as the key transformation with 0.30.3 down-sampling rate in our other experiments.

Augmentation and Inference:

Compared to the original one-branch network, the siamese structure expands the augmentation range of image size in practice. To investigate whether the improvement stems from the rescaling range, we evaluate the baseline model with a larger scale range and Tab. 4 gives the experiment results. It shows that simply increasing the rescaling range cannot improve the accuracy of generated pseudo labels, which proves that the performance improvement comes from the combination of PCM and equivariant regularization instead of data augmentation.

During inference, it is a common practice to employ multi-scale test by aggregating the prediction results from images with different scales to boost the final performance. It can also be regarded as a method to improve the equivariance of predictions. To verify the effectiveness of our propose SEAM, we evaluate the CAMs generated by both single-scale and multi-scale test. Tab. 5 illustrates that our proposed model outperforms baseline with higher peak performance in both single- and multi-scale test.

Source of Improvement:

The improvement of CAM quality mainly stems from more complete activation coverage or fewer over-activated regions. To further analyze the improvement source of our SEAM, we define two metrics to represent the degree of under-activation and over-activation:

Here TPc\mathit{TP}_{c} denotes the pixel number of true positive prediction of class cc, FPc\mathit{FP}_{c} and FNc\mathit{FN}_{c} denote false positive and false negative respectively. These two metrics exclude the background category since the prediction of background is inverse to the foreground. Specifically, if there are more false negative regions when CAMs do not have complete activation coverage, mFNm_{\mathit{FN}} will have a larger value. Relatively, larger mFPm_{\mathit{FP}} means there are more false positive regions, meaning that CAMs are over-activated.

Based on these two metrics, we collect the evaluation results from both baseline and our SEAM, then plot the curves in Fig. 6 which illustrates a large gap between baseline and our method. The SEAM achieves lower mFNm_{\mathit{FN}} and mFPm_{\mathit{FP}}, meaning that the CAMs generated by our approach have more complete activation coverage and fewer over-activated pixels. Therefore, the prediction maps of SEAM better fit the shape of ground truth segmentation. Moreover, the curves of SEAM are more consistent than baseline model over different image scales, which proves that the equivariance regularization works during network learning and contributes to the improvement of CAM.

3 Comparison with State-of-the-arts

To further elevate the accuracy of pseudo pixel-level annotations, we follow the work of to train an AffinityNet based on our revised CAM. The final synthesized pseudo labels achieve 63.61% mIoU on PASCAL VOC 2012 train set. Then we train the classical segmentation model DeepLab with ResNet38 backbone on these pseudo labels in full supervision to achieve final segmentation results. Tab. 6 shows the mIoU of each class on val set and Tab. 7 gives more experiment results of previous approaches. Compared to the baseline method, our SEAM significantly improves the performance on both val and test set with the same training setting. Moreover, our method presents the state-of-the-art performance using only image-level labels on PASCAL VOC 2012 test set. Noting that our performance elevation stems from neither the larger network structure nor the improved saliency detector. The performance improvement mainly comes from the cooperation of additional self-supervision and PCM, which produces better CAMs for the segmentation task. Fig. 7 shows some qualitative results, which verify that our method works well on both large and small objects.

Conclusion

In this paper, we propose a self-supervised equivariant attention mechanism (SEAM) to narrow the supervision gap between fully and weakly supervised semantic segmentation by introducing additional self-supervision. The SEAM embeds self-supervision into weakly supervised learning framework by exploiting equivariant regularization, which forces CAMs predicted from various transformed images to be consistent. To further improve the ability of network for generating consistent CAMs, a pixel correlation module (PCM) is designed, which refines original CAMs by learning inter-pixel similarity. Our SEAM is implemented by a siamese network structure with efficient regularization losses. The generated CAMs not only keep consistent over different transformed inputs but also better fit the shape of ground truth masks. The segmentation network retrained by our synthesized pixel-level pseudo labels achieves state-of-the-art performance on PASCAL VOC 2012 dataset, which proves the effectiveness of our SEAM.

This work was partially supported by National Key R&D Program of China (No. 2017YFA0700800), CAS Frontier Science Key Research Project (No. QYZDJ-SSWJSC009) and Natural Science Foundation of China (Nos. 61806188, 61772496).

References