Mining Cross-Image Semantics for Weakly Supervised Semantic Segmentation

Guolei Sun, Wenguan Wang, Jifeng Dai, Luc Van Gool

Introduction

Recently, modern deep learning based semantic segmentation models​ , trained with massive manually labeled data, achieve far better performance than before. However, the fully supervised learning paradigm has the main limitation of requiring intensive manual labeling effort, which is particularly expensive for annotating pixel-wise ground-truth for semantic segmentation. Numerous efforts are motivated to develop semantic segmentation with weaker forms of supervision, such as bounding boxes​ , scribbles​ , points​ , and image-level labels​ , etc. Among them, a prominent and appealing trend is using only image-level labels to achieve weakly supervised semantic segmentation (WSSS), which demands the least annotation efforts and is followed in this work.

To tackle the task of WSSS with only image-level labels, current popular methods are based on network visualization techniques​ , which discover discriminative regions that are activated for classification. These methods use image-level labels to train a classifier network, from which class-activation maps are derived as pseudo ground-truths for further supervising pixel-level semantics learning. However, it is commonly evidenced that the trained classifier tends to over-address the most discriminative parts rather than entire objects, which becomes the focus of this area. Diverse solutions are explored, typically adopting: image-level operations, such as region hiding and erasing​ , regions growing strategies that expand the initial activated regions​ , and feature-level enhancements that collect multi-scale context from deep features​ .

These efforts generally achieve promising results, which demonstrates the importance of discriminative object pattern mining for WSSS. However, as shown in Fig. 1(a), they typically use only single-image information for object pattern discovering, ignoring the rich semantic context among the weakly annotated data. For example, with the image-level labels, not only the semantics of each individual image can be identified, the cross-image semantic relations, i.e., two images whether sharing certain semantics, are also given and should be used as cues for object pattern mining. Inspired by this, rather than relying on intra-image information only, we further address the value of cross-image semantic correlations for complete object pattern learning and effective class-activation map inference (see Fig. 1(b-c)). In particular, our classifier is equipped with a differentiable co-attention mechanism that addresses semantic homogeneity and difference understanding across training image pairs. More specifically, two kinds of co-attentions are learned in the classifier. The former one aims to capture cross-image common semantics, which enables the classifier to better ground the common semantic labels over the co-attentive regions. The latter one, called contrastive co-attention, focuses on the rest, unshared semantics, which helps the classifier better separate semantic patterns of different objects. These two co-attentions work in a cooperative and complimentary manner, together making the classifier understand object patterns more comprehensively.

In addition to benefiting object pattern learning, our co-attention provides an efficient tool for precise localization map inference (see Fig. 1(c)). Given a training image, a set of related images (i.e., sharing certain common semantics) are utilized by the co-attention for capture richer context and generate more accurate localization maps. Another advantage is that our co-attention based classifier learning paradigm brings an efficient data augmentation strategy, due to the use of training image pairs. Overall, our co-attention boosts object discovering during both the classifier’s training phase as well as localization map inference stage. This provides the possibility of obtaining more accurate pseudo pixel-level annotations, which facilitate final semantic segmentation learning.

Our algorithm is a unified and elegant framework, which generalizes well different WSSS settings. Recently, to overcome the inherent limitation in WSSS without additional human supervision, some efforts resort to extra image-level supervision from simple single-class data readily available from other existing datasets , or cheap web-crawled data . Although they improve the performance to some extent, complicated techniques, such as energy function optimization , heuristic constraints , and curriculum learning , are needed to handle the challenges of domain gap and data noise, restricting their utility. However, due to the use of paired image data for classifier training and object map inference, our method has good tolerance to noise. In addition, our method also handles domain gap naturally, as the co-attention effectively addresses domain-shared object pattern learning and achieves domain adaption as a part of co-attention parameter learning. We conduct extensive experiments on PASCAL VOC 2012 , under three WSSS settings, i.e., learning WSSS with (1) PASCAL VOC image-level supervision only, (2) extra simple single-label data, and (3) extra web data. Our algorithm sets state-of-the-art on each case, verifying its effectiveness and generalizability. Our method also ranked 1st place in the Weakly-supervised Semantic Segmentation Track of CVPR2020 Learning from Imperfect Data (LID) Challenge​ (LID20 ⁣{}_{20\!}), outperforming other competitors by large margins.

Our contributions are three-fold. (1) We address the value of cross-image semantic correlations for complete object pattern learning as well as object location inference, which is achieved by a co-attention classifier that works over paired training samples. (2) Our co-attention classifier mines semantic cues in a more comprehensive manner. In addition to single-image semantics, it mines complimentary supervision from cross-image semantic similarities and differences by the co-attention and contrastive co-attention, respectively. (3) Our approach is general enough to learn WSSS with precise image-level supervision, or with extra simple single-label, or even noisy web-crawled data. It solves inherent challenges of different WSSS settings elegantly, and shows promising results consistently.

Related Work

Weakly Supervised Semantic Segmentation. Recently, lots of WSSS methods have been proposed to alleviate labeling cost. Various weak supervision forms have been explored, such as bounding boxes , scribbles , point supervision , etc. Among them, image-level supervision, due to its less annotation demand, gains most attention and is also adopted in our approach.

Current popular solutions for WSSS with image-level supervision rely on network visualization techniques , especially the Class Activation Map (CAM) , which discovers image pixels that are informative for classification. However, CAM typically only identifies small discriminative parts of objects, making it not an ideal proxy ground-truth for semantic segmentation training. Therefore, numerous efforts are made towards expanding the CAM-highlighted regions to the whole objects. In particular, some representative approaches make use of image-level hiding and erasing operations to drive a classifier to focus on different parts of objects . A few ones instead resort to a regions growing strategy, i.e., view the CAM-activated regions as initial “seeds” and gradually grow the seed regions until cover the complete objects . Meanwhile, some researchers investigate to directly enhance the activated regions on feature-level . When constructing CAMs, they collect multi-scale context, which is achieved by dilated convolution , multi-layer feature fusion , saliency-guided iterative training , or stochastic feature selection . Some others accumulate CAMs from multiple training phases , or self-train a difference detection network to complete the CAMs with trustable information . In addition, a recent trend is to utilize class-agnostic saliency cues to filter out background responses during localization map inference.​​

Since the supervision provided in above problem setting is so weak, another category of approaches explores to leverage more image-level supervision from other sources. There are mainly two types: (1) exploring simple and single-label examples (e.g., images from existing datasets ); or (2) utilizing near-infinite yet noisy web-sourced image or video data (also referred as webly supervised semantic segmentation ). In addition to the common challenge of domain gap between the extra data and target semantic segmentation dataset, the second-type methods need to handle data noise.

Past efforts only consider each image individually, while only few exceptions address cross-image information. simply applies off-the-shelf co-segmentation over the web images to generate foreground priors, instead of ours encoding the semantic relations into network learning and inference. For , although also exploiting correlations within image pairs, the core idea is to use extra information from a support image to supplement current visual representations. Thus the two images are expected to better contain the same semantics, and unmatched semantics would bring negative influences. In contrast, we view both semantic homogeneity and difference as informative cues, driving our classifier to more explicitly identify the common as well as unshared objects, respectively. Moreover, only utilizes single image to infer the activated objects, but our method comprehensively leverages the cross-image semantics in both classifier training and localization map inference stages. More essentially, our framework is neat and flexible, which is not only able to learn WSSS from clean image-level supervision, but general enough to naturally make use of extra noisy web-crawled or simple single-label data, contrarily to previous efforts which are limited to specific training settings and largely dependent on complicated optimization methods or heuristic constraints .

Deterministic Neural Attention. Differentiable attention mechanisms enable a neural network to focus more on relevant elements of the input than on irrelevant parts. With their popularity in the field of natural language processing , attention modeling is rapidly adopted in various computer vision tasks, such as image recognition , domain adaptation , human pose estimation , reasoning among objects , and image generation . Further, co-attention mechanisms become an essential tool in many vision-language applications and sequential modeling tasks, such as visual question answering , visual dialog , vision-language navigation , and video segmentation , showing its effectiveness in capturing the underlying relations between different entities. Inspired by the general idea of attention mechanisms, this work leverages co-attention to mine semantic relations within training image pairs, which helps the classifier network learn complete object patterns and generate precise object localization maps.

Methodology

Problem Setup. Here we follow current popular WSSS pipelines: given a set of training images with image-level labels, a classification network is first trained to discover corresponding discriminative object regions. The resulting object localization maps over the training samples are refined as pseudo ground-truth masks to further supervise the learning of a semantic segmentation network.

Our Idea. Unlike most previous efforts that treat each training image individually, we explore cross-image semantic relations as class-level context for understanding object patterns more comprehensively. To achieve this, two neural co-attentions are designed. The first one drives the classifier to learn common semantics from the co-attentive object regions, while the other one enforces the classifier to focus on the rest objects for unshared semantics classification.

So far the classifier is learned in a standard manner, i.e., only individual-image information is used for semantic learning. One can directly use the activation maps to supervise next-stage semantic segmentation learning, as done in . Differently, our classifier additionally utilizes a co-attention mechanism for further mining cross-image semantics and eventually better localizing objects.

Co-Attention for Cross-Image Common Semantics Mining. Our co-attention attends to the two images, i.e., Im\bm{I}_{m} and In\bm{I}_{n}, simultaneously, and captures their correlations. We first compute the affinity matrix P\bm{P} between Fm\bm{F}_{m} and Fn\bm{F}_{n}:

Then P\bm{P} is normalized column-wise to derive attention maps across Fm ⁣\bm{F}_{m\!} for each position ⁣{}_{\!} in ⁣{}_{\!} Fn\bm{F}_{n}, and ⁣{}_{\!} row-wise ⁣{}_{\!} to ⁣{}_{\!} derive attention maps across Fn ⁣\bm{F}_{n\!} for each position in Fm\bm{F}_{m}:

where softmax is performed column-wise. In this way, An\bm{A}_{n} and Am\bm{A}_{m} store the co-attention maps in their columns. Next, we can compute attention summaries of Fm\bm{F}_{m} (Fn\bm{F}_{n}) in light of each position of Fn\bm{F}_{n} (Fm\bm{F}_{m}):

Through co-attention computation, not only the human face, the most discriminative part of Person, but also other parts, such as legs and arms, are highlighted in Fmm∩n\bm{F}^{m\cap n}_{m} and Fnm∩n\bm{F}^{m\cap n}_{n} (see Fig. 2(b)). When we set the common class labels, i.e., Person, as the supervision signal, the classifier would realize that the semantics preserved in Fmm∩n\bm{F}^{m\cap n}_{m} and Fnm∩n\bm{F}^{m\cap n}_{n} are related and can be used to recognize Person. Therefore, the co-attention, computed across two related images, explicitly helps the classifier associate semantic labels and corresponding object regions and better understand the relations between different object parts. It essentially makes full use of the context across training data.

Intuitively, for the co-attention based common semantic classification, the labels lm ⁣∩ln\bm{l}_{m\!}\cap\bm{l}_{n} shared between Im\bm{I}_{m} and In\bm{I}_{n} are used to supervise learning:

Contrastive Co-Attention for Cross-Image Exclusive Semantics Mining. Aside from the co-attention described above that explores cross-image common semantics, we propose a contrastive co-attention that mines semantic differences between paired images. The co-attention and contrastive co-attention complementarily help the classifier better understand the concept of the objects.

As shown in Fig. 2(a), for Im\bm{I}_{m} and In\bm{I}_{n}, we first derive class-agnostic co-attentions from their co-attentive features, i.e., Fmm∩n ⁣\bm{F}^{m\cap n\!}_{m} and Fnm∩n ⁣\bm{F}^{m\cap n\!}_{n}, respectively:

Compared with the co-attention that investigates common semantics as informative cues for boosting object patterns mining, the contrastive co-attention addresses complementary knowledge from the semantic differences between paired images. Fig. 2(b) gives an intuitive example. After computing the contrastive co-attentions between Im\bm{I}_{m} and In\bm{I}_{n} (Eq. 7), Table and Cow, which are unique in their original images, are highlighted. Based on the contrastive co-attentive features, i.e., Fmm ⁣∖ ⁣n\bm{F}^{m\!\setminus\!n}_{m} and Fnn ⁣∖ ⁣m\bm{F}^{n\!\setminus\!m}_{n}, the classifier is required to accurately recognize Table and Cow classes, respectively. When the common objects are filtered out by the contrastive co-attentions, the classifier has a chance to focus more on the rest image regions and mine the unshared semantics more consciously. This also helps the classifier better discriminate the semantics of different objects, as the semantics of common objects and unshared ones are disentangled by the contrastive co-attention. For example, if some parts of Cow are wrongly recognized as Person-related, the contrastive co-attention will discard these parts in Fnn ⁣∖ ⁣m\bm{F}^{n\!\setminus\!m}_{n}. However, the rest semantics in Fnn ⁣∖ ⁣m\bm{F}^{n\!\setminus\!m}_{n} may be not sufficient enough for recognizing Cow. This will enforce the classifier to better discriminate different objects.

For the contrastive co-attention based unshared semantic classification, the supervision loss is designed as:

More In-Depth Discussion. One can interpret our co-attention classifier from a view of auxiliary-task learning , which is investigated in self-supervised learning field to improve data efficiency and robustness, by exploring auxiliary tasks from inherent data structures. In our case, rather than the task of single-image semantic recognition which has been extensively studied in conventional WSSS methods, we explore two auxiliary tasks, i.e., predicting the common and uncommon semantics from image pairs, for fully mining supervision signals from weak supervision. The classifier is driven to better understand the cross-image semantics by attending to (contrastive) co-attentive features, instead of only relying on intra-image information (see Fig. 2(c)). In addition, such strategy shares a spirit of image co-segmentation. Since the image-level semantics of training set are given, the knowledge about some images share or unshare certain semantics should be used as a cue, or supervision signal, to better locate corresponding objects. Our co-attention based learning pipeline also provides an efficient data augmentation strategy, due to the use of paired samples, whose amount is near the square of the number of single training images.

2 Co-Attention Classifier Guided WSSS Learning

Training Co-Attention Classifier. The overall training loss for our co-attention classifier ensembles the three terms defined in Eq. 1, 5, and 9:

The coefficients of different loss terms are set as 1 in our all experiments. During training, to fully leverage the co-attention to mine the common semantics, we sample two images (Im,In)(\bm{I}_{m},\bm{I}_{n}) with at least one common class, i.e., lm ⁣∩ ⁣ln ⁣≠ ⁣0\bm{l}_{m}\!\cap\!\bm{l}_{n}\!\neq\!\mathbf{0}.

Generating Object Localization Maps. Once our image classifier is trained, we apply it over the training data I ⁣= ⁣{(In,ln)}n\mathcal{I}\!=\!\{(\bm{I}_{n},\bm{l}_{n})\}_{n} to produce corresponding object localization maps, which are essential for semantic segmentation network training. We explore two different strategies to generate localization maps.

These two localization map generation strategies are studied in our experiments (§4.5), and the last one is more favored, as it uses both intra- and inter-image semantics for object inference, and shares a similar data distribution of the training phase. One may notice that the contrastive co-attention is not used here. This is because contrastive co-attentive feature (Eq. 8) is from its original image, which is effective for boosting feature representation learning during classifier training, while contributes little for localization maps inference (with limited cross-image information). Related experiments can be found at §4.5.

Learning Semantic Segmentation Network. After obtaining high-quality localization maps, we generate pseudo pixel-wise labels for all the training samples I\mathcal{I}, which can be used to train arbitrary semantic segmentation network. For pseudo groundtruth generation, we follow current popular pipeline , that uses localization maps to extract class-specific object cues and adopts saliency maps to get background cues. For the semantic segmentation network, as in , we choose DeepLab-LargeFOV .

Learning with Extra Simple Single-Label Images. Some recent efforts are made towards exploring extra simple single-label images from other existing datasets for further boosting WSSS. Though impressive, specific network designs are desired, due to the issue of domain gap between additionally used data and the target complex multi-label dataset, i.e., PASCAL VOC 2012 . Interestingly, our co-attention based WSSS algorithm provides an alternate that addresses the challenge of domain gap naturally. Here we revisit the computation of co-attention in Eq. 2. When Im\bm{I}_{m} and In\bm{I}_{n} are from different domains, the parameter matrix W ⁣P\bm{W}_{\!\bm{P}}, in essence, learns to map them into a unified common semantic space and the co-attentive features can capture domain-shared semantics. Therefore, for such setting, we learn three different parameter matrixes for W ⁣P\bm{W}_{\!\bm{P}}, for the cases where Im\bm{I}_{m} and In\bm{I}_{n} are from (1) the target semantic segmentation domain, (2) the one-label image domain, and (3) two different domains, respectively. Thus the domain adaption is efficiently achieved as a part of co-attention learning. We conduct related experiments in §4.2.

Learning with Extra Web Images. Another trend of methods address webly supervised semantic segmentation, i.e., leveraging web images as extra training samples. Though cheaper, web data are typically noisy. To handle this, previous arts propose diverse effective yet sophisticated solutions, such as multi-stage training and self-paced learning . Our co-attention based WSSS algorithm can be easily extended to this setting and solve data noise elegantly. As our co-attention classifier is trained with paired images, instead of previous methods only relying on each image individually, our model provides a more robust training paradigm. In addition, during localization map inference, a set of extra related images are considered, which provides more comprehensive and accurate cues, and further improves the robustness. We ⁣{}_{\!} experimentally ⁣{}_{\!} demonstrate ⁣{}_{\!} the ⁣{}_{\!} effectiveness ⁣{}_{\!} of ⁣{}_{\!} our ⁣{}_{\!} method ⁣{}_{\!} in ⁣{}_{\!} such ⁣{}_{\!} a ⁣{}_{\!} setting ⁣{}_{\!} in ⁣{}_{\!} §4.3.

3 Detailed Network Architecture

Network Configuration. In line with conventions , our image classifier is based on ImageNet pre-trained VGG-16 . For VGG-16 network, the last three fully-connected layers are replaced with three convolutional layers with 512 channels and kernel size 3 ⁣× ⁣33\!\times\!3, as done in . For the semantic segmentation network, for fair comparison with current top-leading methods , we adopt the ResNet-101 version Deeplab-LargeFOV architecture.

Training Phases of the Co-Attention Classifier and Semantic Segmentation Network. Our co-attention classifier is fully end-to-end trained by minimizing the loss defined in Eq. 10. The training parameters are set as: initial learning rate (0.001) which is reduced by 0.1 after every 6 epochs, batch size (5), weight decay (0.0002), and momentum (0.9). Once the classifier is trained, we generate localization maps and pseudo segmentation masks over all the training samples (see §3.2). Then, with the masks, the semantic segmentation network is trained in a standard way using the hyper-parameter setting in .

Inference ⁣{}_{\!} Phase ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} Semantic ⁣{}_{\!} Segmentation ⁣{}_{\!} Network. Given ⁣{}_{\!} an ⁣{}_{\!} unseen ⁣{}_{\!} test ⁣{}_{\!} image, our ⁣{}_{\!} segmentation ⁣{}_{\!} network ⁣{}_{\!} works ⁣{}_{\!} in ⁣{}_{\!} the ⁣{}_{\!} standard ⁣{}_{\!} semantic ⁣{}_{\!} segmentation ⁣{}_{\!} pipeline ⁣{}_{\!} , i.e., directly ⁣{}_{\!} generating ⁣{}_{\!} segments ⁣{}_{\!} without ⁣{}_{\!} using any other images. Then CRF post-processing is performed to refine predicted masks.

Note that above settings are used in traditional WSSS datasets (i.e., §4.1, §4.2, §4.3). Due to the specific task setup in LID20 ⁣{}_{20\!}​ , corresponding training and testing settings will be detailed in §4.4.

Experiment

Overview. Experiments are first conducted over three different WSSS settings: (1) The most standard paradigm that only allows image-level supervision from PASCAL VOC 2012 (see §4.1). (2) Following , additional single-label images can be used, yet bringing the challenge of domain gap (see §4.2). (3) Webly supervised semantic segmentation paradigm , where extra web data can be accessed (see §4.3). Then, in §4.4, we show the results in WSSS track of LID20 ⁣{}_{20\!}, where our method achieves the champion. Finally, in §4.5, ablation studies are made to assess the effectiveness of essential parts of our algorithm.

Evaluation Metric. In our experiments, the standard intersection over union (IoU) criterion is reported on the val and test sets of PASCAL VOC 2012 . The scores on test set are obtained from official PASCAL VOC evaluation server.

Experimental Setup: We first conduct experiment following the most standard setting that learns WSSS with only image-level labels , i.e., only image-level supervision from PASCAL VOC 2012 is accessible. PASCAL VOC 2012 contains a total of 20 object categories. As in , augmented training data from are also used. Finally, our model is trained on totally 10,582 samples with only image-level annotations. Evaluations are conducted on the val and test sets, which have 1,449 and 1,456 images, respectively.

Experimental Results: Table 1(a) compares our approach and current top-leading WSSS methods (highest mIoU is used for comparison) with image-level supervision, on both PASCAL VOC12 val and test sets. Additionally, we show some segmentation results in Fig. 3. We can observe that our method achieves mIoU scores of 66.2 and 66.9 on val and test sets respectively, outperforming all the competitors. The performance of our method is 87% of the DeepLab-LargeFOV trained with fully annotated data, which achieved an mIoU of 76.3 on val set. When compared to OAA+ , current best-performing method, our approach obtains the improvement of 1.0% on val set. This well demonstrates that the localization maps produced by our co-attention classifier effectively detect more complete semantic regions towards the whole target objects. Note that our network is elegantly trained end-to-end in a single phase. In contrast, many other recent approaches including OAA+ and SSDD , use extra networks to learn auxiliary information (e.g., integral attention , pixel-wise semantic affinity , etc.), or adopt multi-step training .

Experimental Setup: Following , we train our co-attention classifier and segmentation network with PASCAL images and extra single-label images. The extra single-label images are borrowed from the subsets of Caltech-256 and ImageNet CLS-LOC , and whose annotations are within 20 VOC object categories. There are a total of 20,057 extra single-label images.

Experimental Results: The comparisons are shown in Table 1(c). Our method significantly improves the most recent method (i.e., AttnBN ) in this setting by 5.0% and 4.2% in val and test sets, respectively. With the fact that objects of the same category but from different domains share similar visual patterns , our co-attention provides an end-to-end strategy that efficiently captures the common, cross-domain semantics, and learns domain adaption naturally. Even AttnBN is specifically designed for addressing such setting by knowledge transfer, our method still suppresses it by a large margin. Compared with the setting in §4.1 where only PASCAL images are used for training, our method obtains improvements on both val and test sets, verifying that it successfully mines knowledge from extra simple single-label data and copes with domain gap well.

3 Experiment 3: Learn WSSS with Extra Web-Sourced Data

Experimental Setup: We also conduct experiments using both PASCAL VOC images and webly craweled images as training data. We use the web data provided by , which are retrieved from Bing based on class names. The final dataset contains 76,683 images across 20 PASCAL VOC classes.

Experimental Results: Table 1(c) shows the performance comparisons between our method and the previous webly supervised segmentation methods. It shows that our method outperforms all other approaches and sets new state-of-the-arts with mIoU score of 67.7 and 67.5 on PASCAL VOC 2012 val and test sets, respectively. Among the compared methods, Hong et al. utilize richer information of the temporal dynamics provided by additional large-scale videos. In contrast, although only using static image data, our method still outperforms it on the val and test sets by 9.6% and 8.8%, respectively. Compared with Shen et al. which uses the same web data as ours, our method substantially improves it by a clear margin of 3.6% on the test set.

4 Experiment 4: Performance on WSSS Track of LID2020{}_{20\!} Challenge

Experimental Setup: The challenge dataset is built upon ImageNet . It contains 349,319 images with image-level labels from 200 classes. Evaluations are conducted on the val and test sets, which have 4,690 and 10,000 images, respectively. In this challenge, our co-attention image classifier is built upon ResNet-38, as the dataset has 200 classes and a stronger backbone can better learn subtle semantics between classes. The training parameters are set as: initial learning rate (0.005) and the poly policy based training schedule: lr ⁣= ⁣lrinit ⁣× ⁣(1 ⁣− ⁣itermax_iter)γlr\!=\!lr_{init}\!\times\!(1\!-\!\frac{iter}{max\_iter})^{\gamma} with γ\gamma(0.9), batch size (8), weight decay (0.0005), and max epoch (15). During training, the equivariant attention is also adopted. Once our image classifier is trained, we run the classifier and directly use its class-aware activation map (i.e., Sn\bm{S}_{n}) as the object localization map Ln\bm{L}_{n}. Then we generate pseudo pixel-wise labels for all the training samples I\mathcal{I}. Since only image tags can be used, we follow​ : localization maps are first used to train an AffinityNet model, which is then used to generate pseudo ground truth masks and background threshold is set as 0.2. For better segmentation results, we choose ResNet-101 based DeepLab-V3. The parameters are set as below: initial learning rate (0.007) with poly schedule, batch size (48), max epoch (100), and weight decay (0.0001). The segmentation model is trained on 4 Tesla V100 GPUs. During testing, results from multiple scales are averaged, with CRF refinement.

Experimental Results: The final results with the standard mean intersection over union (mIoU) criterion for WSSS track of both LID19 ⁣{}_{19\!} and LID20 ⁣{}_{20\!} challenges are shown in Table​ 3. Both LID19 ⁣{}_{19\!} and LID20 ⁣{}_{20\!} challenge use the same data. In LID19, competitors can use extra saliency annotations to learn saliency models and refine pseudo ground truths. However, in LID20, only image tags can be accessed. For methods shown in the table, top performing methods are included. As can be seen from Table 3, our approach not only outperforms the champion team in LID19, which can use deep learning based saliency models, but also achieves the best performance in LID20 ⁣{}_{20\!} and sets a new state-of-the-art (i.e., mIoU of 46.2 and 45.1 in val and test sets, respectively).

5 Ablation Studies

Inference Strategies. Table 2 shows mIoU scores on PASCAL VOC 2012 val set w.r.t. different inference modes (see §3.2). When using the traditional inference mode “single-round feed-forward”, our method substantially suppresses basic classifier, by improving mIoU score from 61.7 to 64.7. This evidences that co-attention mechanism (trained in an end-to-end manner) in our classifier improves the underlying feature representations and more object regions are identified by the network. We can observe that by using more images to generate localization maps, our method obtains consistent improvement from “Test image only” (64.7), to “Test images and other related images” (66.2). This is because more semantic context are exploited during localization map inference. In addition, using contrastive co-attention for localization map inference doesn’t boost performance (66.2). This is because the contrastive co-attentive features for one image are derived from the image itself. In contrast, co-attentive features are from the other related image, thus can be effective in the inference stage.

(Contrastive) Co-Attention. As seen in Table 4, by only using co-attention (Eq. 5), we already largely suppress the basic classifier (Eq. 1) by 3.8%. When adding additional contrastive co-attention (Eq. 9), we obtain mIoU improvement of 0.7%. Above analysis verify our two co-attentions indeed boost performance.

Number of Related Images for Localization Map Inference. For localization map generation, we use 3 extra related images (§3.2). Here, we study how the number of reference images affect the performance. From Table 5, it is easily observed that when increasing the number of related images from 0 to 3, the performance gets boosted consistently. However, when further using more images, the performance degrades. This can be attributed to the trade-off between useful semantic information and noise brought by related images. From 0 to 3 reference images, more semantic information is used and more integral regions for objects are mined. When further using more related images, useful information reaches its bottleneck and noise, caused by imperfect localization of the classifier, takes over, decreasing performance.

Conclusion

This work proposes a co-attention classification network to discover integral object regions by addressing cross-image semantics. With this regard, a co-attention is exploited to mine the common semantics within paired samples, while a contrastive co-attention is utilized to focus on the exclusive and unshared ones for capturing complimentary supervision cues. Additionally, by leveraging extra context from other related images, the co-attention boosts localization map inference. Further, by exploiting additional single-label images and web images, our approach is proven to generalize well under domain gap and data noise. Experiments over three WSSS settings consistently show promising results. Our method also ranked 1st ⁣{}^{st\!} place in the weakly-supervised semantic segmentation track of LID20 ⁣{}_{20\!} challenge.

References