ESL: Entropy-guided Self-supervised Learning for Domain Adaptation in Semantic Segmentation

Antoine Saporta, Tuan-Hung Vu, Matthieu Cord, Patrick Pérez

Introduction

Intelligent systems like autonomous cars often require in-depth understanding of scenes in which they operate. To this end, most frameworks incorporate semantic segmentation modules to obtain class-label predictions for all input scene pixels. While recent advances in deep convolutional neural networks (CNNs) have significantly boosted segmentation performance, state-of-the-art results are only achieved with full-supervision. While full supervision already enables perception in functional self-driving vehicles, the annotation cost among others limits the operational domains of such systems without additional techniques to make the learning approach more robust and scalable. To mitigate the indispensable need of manual pixel-level annotations required by the task, recent works are trying to leverage cheaper alternative supervision sources such as synthetic datasets. However, models solely trained on synthesized images hardly reach comparable performance on real test cases as ones trained on real images. Such negative effects are originated from the so-called “domain-gap” between synthetic and real data, namely source and target. To alleviate the performance drop caused by train-test distribution discrepancy, most recent works resort to unsupervised domain adaptation (UDA) techniques. In the UDA context, along with the annotated source samples, some unlabeled target images are available at train time. Most UDA works aim at learning domain-invariant representations such that input features to the final classifiers demonstrate insignificant discrepancy across domains. For that purpose, recent works advocate adversarial training as an effective technique. Orthogonal to adversarial approaches are works motivated by entropy minimization .

Pseudo-labeling is a simple yet effective strategy used in semi-supervised learning and recently adopted for domain adaptation . For UDA, the main idea is to accept high-confident pseudo-labels as if they were true labels of target images at train time. To that purpose, classes with maximum softmax score are usually selected. Such a strategy shares the same underlying “cluster assumption” as in entropy minimization methods , i.e., classification decision boundaries should be driven toward low-density regions in the target space. One major concern of training with pseudo-labels is the lack of guarantee for label correctness, which may eventually cause a “confirmation bias” , namely models are biased to previous incorrect pseudo-labels and resist new changes. To address this problem, one can use a curriculum approach to gradually relax the pseudo-labeling constraint from easy-to-hard . Self-supervised learning (SSL) done in an iterative manner has also been proven effective in improving pseudo-label quality over time.

In this work, we explore the use of entropy, instead of maximum softmax score, as a more reliable criterion for constructing a high-quality set of pseudo-labels. We argue that maximum softmax score is more error-prone than entropy in cases where the model is highly uncertain. For example, the right-most image in Fig. 1 visualizes in color the set of pixels that are selected as pseudo-labels based on softmax threshold but do not satisfy the entropy criterion. Most of these selected pixels lie on the boundaries between object classes, regions where segmentation models are naturally less confident. By excluding such likely troublesome pixels, we expect to obtain higher-quality pseudo-labels, which indeed mitigates the adverse confirmation bias. To validate such a hypothesis, starting from the previous state-of-the-art SSL work for UDA, we replace the standard softmax-based pseudo-labeling with our entropy-based strategy. We coin our framework Entropy-guided Self-supervised Learning (ESL). On different benchmarks for UDA in semantic segmentation, the proposed ESL scheme shows consistent improvement over strong SSL baselines and achieves state-of-the-art performance. Extensive experiments and ablation studies bring more insights into the proposed framework.

Related Work

In last few years, UDA has become an active line of research . Most works share the similar objective of minimizing distribution discrepancy between source and target domains. To this end, one can regularize the maximum mean discrepancy (MMD) , match correlations of layer activations or adopt adversarial training to align source-target distributions either at feature-level or on output space . In some set-ups where extra supervision on the source domain is available at train time, e.g. depth, recent UDA frameworks leverage such privileged information to further enhance the domain-invariant representations. For UDA in semantic segmentation, style-transfer-based methods can be used to translate source images into target-like ones. The translated source images, with different visual styles yet preserved semantic labels, indeed bring adaptation effects for models trained on them . In this work, we base our iterative self-supervised framework on top of the state-of-the-art adversarial training methods, namely AdaptSegNet and AdvEnt .

Recent UDA approaches adopt the entropy minimization principle in the attempt to reinforce prediction confidence on the target domain. Imposing the minimum entropy constraint on target predictions is indeed one way to implement the cluster assumption, namely to prevent decision boundaries from crossing high-density regions. The pseudo-labeling strategy in is in the same spirit. For UDA in semantic segmentation, Vu et. al leverage such a principle by either directly penalizing high-entropy target predictions or, in a more implicit way, aligning source-target distributions on the weighted self-information space with adversarial training. Different to previous works, we here use entropy as the confidence measurement to construct high-quality sets of pseudo-labels.

Another efficient UDA approach is self-supervised learning, or self-training . Such frameworks iterate training upon target pseudo-labels extracted from the preceding step. For UDA in semantic segmentation, Li et al. , adopting a simple pseudo-labeling strategy, experimentally demonstrate that the SSL training often converges after two iterations. With the same methodology, the class-balanced self-training (CBST) solves optimization problems to determine class-specific pseudo-label thresholds while making good use of class-proportion priors. In our work, we opt to improve over , not only because this work reports state-of-the-art performance, but also due its superior simplicity.

ESL: Entropy-guided Self-supervised Learning

In this Section, we provide details of the proposed ESL framework. Similar to previous self-supervised works, our learning strategy operates in an iterative manner, in which each iteration involves performing a domain alignment step. Section 3.1 formulates the task and describes alignment techniques that will be used in the experiments. We then introduce our self-supervised algorithm in Section 3.2, followed by details of the entropy-guided pseudo-labeling strategy in Section 3.3.

is minimized over the parameters θF\theta_{F} of FF.

In adversarial learning, a discriminator DD – usually, a small fully-convolutional network – with parameters θD\theta_{D} is appended at some layer of FF and produces domain classification outputs: output 11 for the source domain and output for the target domain. By denoting Ladv\mathcal{L}_{\text{adv}} the cross-entropy loss of this discriminator, the minimization objective of DD over θD\theta_{D} is:

At the opposite, the semantic segmentation network FF must be trained to fool this discriminator DD. Thus, an adversarial loss is added to the segmentation objective and the training objective of the semantic segmentation network FF we minimize over θF\theta_{F} is:

with a weight λadv\lambda_{\text{adv}} for the adversarial term. When training the network, we alternately optimize DD and FF using the previously defined objective functions LD\mathcal{L}_{\text{D}} and LF\mathcal{L}_{\text{F}}, respectively.

2 Self-supervised learning

On top of the regular domain alignment process, additional self-supervision on the target domain can be added to the objective of the segmentation network FF using pseudo-labels on Xt\mathcal{X}_{t}. By noting (xt,y^t)(\bm{x}_{t},\hat{\bm{y}}_{t}) the pairs of target images and pseudo-labels, the objective function of FF with self-supervised learning can be written:

with a weight λSL\lambda_{\text{SL}} for the self-supervised learning term.

Considering we can already train a segmentation network FF by unsupervised domain adaptation without self-supervision, such a semantic segmentation network would give “good” predictions in the target domain. Those predictions can serve to collect pseudo-labels for further training. Under this assumption, the self-supervised learning process could be described as follow:

Train a segmentation network without self-supervision;

Extract pseudo-labels on the target training set using the predictions from the trained network;

Re-train a segmentation network from scratch with additional supervision from the extracted pseudo-labels.

This process could be repeated multiple times, extracting finer pseudo-labeling after each iteration. In this work, we will focus on a single iteration of this algorithm. We discuss in what follows ways to extract pseudo-labels on the target training set using a trained network.

3 Entropy-guided Pseudo-label Extraction

A standard strategy for pseudo-label extraction is to assign classes having maximum prediction softmax score as pseudo-labels. Enforcing this maximum softmax score to be greater than a threshold is a common practice to filter poor predictions from the constructed pseudo-labels. This strategy has already been used in previous state-of-the-art paper in unsupervised domain adaptation for semantic segmentation . Formally, the pseudo-labels extracted using the trained segmentation network FF can be written:

where μ(c)\mu^{(c)} is a threshold over the softmax prediction score for class cc. Pixels with maximum class score below the relevant threshold are assigned a null pseudo-label vector y^t(h,w)=0\hat{\bm{y}}_{t}^{(h,w)}=\bm{0} (not one-hot then). This assignment effectively excludes such pixels from the segmentation loss Lseg(xt,y^t)\mathcal{L}_{\text{seg}}(\bm{x}_{t},\bm{\hat{y}}_{t}) according to its definition in (1). A simple way to define this threshold, as in , would be:

where μ∗\mu^{*} is an hyper-parameter. Such a threshold ensures that we keep at least 50% of the predictions for each class based on their softmax prediction score and that we keep all the predicted labels with a maximum softmax prediction score greater than μ∗\mu^{*} on the easier classes (on which the prediction scores are rather high over the training set).

In the standard pseudo-label extraction strategy previously described, the maximum softmax prediction score is used as a confidence score for the prediction. We argue that the maximum softmax prediction score is not the best measure of confidence of the network. Indeed, while the maximum softmax prediction score may be greater than the given threshold, this measure does not take into account the softmax prediction score over other classes, overlooking potentially high softmax prediction score on those other classes that would question the confidence of the model. Such a behavior is illustrated in Figure 1. As a consequence, we propose an entropy-guided pseudo-label extraction strategy that uses the entropy of the softmax prediction as a measure of confidence. Unlike the maximum softmax prediction score, the entropy actually takes into account the full distribution of the softmax prediction score for each pixel, making this measure more reliable in assessing the confidence of the network. The extracted pseudo-labels can be written as follow:

where the entropy Ext(h,w)\bm{E}_{\bm{x}_{t}}^{(h,w)} is defined as:

and ν(c)\nu^{(c)} is a threshold over the entropy score for pixels of class cc. Similarly to the threshold of the previous method, we can define ν(c)\nu^{(c)} as:

where ν∗\nu^{*} is a hyper-parameter. This threshold ensures we keep at least the 50% most confident predictions in term of entropy for each class for our pseudo-labels. Moreover, we keep all the predictions with an entropy score lower than the hyper-parameter ν∗\nu^{*} for the easier to predict classes on which the segmentation network is more confident.

Experiments

In this section, we present experimental results on models trained with self-supervised learning techniques on various semantic segmentation domain adaptation datasets. We compare them to different baselines, showing that self-supervised learning helps boosting the performance and that models trained with entropy-guided self-supervised learning consistently outperform baselines and models with standard self-supervised learning.

In this work, we consider two synthetic source datasets – SYNTHIA and GTA5 – and two real target datasets – Cityscapes and Mapillary Vistas . For SYNTHIA , we use the SYNTHIA-RAND-CITYSCAPES split composed of 9,400 synthetic images of size 1280 ×\times 760 annotated with pixel-wise semantic labels over 16 classes common with Cityscapes . GTA5 is composed of 24,966 synthetic images of size 1914 ×\times 1052 annotated with 19 classes common with Cityscapes . For the target datasets, Cityscapes is a dataset of street-level images split in a training set, a validation set and a testing set. We exclusively use the training set as the target set for domain adaptation. It contains 2,975 images of size 2048 ×\times 1024. We use the 500-image validation set for testing since ground-truth segmentation maps are missing from the testing dataset. As Cityscapes , Mapillary Vistas is a dataset of street-level images split in a training set, a validation set and a testing set. We only use the 18,000 images training set as target set for domain adaptation. The 2,000-image validation set is used for testing because, as for Cityscapes , the testing set is missing ground-truth segmentation maps.

Architectures and baseline frameworks

We experiment over three state-of-the-art domain adaptation baselines. The three of them are based on DeepLab-V2 for the semantic segmentation module of the architecture but the domain adaptation frameworks of these baselines are different:

AdaptSegNet considers semantic segmentations as structured outputs that contain spatial similarities between source and target domains. For this reason, they adopt an adversarial learning framework on the output space. Moreover, they construct a multi-level adversarial network to perform adaptation at different feature levels.

ADVENT , alternatively, adopts an adversarial learning framework on the entropy of the pixel-wise predictions instead of the raw softmax output predictions as in AdaptSegNet .

BDL , as AdaptSegNet , conducts adaptation on the output space of the semantic segmentation module. The method adds two main components to the framework: first, an image translation module based on CycleGAN to transfer the style of the target domain to source domain images ; second, a bidirectional learning training procedure in which this image translation module and the semantic segmentation module are trained alternately and contribute to each other’s performance. Furthermore, this strategy already incorporates self training using standard pseudo-label extraction. In our experiments, we will focus on two given steps of the sequential model training on GTA5 →\rightarrow Cityscapes, called ‘Step 1’ and ‘Step 2’, which can be found on the authors’ GitHub https://github.com/liyunsheng13/BDL. We apply self-training strategies on those two pretrained models, possibly adding image translation (IT).

Implementation details

Implementations are done with the PyTorch deep learning framework . Training and validation of the models are done on a single NVIDIA 1080TI GPU with 11GB memory. The semantic segmentation models are initialized with the ResNet-101 pre-trained on ImageNet . The semantic segmentation models are trained by a Stochastic Gradient Descent optimizer with learning rate 2.5×10−42.5\times 10^{-4}, momentum 0.90.9 and weight decay of 10−410^{-4}. The discriminators are trained by an Adam optimizer with learning rate 10−410^{-4}. We fix λadv\lambda_{\text{adv}} as 10−310^{-3} and λSL\lambda_{\text{SL}} as 1. Following the recommendation from , we use a value of 0.90.9 for μ∗\mu^{*} in the SSL experiments. Moreover, we adopt a value of 0.10.1 for ν∗\nu^{*} in the ESL experiments and justify this choice in Section 4.3.

2 Results

We report in Table 1 semantic segmentation performance in terms of mIoU (%) on the Cityscapes validation set using GTA5 as source domain. As explained in Section 4.1, ‘BDL (step 1)’ and ‘BDL (step 2)’ represent the two pretrained models which can be found on the authors’ GitHub. We can notice that ESL consistently outperforms SSL on every setup, giving the best performance for each baseline framework. The performance absolute change in terms of mIoU using ESL compared to the baseline state-of-the-art models ranges from +1.0%+1.0\% to +2.2%+2.2\% which is a significant improvement. Along with the quantitative results, Figure 2 displays some samples of pseudo-labels extracted with both SSL and ESL. We first note that these maps are very similar within semantic regions. Indeed, in these regions the softmax prediction is often very peaky with a high maximum score and a low entropy. Nevertheless, marked differences can be observed along the boundaries of these regions. These transition areas are those where the prediction models are the most uncertain. This clearly shows in the fact that most missing pixels in both pseudo-label maps lie on such locations. Around the boundaries however, the prediction entropy tends to get higher even if the softmax prediction score stays high. This results in pixels being rightly excluded from the ESL pseudo-label map while present in the SSL one (column (d) of Figure 2). Reversely, there are much fewer pixels that are added to ESL compared to SSL.

SYNTHIA →→\rightarrow Cityscapes

We report in Table 2 semantic segmentation performance in terms of mIoU (%) on the Cityscapes validation set using SYNTHIA as source domain. Again, ESL consistently outperforms SSL on every setup, even when SSL fails to improve the performance over the baseline.

Mapillary Vistas

We report in Table 3 semantic segmentation performance in terms of mIoU (%) on the Mapillary Vistas validation set using SYNTHIA as source domains and ADVENT as baseline model. Proposed ESL outperforms SSL and improves over the baseline performance.

3 Ablation Studies

Let us describe how to choose the threshold ν∗\nu^{*} such that a good balance is achieved between having as many high confidence predicted labels as possible and avoiding as much as possible noise from incorrect predictions. A quick computation suggests 0.10.1 as a good threshold value. Indeed, considering the maximum softmax prediction score for a given pixel is 0.950.95 on a 19-classes setup, the entropy of the distribution would range from 0.070.07 in the best case to 0.120.12 in the worst case. We confirm this choice experimentally. We show in Table 5 segmentation results on the domain adaptation problem GTA5 →\rightarrow Cityscapes with the ADVENT baseline using different thresholds in the ESL label extraction. Alternately, we show the limit case where the threshold is always selected as the median of the entropy for each class. The result of this experiment is shown on the last row of Table 5. When the threshold is higher than 0.10.1, the incorrect predictions degrade the quality of the pseudo-label maps and induce more noise in the training. When the threshold is lower than 0.10.1, we don’t keep as many confident predictions, slightly reducing the effectiveness of the pseudo-labeling. This experiment confirms that 0.10.1 seems to be a good threshold for ESL.

Incorrect predictions in the pseudo-labels

In order to further motivate the choice of entropy as a confidence measure in the proposed ESL method over the softmax prediction score of SSL, we report in Table 4 the ratio of incorrect predictions (in %\%) in the selected pseudo-labels for every class for both SSL and ESL on the GTA5 →\rightarrow Cityscapes experiment with the ADVENT baseline (the lower the score, the better). Additionally, the last row of the Table displays the relative change in the ratio of incorrect predictions for every class. The results show that ESL performs significantly better than SSL on the easier-to-predict classes (less than 10%10\% incorrect predictions in the selected pseudo-labels). Indeed, ESL reduces the number of incorrect predictions for those classes by a significant margin, ranging from 4.2%4.2\% up to 18.6%18.6\%. This means that ESL induces significantly less noise in the training for those classes. These changes can be explained by the “overconfidence” the model may have in terms of softmax prediction score for easier-to-predict classes, which is not as significant for the entropy as measure of confidence. Overall, ESL decreases the ratio of incorrect predictions for 14 classes out of 19 and globally reduces the ratio of incorrect predictions by 3.7%3.7\% over SSL.

Conclusion

In this paper, we propose an Entropy-guided Self-supervised Learning strategy for semantic segmentation adaptation problems. We demonstrate with a variety of experiments that this method consistently improves the performance of state-of-the-art domain adaptation models on synthetic-to-real problems and outperforms standard softmax-based self-supervised learning approaches.

References