DecoupleNet: Decoupled Network for Domain Adaptive Semantic Segmentation
Xin Lai, Zhuotao Tian, Xiaogang Xu, Yingcong Chen, Shu Liu, Hengshuang Zhao, Liwei Wang, Jiaya Jia
Introduction
Semantic segmentation has made tremendous progress in recent years and it has significantly benefited plenty of applications. However, satisfying performance highly relies on pixel-wise annotations. In this work, to alleviate the data-reliance issue, we focus on unsupervised domain adaptation (UDA), aiming to learn a segmentation network with a labeled source domain dataset (usually a physically synthetic dataset) and an unlabeled target domain dataset.
Due to “domain shift” between the source and target domains, directly adopting the model trained on the source domain causes much performance degradation on the target one. To minimize domain shift, domain-invariant learning was put forward to align distributions of source and target features. Specifically, the features or predictions from different domains are aligned with a discriminator in the manner of adversarial learning, as shown in Fig. 1(a). The discriminator learns to distinguish between source and target features, while the segmentation network learns to generate features that can fool the discriminator.
Domain-invariant learning alleviates domain shift. However, we still observe the following two problems.
(1) Tasks entanglement. The feature distribution alignment and the segmentation task are conducted simultaneously in a single network, as shown in Fig. 1(a). Being distracted by feature distribution alignment, the network cannot focus on semantic segmentation, leading to inferior performance.
(2) Source domain overfitting. Since the training objective involves cross-entropy loss that minimizes the error on the source domain data, the trained model would fit the source domain data well, as shown in Fig. 1(c). However, in UDA, we only care about the performance on the target domain, regardless of how it performs on the source domain. Moreover, as we will discuss in Sec. 3.2, fitting the source domain very well would contrarily compromise the target domain performance.
Based on these two observations, we design DecoupleNet to decouple feature distribution alignment and the segmentation task. As shown in Fig. 1(b), we introduce a copy of shallow encoder layers for the source domain, i.e., , during training. Our goal is to let conduct feature distribution alignment, such that the final model focuses more on the downstream segmentation task. Also, it is notable that is simply discarded during inference, and it only incurs negligible computational costs during training as shown in the supplementary material.
With our new design, the issue of tasks entanglement can be addressed, as shown in Fig. 1(b). Moreover, during training, we only require the model to fit well on the source domain, but never ask the final model to do so. Thus, the final model avoids overfitting in the source domain. As shown in Fig. 1(c), compared to the domain-invariant method (AdaptSegNet ), DecoupleNet alleviates the source domain overfitting problem, and boosts the target domain performance.
In addition, in order to learn more discriminative features for the target domain, we propose the Self-Discrimination (SD) technique by virtue of pseudo labels. Unlike most self-training-based methods , SD does not need another training phase to re-train the whole network from scratch. Instead, pseudo labels are generated at each training iteration and can be employed as an additional supervision in an online manner. Given the fact that directly adopting the noisy pseudo labels to supervise itself could corrupt the existing classifier, we introduce an auxiliary classifier during training to prevent the contamination.
Finally, we propose Online Enhanced Self-Training (OEST) to further boost the performance by extending DecoupleNet to a multi-stage training paradigm. Most existing self-training-based methods directly use the generated pseudo labels without updating them in the re-training process. Contrarily, at each training iteration, OEST updates the pseudo labels by fusing current contextually enhanced predictions, which effectively improves the quality of pseudo labels.
In summary, our contribution is threefold.
We propose DecoupleNet to decouple feature distribution alignment and semantic segmentation. This enables the network to get rid of tasks entanglement and focus more on the segmentation task.
To learn more discriminative features, we put forward Self-Discrimination by introducing an auxiliary classifier. Moreover, we propose Online Enhanced Self-Training to contextually enhance the quality of pseudo labels.
Experiments show that our approach outperforms existing state-of-the-art methods by a large margin. Also, extensive ablation studies verify the effectiveness of each component in our method.
Related Work
Semantic segmentation aims to assign a class label to evey pixel in an image. FCN is a classic semantic segmentation network, which puts forward a fully-convolutional network. Considering that the final output size of FCN is smaller than the input, methods based on encoder-decoder structures are proposed to refine the output. Though the high-level feature has already encoded the semantic information, it cannot well capture the long-range relationship. Dilated convolution , global pooling , pyramid pooling and attention mechanism are used to better incorporate the context. Despite the success, all the models need annotations to accomplish training, which costs much human effort.
0.2 Unsupervised domain adaptation.
Unsupervised dmain adaptation intends to alleviate the data-reliance with a labeled dataset from a different domain. Distance-based methods minimize the distribution distance such as MMD between the source and target domain. With the development of Generative Adversarial Network (GAN) , adversarial learning methods get popular to align the marginal or conditional feature distributions between the source and target domains. Also, methods of factorize the feature into domain-specific and domain-agnostic features.
0.3 UDA in Semantic Segmentation.
AdaptSegNet employs adversarial learning to align predictions between the source and target domain in the output space and method of makes further improvement. Patch-level information is used in to improve the performance and contextual relationship is considered in explicitly. directly minimize the feature distance. In , semi-supervised learning methods, such as entropy minimization, adding perturbation, contrastive learning and randomly dropout, further boost performance. align class-conditioned feature distribution. Methods of provide distinct processing for features from different domains on some modules. On the other hand, image-to-image translation methods are considered in . Recently, self-training-based methods re-train the network with the pseudo labels generated from the initial network, yielding considerable improvement.
Our Method
In this section, we first introduce the preliminary in Sec. 3.1. Then, the key observations are presented as our motivation in Sec. 3.2. Afterwards, DecoupleNet, SD and OEST are elaborated in Sec. 3.3, Sec. 3.4 and Sec. 3.5, respectively.
Formally, we have the source domain images along with ground-truth labels , and the unlabeled target domain images . Our goal is to train a segmentation model that performs well on the target domain.
A representative domain-invariant solution is shown in Fig. 2 (a). The source and target domain images pass forward the segmentation network , which is typically composed of an encoder and a classifier , to obtain the predictions , respectively. It is written as
For the source domain prediction , the cross-entropy loss is employed with its ground-truth label as
where is the number of spatial locations in the source prediction map , is the number of classes, represents the class label at the -th location, and represents the source prediction score of the -th class at the -th location.
As for the target domain prediction , a discriminator is used to align the distributions of the source and target predictions. The adversarial loss is defined as
where is the number of spatial locations in the discriminator output, is the label of the source domain, and we follow LSGAN to use the MSE Loss.
The final loss for the segmentation network is defined as
where controls the weight for . To train the discriminator, the discriminator loss is defined as
where the labels of source and target domain are and , respectively. Training alternates between updating the segmentation network with and the discriminator with .
2 Motivation
The method above aligns the distributions of source and target domain features for domain-invariant learning. However, as shown in Fig. 2(a), since the learning objective involves during training, the trained network has to fit the source domain data very well. The source domain overfitting issue potentially impairs the segmentation performance on the target domain.
We conduct three experiments to verify this fact, and show them in Fig. 2(b)-(d). Unlike the normal case (Fig. 2(a)), we use the target domain ground-truth labels in the toy experiments only to support our idea rather than give a solution. As shown in Fig. 2(e), from Exp. I to II, we apply an extra adversarial loss, so the model performs slightly better on the source domain data. Further, from Exp. II to III, we apply a stronger CE loss on the source domain, so it performs very well on the source domain. However, the results in Fig. 2(e) reveal the fact that the better the model fits on the source domain data, the worse it performs on the target domain. This exactly supports our idea, i.e., overfitting the source domain data actually impairs the final performance on the target domain.
Motivated by the observations, we propose a new framework to decouple the feature distribution alignment from the segmentation task. It alleviates the issue of source domain overfitting, and enables the final model to focus more on target-domain semantic segmentation.
3 DecoupleNet
The framework of DecoupleNet is shown in Fig. 3. We first split the feature encoder into two parts, i.e., and . Besides, we maintain another module , which shares the same architecture with . The source and target domain images are fed into the source blocks and target blocks to yield the shallow features , respectively. They further pass through the shared blocks to get the features . Afterwards, they are passed into the classifier to obtain the predictions . Initially, we have
Then, we adopt cross-entropy loss for the labeled source domain data as
Besides, we require the distribution of the source-domain shallow features to align towards that of the target domain, i.e., , since our goal is to let the source blocks bear the responsibility of feature distribution alignment. Specifically, adversarial learning is adopted for the shallow feature alignment with an additional discriminator and an adversarial loss as
where denotes the number of locations in the discriminator output, and is the label of the target domain.
The design of DecoupleNet is with the following consideration. Basically, the source domain images differ from the target ones mainly in low-level information, such as illumination and texture. Also, it is known that the shallow layers in a network often do well in capturing the low-level information. With these two facts, it is natural to let the source blocks align the source-domain shallow features towards the target ones.
Practically, the shallow feature distribution alignment by may be imperfect, and the shallow features for the source and target domains may still be slightly mismatched. Therefore, we also use the adversarial loss in the output space, as defined in Eq. (3). In this way, we have the final loss as follows for training the segmentation network.
where and control the contributions of the corresponding loss. It is notable that the incorporation of only brings minor improvement (+0.3% mIoU), as shown in Exp. 5 and 6 of Table 3. This shows the feature alignment is mainly attributed to . The only serves as a complement.
To train the discriminators, as shown in Fig. 3(b), we follow the previous work to yield the discriminator loss as
During inference, as shown in Fig. 3(c), we adopt as the final model. All other modules are simply discarded. Note that we do not introduce extra parameters during inference.
(1) The source blocks now bear the responsibility of feature distribution alignment. Being less distracted by feature alignment, the final model (i.e., ) focuses more on the segmentation task. (2) Though the source domain branch needs to directly fit the source domain data with , the final model is never required to perform well on the source domain during training. This alleviates the source domain overfitting problem and facilitates performance enhancement on the target domain, as shown in Fig. 1(c).
4 Self-Discrimination
Despite the effectiveness of DecoupleNet, the target blocks is updated only according to , which may not be strong enough to learn optimal parameters for . Also, without a proper learning objective for the target domain, the features may not be discriminative enough. To this end, we propose Self-Discrimination (SD) to provide more supervision on the target domain branch.
It is notable that the class-wise thresholds are initialized to zero at the start of training. Then, it is updated with the current predictions at each iteration. The implementation details are given in the supplementary material.
Basically, is a cross-entropy loss applied to . It has a nice property that can adaptively scale the gradients with the current prediction error. Hence, it is capable of yielding more discriminative target features . To verify the effectiveness, we compare the t-SNE visualizations with and without SD in the supplementary material. During inference, we only use the main classifier, and the auxiliary classifier is simply discarded.
Remarkably, the accuracy of the pseudo labels is more than 80%, and continues to increase during training, as shown in Fig. 4. Therefore, although there might be wrong supervision from pseudo labels, the benefits brought by SD still outweigh the risks.
Finally, we incorporate into the final segmentation loss as
The auxiliary classifier plays an important role in SD. If we directly apply the self-discrimination loss on the main prediction without the auxiliary classifier, the noisy pseudo labels may corrupt the normal training of the main classifier with and cause large performance degradation, as shown in Exp. 1 and 2 of Table 6. In contrast, introducing an auxiliary classifier avoids the side effect on the main classifier.
5 Online Enhanced Self-Training
To further boost performance, we extend DecoupleNet from a single stage to a multi-stage self-training paradigm. Most existing self-training-based methods generate pseudo labels in the re-labeling phase and directly use them to provide supervision without further update in the re-training phase. Generally, the predictions get more and more accurate in the re-training process, so fixing the generated pseudo labels may lead to inferior performance. ProDA uses prototypes to denoise the pseudo labels. But it requires to maintain an extra momentum encoder and needs to update the prototypes at each iteration. In contrast, we propose a simple yet effective method, i.e., Online Enhanced Self-Training (OEST), to contextually enhance the pseudo labels via a simple average operation at each iteration.
The framework of OEST is given in Fig. 5. After the first-stage training explained in Sections 3.3 and 3.4, we generate the pseudo soft labels by making predictions on each target domain training image using the trained model. Then, in the re-training process, we pass the target domain image crops with strong data augmentation (e.g., color jitter) into the segmentation network to yield their predictions . In addition, we forward their corresponding full images with weak data augmentation (e.g., random horizontal flip) to obtain the full predictions as
Experiment
Following previous work , we use the ResNet-101 and DeepLabv2 as our base model. To split the feature encoder, we take {layer, layer} as the target blocks and the rest as the shared blocks for GTA5 dataset, while {layer, layer, layer} as and the rest as for Synthia dataset. Note that layer refers to {conv1, bn1, relu, maxpool}. More details are given in the supplementary material.
1.2 Datasets.
Following most previous work, evaluation is performed on GTA5 Cityscapes, Synthia Cityscapes and Cityscapes Cross-City. The details of the datasets are given in the supplementary material.
2 Results
The comparison with existing state-of-the-art methods is given in Table 1, Table 2. Clearly, our method outperforms others by a large margin. Previous methods neglect the adverse effect brought by entanglement of feature distribution alignment and the segmentation task. Contrarily, DecoupleNet decouples these two tasks, hence boosting the performance.
What’s more, equipped with OEST, our method demonstrates stronger performance. Notably, our method even surpasses ProDA by 3.1 points on GTA5Cityscapes and 2.8 points on SynthiaCityscapes, achieving a new state of the art. Furthermore, following ProDA to distill the SimCLR initialized student, our method still outperforms ProDA on both benchmarks. On Cityscapes Cross-City, our method also manifests competitive results given in the supplementary material.
3 Ablation Study
By comparing Exp. 2 and 4 in Table 3, we can see DecoupleNet outperforms the domain-invariant method (AdaptSegNet ) by 3.0% mIoU, which reveals the effectiveness of DecoupleNet. Note that except the decoupled architecture and , Exp. 2 and 4 are kept all the same for fair comparison.
Notably, we emphasize that brings only slight improvement (% mIoU) by comparing Exp. 6 and 7 in Table 3. On the other hand, brings large performance boost (% mIoU), through the comparison between Exp. 5 and 7. This shows that the huge performance boost by DecoupleNet mainly comes from the decoupled network architecture and , rather than . only serves as a complement for the imperfect alignment by . This demonstrates the effectiveness of DecoupleNet from another perspective.
Besides, we investigate the effect of decoupled layers (i.e., the architecture of or ) in Table 4. Making it too shallow leads to insufficient capability for feature alignment, while making it too deep may interfere segmentation.
Also, we highlight the importance of the alignment direction of and in Table 5. performs the best. We explain that this prevents the segmentation network from being distracted by the feature alignment.
3.2 Self-Discrimination.
By comparing Exp. 4 and 7 in Table 3, we observe a performance boost of % mIoU brought by SD, which clearly demonstrates its effectiveness. Also, it is notable that when we directly apply SD on the domain-invaraint method (i.e., AdaptSegNet ), the performance still continues to improve by a large margin, through the comparison between Exp. 2 and 3 in Table 3. It shows that SD is not limited to DecoupleNet and can serve as a plugin to existing methods by providing an additional supervision.
In addition, we show the t-SNE visualizations of the target domain features with and without SD in the supplementary material. It reveals the fact that the model tends to learn more discriminative target domain features with SD.
Moreover, to show the necessity of the auxiliary classifier, we make comparison in Table 6. For the model w/o auxiliary classifier (Exp. 2), we directly apply on the main predictions , which leads to large degradation (% mIoU) compared to Exp. 1. We conjecture that the supervision signal from the noisy pseudo labels may interfere the normal training of the main classifier with source domain ground-truth labels. Further, Exp. 1 and 3 in Table 6 show the effectiveness of the class-wise thresholds, since it alleviates the class-imbalance issue on pseudo labels.
3.3 Online Enhanced Self-Training.
As shown in Table 7, we compare the models with various fusion methods. The comparison between ‘avg (full)’ and ‘avg (crop)’ show the effectiveness of contextual enhancement via full predictions. Moreover, ‘fix’ is inferior to ‘avg (full)’ by 1.1% mIoU, which shows that online updating pseudo labels with current predictions indeed improves the quality of pseudo labels and brings performance boost. As for ‘pred only’, it totally corrupts the training potentially due to the instability of the online prediction.
Conclusion
We have observed two issues of existing domain-invariant learning methods – (1) tasks entanglement and (2) source domain overfitting. We propose DecoupleNet to enable the final model to focus more on the segmentation task. Moreover, Self-Discrimination is put forward to learn more discriminative target features. Finally, we design OEST to contextually enhance the pseudo labels.