Source-Free Domain Adaptation for Semantic Segmentation
Yuang Liu, Wei Zhang, Jun Wang
Introduction
Semantic segmentation has been a critical computer vision task, which aims to segment and parse a scene image into different image regions associated with semantic categories. It is critical for precisely understanding the visual scene and can be applied to numerous potential applications, such as autonomous driving , visual grounding , and image editing . But the success of current segmentation techniques depends on large-scale densely-labeled datasets that are prohibitively expensive to be collected in reality. For instance, it takes about 90 minutes to manually annotate a Cityscapes image. An intuitive method to address this issue is transferring knowledge from existing models trained on source datasets to the unlabeled target domain. However, it tends to be hindered by the issue of domain shift which is caused by various data distributions in source and target domains.
Unsupervised domain adaptation (UDA) for semantic segmentation has been proposed to address this issue and generalize the well-trained models on an unlabeled target domain, avoiding expensive data annotation. All the methods suppose that both the well-trained source models and labeled source datasets are available. This is because source data plays a vital role in retaining valuable source knowledge during adaptation training and reducing the cross-domain discrepancy iteratively. However, in some crucial areas like autonomous driving, the source datasets may be private and commercial, making only the source models and unlabeled target datasets available. Due to the lack of supervision of the source domain and the uncertainty of target pseudo-labels, none of these UDA methods can work in such source-free scenarios.
With these insights, we formulate a new but important problem — source-free domain adaptation for semantic segmentation, in which only a well-trained source model and an unlabeled target domain dataset are available for adaptation. Recently, a tiny number of source-free UDA methods have been developed to tackle a similar issue on image classification. However, the image-level computer vision task just associates the label with a whole image, which is fundamentally different from image segmentation that belongs to a pixel-level task with each pixel associated with a semantic label. As shown in Figure 1, the pseudo-labels of one target image contains multiple classes shifting on diverse distributions. As such, it is nontrivial for the above methods to leverage clustering for each class adaptation. Since considering that the source domain knowledge cannot be preserved and utilized without source data, so we attempt to recover and transfer the source domain knowledge by introduced data-free knowledge distillation approaches that are originally for model compression.
In this work, we propose a novel source-free unsupervised domain adaptation framework for segmentation, namely SFDA. Our framework alternatively works in two stages: knowledge transfer and model adaptation. Due to unavailable source data and uncertain target pseudo-labels, recovering and preserving the source knowledge learned by a source model is vital during adaptation training. This is because the uncertain supervision information in target pseudo-labels will tend to deviate the target model from the working domain. As such, in the knowledge transfer stage, we leverage a generator to estimate the source domain (working domain) and synthesize fake samples similar to the real source data in distribution, which can be used to transfer the domain knowledge from a well-trained source model to a target model. The key to semantic segmentation networks lies in capturing contextual feature relationships. With this intuition, a dual attention distillation (DAD) mechanism is introduced to help the generator synthesize samples with meaningful semantic context, which is beneficial to efficient pixel-level domain knowledge transfer. Moreover, the source model could work well on partial target domain and predict correct labels. Therefore we propose an entropy-based intra-domain patch-level self-supervision module (IPSM) to leverage the correctly segmented patches as self-supervision during the model adaptation stage.
Our main contributions can be summarized as follows:
We propose the novel SFDA framework that combines knowledge transfer and model adaptation without requiring any source data and target labels. To our best knowledge, this is the first attempt to address the problem of source-free UDA for semantic segmentation.
A novel dual attention distillation mechanism is designed specifically for segmentation to transfer and retain the contextual information, and the intra-domain patch-level self-supervision module is introduced to exploit patch-level knowledge in target domain.
We demonstrate the effectiveness of our framework on synthetic-to-real and cross-city segmentation scenarios. In particular, it can even achieve competitive results with the state-of-the-art source-driven UDA approaches under the source-free setting.
Related Work
UDA for Semantic Segmentation. Existing UDA methods for segmentation can be mainly divided into three categories. To reduce the cross-domain discrepancy, numerous UDA methods focus on distribution consistency by introducing adversarial learning. Inspired by image-to-image translation , a category of UDA methods has been proposed to generate target images conditioned on source data . In addition, self-supervision with target pseudo-labels is a relatively simple but efficient approach , but it requires source data for supervision. In summary, all the above UDA methods for segmentation assume that the densely-annotated source dataset is available during adaptation, ignoring the data privacy and inaccessibility issues in practice. To the best of our knowledge, we are the first to consider the source-free unsupervised domain adaptation issue for image segmentation.
Knowledge Distillation (KD). Knowledge distillation is originally developed to transfer knowledge from a large teacher network to a compact student network . Since then, a variety of KD methods has been presented for model compression , domain adaptation , and multi-modal learning . More recently, data-free knowledge distillation has drawn surging attention, due to the inevitable data privacy issue. In , activation records are used to reconstruct training samples for training a compact student model. Analogously, Batch Normalization Statistics (BNS) stored in Batch Normalization (BN) layers can be used to reconstruct training samples as well. Most of the data-free KD methods based on generative adversarial networks . They all focus on generating fake samples for transferring knowledge from teacher to student networks without original training data mainly on classification tasks. In this work, we extend the data-free knowledge distillation methods to segmentation and tackle the source-free domain adaptation challenge.
Methodology
where is the supervised training loss for preserving source domain knowledge, usually cross-entropy or focal loss. And is the self-supervision loss for the target domain based on pseudo-labels, such as entropy minimization , maximum square loss (MaxSquare) , \etc. In this work, we adopt the maximum square loss as an assistance during adaptation, which is defined as:
where is the probability of category for one target image pixel and is the number of semantic categories.
In source-free scenarios, the annotated source dataset is unavailable, so the supervised learning process to preserving source knowledge will abort. Fortunately, the source domain knowledge has been permanently retained in the source model. We can consider source-free UDA as a knowledge transfer and adaptation problem, shown in Figure 2. The orange or blue ellipse areas represent the feature space of the source and target domain. Due to the learning bias, the source model can only work well in the source domain, making it necessary to estimate the source domain (marked as green ellipse) and transfer the knowledge to target model during adaptation. Following the above principle analysis, a source-free UDA framework combining knowledge transfer and adaptation is proposed for semantic segmentation.
2 Source-Free Domain Knowledge Transfer
Following BNS-guided data-free knowledge distillation , the feature distribution of estimated source samples is supposed to satisfy the batch normalization statistics of the source segmentation model. Hence, we apply a BNS constraint on the generator:
Analogously, we can define the discrepancy between the source and target models as follows:
in which obtains the feature map extracted from the backbone of target model with the target data as input. and are the spatial and channel attention maps extracted from the feature maps, which will be defined at Sec 3.2.2. The motivation behind this equation is that the data generated by the generator is not enough to restore the contextual relationships of the source data, due to the lack of necessary prior information. Fortunately, the unlabeled target data has a similar domain-agnostic semantic structure with the real source data to a certain extent. This provides valuable knowledge for the generator to synthesize fake images. So we adopt Kullback-Leibler (KL) divergence to measure the distribution distance of the dual attention maps of fake source and target data, then minimize it in optimization.
2.2 Dual Attention Module
where measures the impact of the -th position on the -th position.
where measures the impact of the -th channel on the -th channel.
After obtaining the spatial and channel attention maps, the dual attention map of sample can be calculated by concatenating the two attention maps:
To transform the spatial and channel attention maps to the same shape, they are multiplied by the original feature , respectively.
2.3 Objective Function
In this way, we have introduced all the necessary components for source-free domain knowledge transfer (SFKT). The generator in our framework aims to synthesize valuable fake samples for transferring source knowledge from the source model to the target model. First, it is supposed to make the fake samples comply with the BNS constraints. Second, the generator explores the discrepancy space by maximizing the discrepancy between the source and target models to drive the search for new knowledge. In addition, it’s better to take advantage of the prior attention information in the target domain by minimizing . Hence, the total objective function of generator is formulated as:
where , and are hyper-parameters for balancing the MAE loss and the two DAD losses.
The target model learns from two aspects: the target pseudo-labels and two-level knowledge from the source model. We hope that while reducing the uncertainty of the target domain, the target also preserves the source domain information to guide adaptive learning by minimizing the output and attention discrepancy (two-level) with the source model. The objective function of target model in knowledge transfer stage is as follows:
3 Self-supervised Model Adaptation
Since it is hard for the generator to guarantee to continuously restore and transfer the information precisely covering the source domain, we draw inspiration from the self-supervision mechanism and consider taking advantage of the valuable information output by target model for target data. Through analyzing the prediction of the initial target model on the target domain, we found that its prediction on most patches are correct, in which there are useful supervision information for learning on uncertain or error patches.
To take advantage of the pseudo-labels in UDA-based segmentation, Pan et al. proposed an unsupervised inter-domain and intra-domain adaptation method, which first separates the target domain into easy and hard splits using an entropy-based ranking function, and then decreases the inter-domain or intra-domain gap via an adversarial mechanism. However, in reality, the gap between the source and the target domain is too large, making it difficult to filter out a sufficient number of easy splits in the target domain for intra-domain supervision. What makes matters worse is that the source domain is unavailable in our setting.
Then, the mean entropy score of each prediction map for the target image is defined as:
In a batch containing (even number) target images , entropy-ranking is executed on patch entropy maps at the same position or class. The prediction maps in each class with lower entropy are assigned to the easy group , while the other are assigned to the hard group . This process is given as follows:
After obtaining the prediction maps of hard and easy patches, we train a discriminator . aims to discriminates easy and hard patches, while is trained to fool from the side of hard patches to reduce the gap between patches. The adversarial learning loss to optimize and is given by:
3.2 Objective Function
Upon this, we extend the objective function in Equation 12 by adding the adversarial loss \wrtIPSM and the self-supervision loss. As a result, we define the following objective function to train the target and source models (\ie, and ) with shared weights:
where is the hyper-parameter to control the adversarial loss. The detailed training algorithm is presented in the supplementary material.
Experiments
Datasets We evaluate our SFDA framework on semantic segmentation under two different settings: synthetic-to-real and cross-city. For the former setting, we follow previous work by considering Cityscapes as the target domain, and GTA5 or SYNTHIA as the source domain. For the latter setting, Cityscapes dataset is used as the source domain and NTHU dataset is as the target domain.
Cityscapes provides 3,975 images with fine-grained segmentation annotations. The synthetic dataset GTA5 contains 24,966 annotated images with a resolution of 1,9141,052 taken from the GTA5 game. SYNTHIA is used as another synthetic dataset, which contains 9,400 fully annotated 1,280760 RGB images. The NTHU dataset contains four different cities: Rio, Rome, Tokyo, and Taipei.
Metrics The semantic segmentation performance is evaluated on every category using Intersection-over-Union (IoU) ratio and Pixel-Accuracy (PA). For the whole test set, we calculate Mean Intersection-over-Union (mIoU) and Mean Pixel-Accuracy (mPA).
1.2 Implementation Details
Two kinds of segmentation networks are adopted in our experiments. One is DeepLabV3 with the ResNet-50 pre-trained on ImageNet , and the other is SegNet with the pre-trained VGG-16 backbone. Considering SegNet in an encoder-decoder architecture, the DAM is connected behind the encoder. When calculating the dual attention maps of target images, an adaptive pooling is applied before DAM. For the generator and the discriminator , we use an architecture similar to but extend to a conditional version. The input channel of is set to be consistent with the output channel of prediction maps. The latent space dimension for and label embedding dimension for both are 256. The architectures of the generator and the discriminator are detailed in the supplementary material.
We implement the proposed framework using the PyTorch toolbox on two GTX 2080Ti GPUs. To train the segmentation networks, we use the Stochastic Gradient Descent (SGD) optimizer with Nesterov acceleration where the momentum is 0.9 and the weight decay is . The initial learning rate is set to and is decreased using the polynomial decay with a power of 0.9 as mentioned in . For training the generator and discriminator, Adam optimizer with an initial learning rate of 0.1 is adopted. Due to the difficulty in generating high-resolution images, we resize the images to 512256 for all datasets. Thanks to full-convolutional segmentation networks, we can set the resolution of synthetic samples to 256128, which is lower than target data but enough for transferring knowledge. To get a high-quality source model for adaptation, we pre-train the source models for 30 epochs on Cityscapes while for 20 epochs on GTA5 or SYNTHIA. In source-free adaptation, the target model, the generator, and the discriminator are jointly trained on a target domain for 120 epochs with a batch size of 8.
As for hyper-parameters, and are set to 1.0 and 0.5 by default, respectively. Notably we set to balance two DAD losses. We set to 0.01 in all experiments if not particularly indicated. The number of patches, \ie, in IPSM is reasonable to choose from .
2 Comparison
Synthetic-to-Real Adaptation: (1) GTA5 Cityscapes. Figure 6 shows the qualitative results on GTA5 Cityscapes. In order to show the versatility of SFKT and the contribution of IPSM, we remove the IPSM part in our architecture, namely ‘SFDA (w/o IPSM)’. It is obvious that even without source data, our method outperforms traditional MinEnt method. What’s more, with the enhancement of IPSM, our full method can make up for errors in some areas through self-supervision, shown in the yellow dashed box. We present adaptation results in Table 1 with comparisons to the state-of-the-art source-driven domain adaptation methods.
(2) SYNTHIA Cityscapes. Following the evaluation setting in , we present the results of IoU and mIoU w.r.t. 16-class and 13-class segmentation in Table 2, respectively. Our architecture is used with DeepLabV3, and even outperforms the source-driven UDA methods with the assistance of IPSM. Besides, our method achieves competitive performance for the small object segmentation, such as traffic light, traffic sign, and motorbike.
Cross-City Adaptation: To show the effectiveness of our methods for smaller domain shift, we conduct the experiment on Cityscapes NTHU with DeepLabV3 architecture. Table 3 shows the comparisons of our method with other source-driven UDA methods. Compared to the best UDA method MaxSquare, our method with IPSM achieves competitive performance on four city datasets. In addition, we distill source domain knowledge via SFKT from well-trained source model into a new model and evaluate it on target domain without adaptation, shown as ‘transfer only’ in the table. The results demonstrate that the knowledge we obtained via SFKT is still valuable on the target, although the effect is not as good as ‘source only’.
3 Ablation Study
To show the detailed contributions of the components in SFKT, we conduct ablation experiments on three datasets, shown in Table 4. The results demonstrate that the DAD losses in source-free domain knowledge transfer is more effective than the commonly used BNS loss, and the fusion of them could further improve the performance.
The visualization of semantic maps and fake samples synthesized in the knowledge transfer stage are shown in Figure 7. The left two columns are the fake samples synthesized by generator and corresponding semantic maps predicted by DeepLabV3 pre-trained on Cityscapes. The right two columns are several semantic maps predicted without DAD or BNS loss. On one hand, the output semantic maps are similar to the real-world street view structure without DAD, but it is hard to pay attention to some small objects or refined segmentation. On the other hand, the generator captures the discrepancy between two models, but cannot preserve the original semantic distribution of source domain without BNS loss, which is vital for segmentation tasks. Although the fake samples cannot be recognized by humans, they have similar representations and outputs in convolutional neural networks with the source domain data. Hence, the fake samples become the key to transfer source domain knowledge.
4 Hyper-parameter Analysis
Firstly, we discuss the influence of and (), the weights for the MAE loss and the DAD losses, respectively, for DeepLabV3 on GTA5 Cityscapes. Given , we adjust from 0.1 to 2.0, and show the results in Table 5. Since the MAE loss of the source prediction output is similar to the target segmentation loss when supervised by target pseudo-labels, should be close to 1.0. Otherwise, there will be disagreements with , resulting in bias during adaptation.
Analogously, given , we adjust from 0.01 to 1.0, and the results are shown in Table 6. Different from , controls the weights of the DAD losses in intermediate layers, so they are supposed to be smaller than . If too many weights are allocated to the DAD losses, they will limit the learning capacity of the intermediate layers.
We show the sensitivity analysis of parameters in Figure 8, from which we observe that too large or too small is not suitable for IPSM, and 3 to 5 is reasonable. Note that when , it means IPSM is not adopted in training.
Conclusion
In this paper, we have presented a novel source-free domain adaptation framework (SFDA) for semantic segmentation. It aims to preserve the source domain knowledge from a fixed source model via knowledge transfer. Specifically, a dual attention distillation method is designed to capture and transfer pixel-level semantic information for segmentation tasks. Moreover, during model adaptation, an intra-domain patch-level self-supervision mechanism is introduced to take advantage of valuable knowledge at patch-level pseudo-labels in a target domain. We conduct extensive experiments and ablation studies to validate the effectiveness of the proposed framework on different segmentation tasks, showing it performs favorably against existing source-driven UDA methods. However, our approach does not support high-resolution image segmentation tasks due to the limitation of generative fake sample synthesis, which will be tackled in future work.
Acknowledgement
This work was supported in part by National Natural Science Foundation of China under Grant (No. 62072182).