CrDoCo: Pixel-level Domain Transfer with Cross-Domain Consistency
Yun-Chun Chen, Yen-Yu Lin, Ming-Hsuan Yang, Jia-Bin Huang
Introduction
Deep convolutional neural networks (CNNs) are extremely data hungry. However, for many dense prediction tasks (e.g., semantic segmentation, optical flow estimation, and depth prediction), collecting large-scale and diverse datasets with pixel-level annotations is difficult since the labeling process is often expensive and labor intensive (see Figure 1). Developing algorithms that can transfer the knowledge learned from one labeled dataset (i.e., source domain) to another unlabeled dataset (i.e., target domain) thus becomes increasingly important. Nevertheless, due to the domain-shift problem (i.e., the domain gap between the source and target datasets), the learned models often fail to generalize well to new datasets.
To address these issues, several unsupervised domain adaptation methods have been proposed to align data distributions between the source and target domains. Existing methods either apply feature-level or pixel-level adaptation techniques to minimize the domain gap between the source and target datasets. However, aligning marginal distributions does not necessarily lead to satisfactory performance as there is no explicit constraint imposed on the predictions in the target domain (as no labeled training examples are available). While several methods have been proposed to alleviate this issue via curriculum learning or self-paced learning , the problem remains challenging since these methods may only learn from cases where the current models perform well.
In this paper, we present CrDoCo, a pixel-level adversarial domain adaptation algorithm for dense prediction tasks. Our model consists of two main modules: 1) an image-to-image translation network and 2) two domain-specific task networks (one for source and the other for target). The image translation network learns to translate images from one domain to another such that the translated images have a similar distribution to those in the translated domain. The domain-specific task network takes images of source/target domain as inputs to perform dense prediction tasks. As illustrated in Figure 2, our core idea is that while the original and the translated images in two different domains may have different styles, their predictions from the respective domain-specific task network should be exactly the same. We enforce this constraint using a cross-domain consistency loss that provides additional supervisory signals for facilitating the network training, allowing our model to produce consistent predictions. We show the applicability of our approach to multiple different tasks in the unsupervised domain adaptation setting.
First, we present an adversarial learning approach for unsupervised domain adaptation which is applicable to a wide range of dense prediction tasks. Second, we propose a cross-domain consistency loss that provides additional supervisory signals for network training, resulting in more accurate and consistent task predictions. Third, extensive experimental results demonstrate that our method achieves the state-of-the-art performance against existing unsupervised domain adaptation techniques. Our source code is available at https://yunchunchen.github.io/CrDoCo/
Related Work
Unsupervised domain adaptation methods can be categorized into two groups: 1) feature-level adaptation and 2) pixel-level adaptation. Feature-level adaptation methods aim at aligning the feature distributions between the source and target domains through measuring the correlation distance , minimizing the maximum mean discrepancy , or applying adversarial learning strategies in the feature space. In the context of image classification, several methods have been developed to address the domain-shift issue. For semantic segmentation tasks, existing methods often align the distributions of the feature activations at multiple levels . Recent advances include applying class-wise adversarial learning or leveraging self-paced learning policy for adapting synthetic-to-real or cross-city adaptation , adopting curriculum learning for synthetic-to-real foggy scene adaptation , or progressively adapting models from daytime scene to nighttime . Another line of research focuses on pixel-level adaptation . These methods address the domain gap problem by performing data augmentation in the target domain via image-to-image translation or style transfer methods.
Most recently, a number of methods tackle joint feature-level and pixel-level adaptation in image classification , semantic segmentation , and single-view depth prediction tasks. These methods utilize image-to-image translation networks (e.g., the CycleGAN ) to translate images from source domain to target domain with pixel-level adaptation. The translated images are then passed to the task network followed by a feature-level alignment.
While both feature-level and pixel-level adaptation have been explored, aligning the marginal distributions without enforcing explicit constraints on target predictions would not necessarily lead to satisfactory performance. Our model builds upon existing techniques for feature-level and pixel-level adaptation . The key difference lies in our cross-domain consistency loss that explicitly penalizes inconsistent predictions by the task networks.
Cycle consistency constraints have been successfully applied to various problems. In image-to-image translation, enforcing cycle consistency allows the network to learn the mappings without paired data . In semantic matching, cycle or transitivity based consistency loss help regularize the network training . In motion analysis, forward-backward consistency check can be used for detecting occlusion or learning visual correspondence . Similar to the above methods, we show that enforcing two domain-specific networks to produce consistent predictions leads to substantially improved performance.
Training the model on large-scale synthetic datasets has been extensively studied in semantic segmentation , multi-view stereo , depth estimation , optical flow , amodal segmentation , and object detection . In our work, we show that the proposed cross-domain consistency loss can be applied not only to synthetic-to-real adaptation but to real-to-real adaptation tasks as well.
Method
In this section, we first provide an overview of our approach. We then describe the proposed loss function for enforcing cross-domain consistency on dense prediction tasks. Finally, we describe other losses that are adopted to facilitate network training.
We consider the task of unsupervised domain adaptation for dense prediction tasks. In this setting, we assume that we have access to a source image set , a source label set , and an unlabeled target image set . Our goal is to learn a task network that can reliably and accurately predict the dense label for each image in the target domain.
To achieve this task, we present an end-to-end trainable network which is composed of two main modules: 1) the image translation network and and 2) two domain-specific task networks and . The image translation network translates images from one domain to the other. The domain-specific task network takes input images to perform the task of interest.
As shown in Figure 3, the proposed network takes an image from the source domain and another image from the target domain as inputs. We first use the image translation network to obtain the corresponding translated images (in the target domain) and (in the source domain). We then pass and to , and to to obtain their task predictions.
2 Objective function
where and are the task predictions for and , respectively, while denotes the number of classes.
4 Other losses
To guide the training of the two task networks and using labeled data, for each image-label pair (, ) in the source domain, we first translate the source domain image to by passing to (i.e., ). Similarly, images before and after translation should have the same ground truth label. Namely, the label for is identical to that of which is .
Based on the aforementioned loss functions, we aim to solve for a target domain task network by optimizing the following min-max problem:
5 Implementation details
Experimental Results
We present experimental results for semantic segmentation in two different settings: 1) synthetic-to-real: adapting from synthetic GTA5 and SYNTHIA datasets to real-world images from Cityscapes dataset and 2) real-to-real: adapting the Cityscapes dataset to different cities .
The GTA5 dataset consists of synthetic images with pixel-level annotations of categories (compatible with the Cityscapes dataset ). Following Hoffman et al. , we use the GTA5 dataset and adapt the model to the Cityscapes training set with images.
We evaluate our model on the Cityscapes validation set with images using the mean intersection-over-union (IoU) and the pixel accuracy as the evaluation metrics.
We evaluate our proposed method using two task networks: 1) dilated residual network- (DRN-) and 2) FCNs-VGG . For the DRN-, we initialize our task network from Hoffman et al. . For the FCNs-VGG, we initialize our task network from Sankaranarayanan et al. .
1.2 SYNTHIA to Cityscapes
We use the SYNTHIA-RAND-CITYSCAPES set as the source domain which contains images compatible with the Cityscapes annotated classes. Following Dundar et al. , we evaluate images on the Cityscapes validation set with classes.
1.3 Cityscapes to Cross-City
In addition to the synthetic-to-real adaptation, we conduct an experiment on the Cross-City dataset which is a real-to-real adaptation. The dataset contains four different cities: Rio, Rome, Tokyo, and Taipei, where each city has images without annotations and images with pixel-level ground truths for classes. Following Tsai et al. , we use the Cityscapes training set as our source domain and adapt the model to each target city using images, and use the annotated images for evaluation.
We compare our approach with the Cross-City , the CBST , and the AdaptSegNet . Table 2 shows that our method achieve state-of-the-art performance on two out of four cities. Note that the results in AdaptSegNet are obtained by using a ResNet- . We run their publicly available code with the default settings and report the results using the ResNet- as the feature backbone for a fair comparison. Under the same experimental setting, our approach compares favorably against state-of-the-art methods. Furthermore, we show that enforcing cross-domain consistency constraints, our method effectively and consistently improves the results evaluated on all four cities.
2 Single-view depth estimation
To show that our formulation is not limited to semantic segmentation, we present experimental results for single-view depth prediction task. Specifically, we use SUNCG as the source domain and adapt the model to the NYUDv2 dataset.
To generate the paired synthetic training data, we rendered RGB images and depth map from the SUNCG dataset , which contains 3D houses with various room types. Following Zheng et al. , we choose the camera locations, poses and parameters based on the distribution of real NYUDv2 dataset and retain valid depth maps using the criteria described by Song et al. . In total, we generate valid views from different houses.
We use the root mean square error (RMSE) and the log scale version (RMSE log.), the squared relative difference (Sq. Rel.) and the absolute relative difference (Abs. Rel.), and the accuracy measured by thresholding ( threshold).
We initialize our task network from the unsupervised version of Zheng et al. .
3 Optical flow estimation
We show evaluations of the model trained on a synthetic dataset (i.e., MPI Sintel ) and test the adapted model on real-world images from the KITTI 2012 and KITTI 2015 datasets.
The MPI Sintel dataset consists of images rendered from artificial scenes. There are two versions: 1) the final version consists of images with motion blur and atmospheric effects, and 2) the clean version does not include these effects. We use the clean version as the source dataset. We report two results obtained by 1) using the KITTI 2012 as the target dataset and 2) using the KITTI 2015 as the target dataset.
We adopt the average endpoint error (AEPE) and the F1 score for both KITTI 2012 and KITTI 2015 to evaluate the performance.
Our task network is initialized from the PWC-Net (without finetuning on the KITTI dataset).
4 Limitations
Our method is memory-intensive as the training involves multiple networks at the same time. Potential approaches to alleviate this issue include 1) adopting partial sharing on the two task networks, e.g., share the last few layers of the two task networks, and 2) sharing the encoders in the image translation network (i.e., and ).
Conclusions
We have presented a simple yet surprisingly effective loss for improving pixel-level unsupervised domain adaption for dense prediction tasks. We show that by incorporating the proposed cross-domain consistency loss, our method consistently improves the performances over a wide range of tasks. Through extensive experiments, we demonstrate that our method is applicable to a wide variety of tasks.
This work was supported in part by NSF under Grant No. 1755785, No. 1149783, Ministry of Science and Technology (MOST) under grants 107-2628-E-001-005-MY3 and 108-2634-F-007-009, and gifts from Adobe, Verisk, and NEC. We thank the support of NVIDIA Corporation with the GPU donation.