How does the Combined Risk Affect the Performance of Unsupervised Domain Adaptation Approaches?
Li Zhong, Zhen Fang, Feng Liu, Jie Lu, Bo Yuan, Guangquan Zhang
Introduction
Domain Adaptation (DA) aims to train a target-domain classifier with samples from source and target domains (Lu et al. 2015). When the labels of samples in the target domain are unavailable, DA is known as unsupervised DA (UDA) (Zhong et al. 2020; Fang et al. 2020), which has been applied to address diverse real-world problems, such as computer version (Zhang et al. 2020c; Dong et al. 2019, 2020b), natural language processing (Lee and Jha 2019; Guo, Pasunuru, and Bansal 2020), and recommender system (Zhang et al. 2017; Yu, Wang, and Yuan 2019; Lu et al. 2020)
Significant theoretical advances have been achieved in UDA. Pioneering theoretical work was proposed by Ben-David et al. (2007). This work shows that the target risk is upper bounded by three terms: source risk, marginal distribution discrepancy, and combined risk. This earliest learning bound has been extended from many perspectives, such as considering more surrogate loss functions (Zhang et al. 2019a) or distributional discrepancies (Mohri and Medina 2012; Shen et al. 2018) (see (Redko et al. 2020) as a survey). Recently, Zhang et al. (2019a) proposed a new distributional discrepancy termed Margin Disparity Discrepancy and developed a tighter and more practical UDA learning bound.
The UDA learning bounds proposed by (Ben-David et al. 2007, 2010) and the recent UDA learning bounds proposed by (Shen et al. 2018; Xu et al. 2020; Zhang et al. 2020b) consist of three terms: source risk, marginal distribution discrepancy, and combined risk. Minimizing the source risk aims to obtain a source-domain classifier, and minimizing the distribution discrepancy aims to learn domain-invariant features so that the source-domain classifier can perform well on the target domain. The combined risk embodies the adaptability between the source and target domains (Ben-David et al. 2010). In particularly, when the hypothesis space is fixed, the combined risk is a constant.
Based on the UDA learning bounds where the combined risk is assumed to a small constant, many existing UDA methods focus on learning domain-invariant features (Fang et al. 2019; Dong et al. 2020c, a; Liu et al. 2019) by minimizing the estimators of the source risk and the distribution discrepancy. In the learned feature space, the source and target distributions are similar while the source-domain classifier is required to achieve a small error. Furthermore, the generalization error of the source-domain classifier is expected to be small in the target domain.
However, the combined risk may increase when learning the domain-invariant features, and the increase of the combined risk may degrade the performance of the source-domain classifier in the target domain. As shown in Figure 1, we calculate the value of the combined risk and accuracy on a real-world UDA task (see the green line). The performance worsens with the increased combined risk. Zhao et al. (2019) also pointed out the increase of combined risk causes the failure of source-domain classifier on the target domain.
To investigate how the combined risk affect the performance on the domain-invariant features, we rethink and develop the UDA learning bounds by introducing feature transformations. In the new bound (see Eq. (5)), the combined risk is a function related to feature transformation but not a constant (compared to bounds in (Ben-David et al. 2010)). We also reveal that the combined risk is deeply related to the conditional distribution discrepancy (see Theorem 3). Theorem 3 shows that, the conditional distribution discrepancy will increase when the combined risk increases. Hence, it is hard to achieve satisfied target-domain accuracy if we only focus on learning domain-invariant features and omit to control the combined risk.
To estimate the combined risk, the key challenge takes root in the unavailability of the labeled samples in the target domain. A simple solution is to leverage the pseudo labels with high confidence in the target domain to estimate the combined risk. However, since samples with high confidence are insufficient, the value of the combined risk may still increase (see the green line in Figure 1). Inspired by semi-supervised learning methods, an advanced solution is to directly use the mixup technique to augment pseudo-labeled target samples, which can slightly help us estimate the combined risk better than the simple solution (see the orange line in Figure 1).
However, the target-domain pseudo labels provided by the source-domain classifier may be inaccurate due to the discrepancy between domains, which causes that mixup may not perform well with inaccurate labels. To mitigate the issue, we propose enhanced mixup (e-mixup) to substitute mixup to compute a proxy of the combined risk. The purple line in Figure 1 shows that the proxy based on e-mixup can significantly boost the performance. Details of the proxy is shown in section Motivation.
To the end, we design a novel UDA method referred to E-MixNet. E-MixNet learns the target-domain classifier by simultaneously minimizing the source risk, the marginal distribution discrepancy, and the proxy of combined risk. Via minimizing the proxy of combined risk, we control the increase of combined risk effectively, thus, control the conditional distribution discrepancy between two domains.
We conduct experiments on three public datasets (Office-31, Office-Home, and Image-CLEF) and compare E-MixNet with a series of existing state-of-the-art methods. Furthermore, we introduce the proxy of the combined risk into four representative UDA methods (i.e., DAN (Long et al. 2015), DANN (Ganin et al. 2016), CDAN (Long et al. 2018), SymNets (Zhang et al. 2019b)). Experiments show that E-MixNet can outperform all baselines, and the four representative methods can achieve better performance if the proxy of the combined risk is added into their loss functions.
Problem Setting and Concepts
In this section, we introduce the definition of UDA, then introduce some important concepts used in this paper.
Given random variables , , the source and target domains are joint distributions and with .
Then, we propose the UDA problem as follows.
Given independent and identically distributed (i.i.d.) labeled samples drawn from the source domain and i.i.d. unlabeled samples drawn from the target marginal distribution , the aim of UDA is to train a classifier with and such that classifies accurately target data .
Lastly, we define the disparity discrepancy based on double losses, which will be used to design our method.
Compared with the classical discrepancy distance (Mansour, Mohri, and Rostamizadeh 2009):
double loss disparity discrepancy is tighter and more flexible.
Theoretical Analysis
Here we introduce our main theoretical results. All proofs can be found at https://github.com/zhonglii/E-MixNet.
Double Loss DA Learning Bound
Note that there exist methods, such as MDD (Zhang et al. 2019a), whose source and target losses are different. To understand these UDA methods and bridge the gap between theory and algorithms, we develop the classical DA learning bound to a more general scenario.
Proposed Method: E-MixNet
Here we introduce motivation and details of our method.
Theorem 3 has shown that the combined risk is related to the conditional distribution discrepancy. As the increase of the combined risk, the conditional distribution discrepancy is increased. Hence, omitting the importance of the combined risk may make negative impacts on the target-domain accuracy. Figure 1 (blue line) verifies our observation.
To control the combined risk, we consider the following problem.
However, the combined risk may still increase as shown in Fig 1 (green line). The reason may be that the target samples, whose pseudo labels with high confidence, are insufficient.
Inspired by semi-supervised learning, an advanced solution is to use mixup technique (Zhang et al. 2018) to augment pseudo-labeled target samples. Mixup produces new samples by a convex combination: given any two samples , ,
The aforementioned issue can be mitigated, since mixup can be regarded as a data augmentation (Zhang et al. 2018). However, the target-domain pseudo labels provided by the source-domain classifier may be inaccurate due to the discrepancy between domains, which causes that mixup may not perform well with inaccurate labels. We propose enhanced mixup (e-mixup) to substitute the mixup to compute the proxy. E-mixup introduces the pure true-labeled source-samples to mitigate the issue caused by bad pseudo labels.
The purple line in Figure 1 and ablation study show that e-mixup can further boost performance.
Algorithm
The optimization of the combined risk plays a crucial role in UDA. Accordingly, we propose a method based on the aforementioned analyses to solve UDA more deeply.
According to the theoretical bound in Eq. (5), we need to solve the following problem
Minimizing double loss disparity discrepancy is a minimax game, since the double losses disparity discrepancy is defined as the supremum over hypothesis space . Thus, we revise the above problem as follows:
where are parameters to make our model more flexible,
To solve the problem (10), we construct a deep method. The network architecture is shown in Fig. 2(b), which consists of a generator , a discriminator , and two classifiers . Next, we introduce the details about our method.
here is the -th coordinate function of function .
Source risk. Given the source samples , then
where is the label corresponding to one-hot vector .
Double loss disparity discrepancy. Given the source and target samples , then
Training Procedure
Finally, the UDA problem can be solved by the following minimax game.
The training procedure is shown in Algorithm 2.
Experiments
We evaluate E-Mixnet on three public datasets, and compare it with several existing state-of-the-art methods. Codes will be available at https://github.com/zhonglii/E-MixNet.
Three common UDA datasets are used to evaluate the efficacy of E-MixNet.
Office-31 (Saenko et al. 2010) is an object recognition dataset with images, which consists of three domains with a slight discrepancy: amazon (A), dslr (D) and webcam (W). Each domain contains kinds of objects. So there are domain adaptation tasks on Office-31: A D, A W, D A, D W, W A, W D.
Office-Home (Venkateswara et al. 2017) is an object recognition dataset with image, which contains four domains with more obvious domain discrepancy than Office-31. These domains are Artistic (A), Clipart (C), Product (P), Real-World (R). Each domain contains kinds of objects. So there are domain adaptation tasks on Office-Home: A C, A P, A R, …, R P.
ImageCLEF-DAhttp://imageclef.org/2014/adaptation/ is a dataset organized by selecting the 12 common classes shared by three public datasets (domains): Caltech-256 (C), ImageNet ILSVRC 2012 (I), and Pascal VOC 2012 (P). We permute all three domains and build six transfer tasks: IP, PI, IC, CI, CP, PC.
Experimental Setup
Following the standard protocol for unsupervised domain adaptation in (Ganin et al. 2016; Long et al. 2018), all labeled source samples and unlabeled target samples are used in the training process and we report the average classification accuracy based on three random experiments. in Eq. (15) is selected from 2, 4, 8, and it is set to 2 for Office-Home, 4 for Office-31, and 8 for Image-CLEF.
ResNet-50 (He et al. 2016) pretrained on ImageNet is employed as the backbone network (). , and are all two fully connected layers where the hidden unit is 1024. Gradient reversal layer between G and is employed for adversarial training. The algorithm is implemented by Pytorch. The mini-batch stochastic gradient descent with momentum 0.9 is employed as the optimizer, and the learning rate is adjected by , where i linearly increase from 0 to 1 during the training process, , . We follow (Zhang et al. 2019a) to employ a progressive strategy for : , is set to 0.1. The in e-mixup is set to 0.6 in all experiments.
Results
The results on Office-31 are reported in Tabel 1. E-MixNet achieves the best results and exceeds the baselines for 4 of 6 tasks. Compared to the competitive baseline MDD, E-MixNet surpasses it by 4.3% for the difficult task D A.
The results on Image-CLEF are reported in Table 2. E-MixNet significantly outperforms the baselines for 5 of 6 tasks. For the hard task C P, E-MixNet surpasses the competitive baseline SymNets by 2.7%.
The results on Office-Home are reported in Table 3. Despite Office-Home is a challenging dataset, E-MixNet still achieves better performance than all the baselines for 9 of 12 tasks. For the difficult tasks A C, P A, and R C, E-MixNet has significant advantages.
In order to further verify the efficacy of the proposed proxy of the combined risk, we add the proxy into the loss functions of four representative UDA methods. As shown in Fig. 2(a), we add a new classifier that is the same as the classifier in the original method to formulate the proxy of the combined risk. The results are shown in Table 4. The four methods can achieve better performance after optimizing the proxy. It is worth noting that DANN obtains a 5.5% percent increase. The experiments adequately demonstrate the combined risk plays a crucial role for methods that aim to learn a domain-invariant representation and the proxy can indeed curb the increase of the combined risk.
Ablation Study and Parameter Analysis
Parameter analysis. Here we aim to study how the parameter affects the performance and the efficiencies of mean square error (MSE) and cross-entropy for the proxy of combined risk. Firstly, as shown in Fig. 3(a), a relatively larger can obtain better performance and faster convergence. Secondly, when mixup behaves between two samples, the accuracy of the pseudo labels of the target samples are much important. To against the adversarial samples, MSE is employed to substitute cross-entropy. As shown in Fig. 3(b), MSE can obtain more stable and better performance. Furthermore, -distance is also an important indicator showing the distribution discrepancy, which is defined as where is the test error. As shown in Fig. 3 (c). E-MixNet achieves a better performance of adaptation, implying the efficiency of the proposed proxy.
Conclusion
Though numerous UDA methods have been proposed and have achieved significant success, the issue caused by combined risk has not been brought to the forefront and none of the proposed methods solve the problem. Theorem 3 reveals that the combined risk is deeply related to the conditional distribution discrepancy and plays a crucial role for transfer performance. Furthermore, we propose a method termed E-MixNet, which employs enhanced mixup to calculate a proxy of the combined risk. Experiments show that our method achieves a comparable performance compared with existing state-of-the-art methods and the performance of the four representative methods can be boosted by adding the proxy into their loss functions.
Acknowledgments
The work presented in this paper was supported by the Australian Research Council (ARC) under DP170101632 and FL190100149. The first author particularly thanks the support of UTS-AAII during his visit.