A DIRT-T Approach to Unsupervised Domain Adaptation
Rui Shu, Hung H. Bui, Hirokazu Narui, Stefano Ermon
Introduction
The development of deep neural networks has enabled impressive performance in a wide variety of machine learning tasks. However, these advancements often rely on the existence of a large amount of labeled training data. In many cases, direct access to vast quantities of labeled data for the task of interest (the target domain) is either costly or otherwise absent, but labels are readily available for related training sets (the source domain). A notable example of this scenario occurs when the source domain consists of richly-annotated synthetic or semi-synthetic data, but the target domain consists of unannotated real-world data (Sun & Saenko, 2014; Vazquez et al., 2014). However, the source data distribution is often dissimilar to the target data distribution, and the resulting significant covariate shift is detrimental to the performance of the source-trained model when applied to the target domain (Shimodaira, 2000).
Solving the covariate shift problem of this nature is an instance of domain adaptation (Ben-David et al., 2010b). In this paper, we consider a challenging setting of domain adaptation where 1) we are provided with fully-labeled source samples and completely-unlabeled target samples, and 2) the existence of a classifier in the hypothesis space with low generalization error in both source and target domains is not guaranteed. Borrowing approximately the terminology from Ben-David et al. (2010b), we refer to this setting as unsupervised, non-conservative domain adaptation. We note that this is in contrast to conservative domain adaptation, where we assume our hypothesis space contains a classifier that performs well in both the source and target domains.
To tackle unsupervised domain adaptation, Ganin & Lempitsky (2015) proposed to constrain the classifier to only rely on domain-invariant features. This is achieved by training the classifier to perform well on the source domain while minimizing the divergence between features extracted from the source versus target domains. To achieve divergence minimization, Ganin & Lempitsky (2015) employ domain adversarial training. We highlight two issues with this approach: 1) when the feature function has high-capacity and the source-target supports are disjoint, the domain-invariance constraint is potentially very weak (see Section 3), and 2) good generalization on the source domain hurts target performance in the non-conservative setting.
Saito et al. (2017) addressed these issues by replacing domain adversarial training with asymmetric tri-training (ATT), which relies on the assumption that target samples that are labeled by a source-trained classifier with high confidence are correctly labeled by the source classifier. In this paper, we consider an orthogonal assumption: the cluster assumption (Chapelle & Zien, 2005), that the input distribution contains separated data clusters and that data samples in the same cluster share the same class label. This assumption introduces an additional bias where we seek decision boundaries that do not go through high-density regions. Based on this intuition, we propose two novel models: 1) the Virtual Adversarial Domain Adaptation (VADA) model which incorporates an additional virtual adversarial training (Miyato et al., 2017) and conditional entropy loss to push the decision boundaries away from the empirical data, and 2) the Decision-boundary Iterative Refinement Training with a Teacher (DIRT-T) model which uses natural gradients to further refine the output of the VADA model while focusing purely on the target domain. We demonstrate that
In conservative domain adaptation, where the classifier is trained to perform well on the source domain, VADA can be used to further constrain the hypothesis space by penalizing violations of the cluster assumption, thereby improving domain adversarial training.
In non-conservative domain adaptation, where we account for the mismatch between the source and target optimal classifiers, DIRT-T allows us to transition from a joint (source and target) classifier (VADA) to a better target domain classifier. Interestingly, we demonstrate the advantage of natural gradients in DIRT-T refinement steps.
We report results for domain adaptation in digits classification (MNIST-M, MNIST, SYN DIGITS, SVHN), traffic sign classification (SYN SIGNS, GTSRB), general object classification (STL-10, CIFAR-10), and Wi-Fi activity recognition (Yousefi et al., 2017). We show that, in nearly all experiments, VADA improves upon previous methods and that DIRT-T improves upon VADA, setting new state-of-the-art performances across a wide range of domain adaptation benchmarks. In adapting MNIST SVHN, a very challenging task, we out-perform ATT by over .
Related Work
Given the extensive literature on domain adaptation, we highlight several works most relevant to our paper. Shimodaira (2000); Mansour et al. (2009) proposed to correct for covariate shift by re-weighting the source samples such that the discrepancy between the target distribution and re-weighted source distribution is minimized. Such a procedure is problematic, however, if the source and target distributions do not contain sufficient overlap. Huang et al. (2007); Long et al. (2015); Ganin & Lempitsky (2015) proposed to instead project both distributions into some feature space and encourage distribution matching in the feature space. Ganin & Lempitsky (2015) in particular encouraged feature matching via domain adversarial training, which corresponds approximately to Jensen-Shannon divergence minimization (Goodfellow et al., 2014). To better perform non-conservative domain adaptation, Saito et al. (2017) proposed to modify tri-training (Zhou & Li, 2005) for domain adaptation, leveraging the assumption that highly-confident predictions are correct predictions (Zhu, 2005). Several of aforementioned methods are based on Ben-David et al. (2010a)’s theoretical analysis of domain adaptation, which states the following,
(Ben-David et al., 2010a) Let be the hypothesis space and let and be the two domains and their corresponding generalization error functions. Then for any ,
where denotes the -distance between the domains and ,
Intuitively, measures the extent to which small changes to the hypothesis in the source domain can lead to large changes in the target domain. It is evident that relates intimately to the complexity of the hypothesis space and the divergence between the source and target domains. For infinite-capacity models and domains with disjoint supports, is maximal.
A critical component to our paper is the cluster assumption, which states that decision boundaries should not cross high-density regions (Chapelle & Zien, 2005). This assumption has been extensively studied and leveraged for semi-supervised learning, leading to proposals such as conditional entropy minimization (Grandvalet & Bengio, 2005) and pseudo-labeling (Lee, 2013). More recently, the cluster assumption has led to many successful deep semi-supervised learning algorithms such as semi-supervised generative adversarial networks (Dai et al., 2017), virtual adversarial training (Miyato et al., 2017), and self/temporal-ensembling (Laine & Aila, 2016; Tarvainen & Valpola, 2017). Given the success of the cluster assumption in semi-supervised learning, it is natural to consider its application to domain adaptation. Indeed, Ben-David & Urner (2014) formalized the cluster assumption through the lens of probabilistic Lipschitzness and proposed a nearest-neighbors model for domain adaptation. Our work extends this line of research by showing that the cluster assumption can be applied to deep neural networks to solve complex, high-dimensional domain adaptation problems. Independently of our work, French et al. (2017) demonstrated the application of self-ensembling to domain adaptation. However, our work additionally considers the application of the cluster assumption to non-conservative domain adaptation.
Limitation of Domain Adversarial Training
Before describing our model, we first highlight that domain adversarial training may not be sufficient for domain adaptation if the feature extraction function has high-capacity. Consider a classifier , parameterized by , that maps inputs to the -simplex (denote as ), where is the number of classes. Suppose the classifier can be decomposed as the composite of an embedding function and embedding classifier . For the source domain, let be the joint distribution over input and one-hot label and let be the marginal input distribution. are analogously defined for the target domain. Let be the loss functions
where the supremum ranges over discriminators . Then is the cross-entropy objective and is a domain discriminator. Domain adversarial training minimizes the objective
where is a weighting factor. Minimization of encourages the learning of a feature extractor for which the Jensen-Shannon divergence between and is small.In practice, the minimization of requires solving a mini-max optimization problem. We discuss this in more detail in Appendix C Ganin & Lempitsky (2015) suggest that successful adaptation tends to occur when the source generalization error and feature divergence are both small.
It is easy, however, to construct situations where this suggestion fails. In particular, if has infinite-capacity and the source-target supports are disjoint, then can employ arbitrary transformations to the target domain so as to match the source feature distribution (see Appendix E for formalization). We verify empirically that, for sufficiently deep layers, jointly achieving small source generalization error and feature divergence does not imply high accuracy on the target task (Table 5). Given the limitations of domain adversarial training, we wish to identify additional constraints that one can place on the model to achieve better, more reliable domain adaptation.
Constraining via Conditional Entropy Minimization
In this paper, we apply the cluster assumption to domain adaptation. The cluster assumption states that the input distribution contains clusters and that points in the same cluster come from the same class. This assumption has been extensively studied and applied successfully to a wide range of classification tasks (see Section 2). If the cluster assumption holds, the optimal decision boundaries should occur far away from data-dense regions in the space of (Chapelle & Zien, 2005). Following Grandvalet & Bengio (2005), we achieve this behavior via minimization of the conditional entropy with respect to the target distribution,
Intuitively, minimizing the conditional entropy forces the classifier to be confident on the unlabeled target data, thus driving the classifier’s decision boundaries away from the target data (Grandvalet & Bengio, 2005). In practice, the conditional entropy must be empirically estimated using the available data. However, Grandvalet & Bengio (2005) note that this approximation breaks down if the classifier is not locally-Lipschitz. Without the locally-Lipschitz constraint, the classifier is allowed to abruptly change its prediction in the vicinity of the training data points, which 1) results in a unreliable empirical estimate of conditional entropy and 2) allows placement of the classifier decision boundaries close to the training samples even when the empirical conditional entropy is minimized. To prevent this, we propose to explicitly incorporate the locally-Lipschitz constraint via virtual adversarial training (Miyato et al., 2017) and add to the objective function the additional term
which enforces classifier consistency within the norm-ball neighborhood of each sample . Note that virtual adversarial training can be applied with respect to either the target or source distributions. We can combine the conditional entropy minimization objective and domain adversarial training to yield
a basic combination of domain adversarial training and semi-supervised training objectives. We refer to this as the Virtual Adversarial Domain Adaptation (VADA) model. Empirically, we observed that the hyperparameters are easy to choose and work well across multiple tasks (Appendix B).
-Distance Minimization. VADA aligns well with the theory of domain adaptation provided in Theorem 1. Let the loss,
denote the degree to which the target-side cluster assumption is violated. Modulating enables VADA to trade-off between hypotheses with low target-side cluster assumption violation and hypotheses with low source-side generalization error. Setting allows rejection of hypotheses with high target-side cluster assumption violation. By rejecting such hypotheses from the hypothesis space , VADA reduces and yields a tighter bound on the target generalization error. We verify empirically that VADA achieves significant improvements over existing models on multiple domain adaptation benchmarks (Table 1).
Decision-boundary Iterative Refinement Training
In non-conservative domain adaptation, we assume the following inequality,
where are generalization error functions for the source and target domains. This means that, for a given hypothesis class , the optimal classifier in the source domain does not coincide with the optimal classifier in the target domain.
We assume that the optimality gap in Eq. 10 results from violation of the cluster assumption. In other words, we suppose that any source-optimal classifier drawn from our hypothesis space necessarily violates the cluster assumption in the target domain. Insofar as VADA is trained on the source domain, we hypothesize that a better hypothesis is achievable by introducing a secondary training phase that solely minimizes the target-side cluster assumption violation.
Under this assumption, the natural solution is to initialize with the VADA model and then further minimize the cluster assumption violation in the target domain. In particular, we first use VADA to learn an initial classifier . Next, we incrementally push the classifier’s decision boundaries away from data-dense regions by minimizing the target-side cluster assumption violation loss in Eq. 9. We denote this procedure Decision-boundary Iterative Refinement Training (DIRT).
Stochastic gradient descent minimizes the loss by selecting gradient steps according to the following objective,
which defines the neighborhood in the parameter space. This notion of neighborhood is sensitive to the parameterization of the model; depending on the parameterization, a seemingly small step may result in a vastly different classifier. This contradicts our intention of incrementally and locally pushing the decision boundaries to a local conditional entropy minimum, which requires that the decision boundaries of stay close to that of . It is therefore important to define a neighborhood that is parameterization-invariant. Following Pascanu & Bengio (2013), we instead select using the following objective,
Each optimization step now solves for a gradient step that minimizes the conditional entropy, subject to the constraint that the Kullback-Leibler divergence between and is small for . The corresponding Lagrangian suggests that one can instead minimize a sequence of optimization problems
that approximates the application of a series of natural gradient steps.
In practice, each of the optimization problems in Eq. 14 can be solved approximately via a finite number of stochastic gradient descent steps. We denote the number of steps taken to be the refinement interval . Similar to Tarvainen & Valpola (2017), we use the Adam Optimizer with Polyak averaging (Polyak & Juditsky, 1992). We interpret as a (sub-optimal) teacher for the student model , which is trained to stay close to the teacher model while seeking to reduce the cluster assumption violation. As a result, we denote this model as Decision-boundary Iterative Refinement Training with a Teacher (DIRT-T).
Weakly-Supervised Learning. This sequence of optimization problems has a natural interpretation that exposes a connection to weakly-supervised learning. In each optimization problem, the teacher model pseudo-labels the target samples with noisy labels. Rather than naively training the student model on the noisy labels, the additional training signal allows the student model to place its decision boundaries further from the data. If the clustering assumption holds and the initial noisy labels are sufficiently similar to the true labels, conditional entropy minimization can improve the placement of the decision boundaries (Reed et al., 2014).
Domain Adaptation. An alternative interpretation is that DIRT-T is the recursive extension of VADA, where the act of pseudo-labeling of the target distribution constructs a new “source” domain (i.e. target distribution with pseudo-labels). The sequence of optimization problems can then be seen as a sequence of non-conservative domain adaptation problems in which but , where and is the true conditional label distribution in the target domain. Since is strictly zero in this sequence of optimization problems, domain adversarial training is no longer necessary. Furthermore, if minimization does improve the student classifier, then the gap in Eq. 10 should get smaller each time the source domain is updated.
Experiments
In principle, our method can be applied to any domain adaptation tasks so long as one can define a reasonable notion of neighborhood for virtual adversarial training (Miyato et al., 2016). For comparison against Saito et al. (2017) and French et al. (2017), we focus on visual domain adaptation and evaluate on MNIST, MNIST-M, Street View House Numbers (SVHN), Synthetic Digits (SYN DIGITS), Synthetic Traffic Signs (SYN SIGNS), the German Traffic Signs Recognition Benchmark (GTSRB), CIFAR-10, and STL-10. For non-visual domain adaptation, we evaluate on Wi-Fi activity recognition.
Architecture We use a small CNN for the digits, traffic sign, and Wi-Fi domain adaptation experiments, and a larger CNN for domain adaptation between CIFAR-10 and STL-10. Both architectures are available in Appendix A. For fair comparison, we additionally report the performance of source-only baseline models and demonstrate that the significant improvements are attributable to our proposed method.
Replacing gradient reversal. In contrast to Ganin & Lempitsky (2015), which proposed to implement domain adversarial training via gradient reversal, we follow Goodfellow et al. (2014) and instead optimize via alternating updates to the discriminator and encoder (see Appendix C).
Instance normalization. We explored the application of instance normalization as an image pre-processing step. This procedure makes the classifier invariant to channel-wide shifts and rescaling of pixel intensities. A discussion of instance normalization for domain adaptation is provided in Appendix D. We show in Figure 3 the effect of applying instance normalization to the input image.
Hyperparameters. For each task, we tuned the four hyperparameters by randomly selecting labeled target samples from the training set and using that as our validation set. We observed that extensive hyperparameter-tuning is not necessary to achieve state-of-the-art performance. In all experiments with instance-normalized inputs, we restrict our hyperparameter search for each task to . We fixed . Note that the decision to turn on or off that can often be determined a priori. A complete list of the hyperparameters is provided in Appendix B.
2 Model Evaluation
MNIST MNIST-M. We first evaluate the adaptation from MNIST to MNIST-M. MNIST-M is constructed by blending MNIST digits with random color patches from the BSDS500 dataset.
MNIST SVHN. The distribution shift is exacerbated when adapting between MNIST and SVHN. Whereas MNIST consists of black-and-white handwritten digits, SVHN consists of crops of colored, street house numbers. Because MNIST has a significantly lower intrinsic dimensionality that SVHN, the adaptation from MNIST SVHN is especially challenging when the input is not pre-processed via instance normalization. When instance normalization is applied, we achieve a strong state-of-the-art performance and an equally impressive margin-of-improvement over source-only of . Interestingly, by reducing the refinement interval and taking noisier natural gradient steps, we were occasionally able to achieve accuracies as high as . However, due to the high-variance associated with this, we omit reporting this configuration in Table 1.
SYN DIGITS SVHN. The adaptation from SYN DIGITS SVHN reflect a common adaptation problem of transferring from synthetic images to real images. The SYN DIGITS dataset consist of images generated from Windows fonts by varying the text, positioning, orientation, background, stroke color, and the amount of blur.
SYN SIGNS GTSRB. This setting provides an additional demonstration of adapting from synthetic images to real images. Unlike SYN DIGITS SVHN, SYN SIGNS GTSRB contains 43 classes instead of 10.
STL CIFAR. Both STL-10 and CIFAR-10 are 10-class image datasets. These two datasets contain nine overlapping classes. Following the procedure in French et al. (2017), we removed the non-overlapping classes (“frog” and “monkey”) and reduce to a 9-class classification problem. We achieve state-of-the-art performance in both adaptation directions. In STL CIFAR, we achieve a margin-of-improvement and a performance accuracy of . Note that because STL-10 contains a very small training set, it is difficult to estimate the conditional entropy, thus making DIRT-T unreliable for CIFAR STL.
Wi-Fi Activity Recognition. To evaluate the performance of our models on a non-visual domain adaptation task, we applied VADA and DIRT-T to the Wi-Fi Activity Recognition Dataset (Yousefi et al., 2017). The Wi-Fi Activity Recognition Dataset is a classification task that takes the Wi-Fi Channel State Information (CSI) data stream as input to predict motion activity within an indoor area as output . Domain adaptation is necessary when the training and testing data are collected from different rooms, which we denote as Rooms A and B. Table 2 shows that VADA significantly improves classification accuracy compared to Source-Only and DANN by and respectively. However, DIRT-T does not lead to further improvements on this dataset. We perform experiments in Appendix F which suggests that VADA already achieves strong clustering in the target domain for this dataset, and therefore DIRT-T is not expected to yield further performance improvement.
Overall. We achieve state-of-the-art results across all tasks. For a fairer comparison against ATT and the -model, Table 3 provides the improvement margin over the respective source-only performance reported in each paper. In four of the tasks (MNIST MNIST-M, SVHN MNIST, MNIST SVHN, STL CIFAR), we achieve substantial margin of improvement compared to previous models. In the remaining three tasks, our improvement margin over the source-only model is competitive against previous models. Our closest competitor is the -model. However, unlike the -model, we do not perform data augmentation.
It is worth noting that DIRT-T consistently improves upon VADA. Since DIRT-T operates by incrementally pushing the decision boundaries away from the target domain data, it relies heavily on the cluster assumption. DIRT-T’s empirical success therefore demonstrates the effectiveness of leveraging the cluster assumption in unsupervised domain adaptation with deep neural networks.
3 Analysis of VADA and DIRT-T
3.2 Role of Teacher Model in DIRT-T
When considering Eq. 14, it is natural to ask whether defining the neighborhood with respect to the classifier is truly necessary. In Figure 4, we demonstrate in SVHN MNIST and STL CIFAR that removal of the KL-term negatively impacts the model. Since the MNIST data manifold is low-dimensional and contains easily identifiable clusters, applying naive gradient descent (Eq. 12) can also boost the test accuracy during initial training. However, without the KL constraint, the classifier can sometimes deviate significantly from the neighborhood of the previous classifier, and the resulting spikes in the KL-term correspond to sharp drops in target test accuracy. In STL CIFAR, where the data manifold is much more complex and contains less obvious clusters, naive gradient descent causes immediate decline in the target test accuracy.
3.3 Visualization of Representation
We further analyze the behavior of VADA and DIRT-T by showing T-SNE embeddings of the last hidden layer of the model trained to adapt from MNIST SVHN. In Figure 5, source-only training shows strong clustering of the MNIST samples (blue) and performs poorly on SVHN (red). VADA offers significant improvement and exhibits signs of clustering on SVHN. DIRT-T begins with the VADA initialization and further enhances the clustering, resulting in the best performance on MNIST SVHN.
4 Domain Adversarial Training: Layer Ablation
In Table 5, we applied domain adversarial training to various layers of a Domain Adversarial Neural Network (Ganin & Lempitsky, 2015) trained to adapt MNIST SVHN. With the exception of layers and , which experienced training instability, the general observation is that as the layer gets deeper, the additional capacity of the corresponding embedding function allows better matching of the source and target distributions without hurting source generalization accuracy. This demonstrates that the combination of low divergence and high source accuracy does not imply better adaptation to the target domain. Interestingly, when the classifier is regularized to be locally-Lipschitz via VADA, the combination of low divergence and high source accuracy appears to correlate more strongly with better adaptation.
Conclusion
In this paper, we presented two novel models for domain adaptation inspired by the cluster assumption. Our first model, VADA, performs domain adversarial training with an added term that penalizes violations of the cluster assumption. Our second model, DIRT-T, is an extension of VADA that recursively refines the VADA classifier by untethering the model from the source training signal and applying approximate natural gradients to further minimize the cluster assumption violation. Our experiments demonstrate the effectiveness of the cluster assumption: VADA achieves strong performance across several domain adaptation benchmarks, and DIRT-T further improves VADA performance. Our proposed models open up several possibilities for future work. One possibility is to apply DIRT-T to weakly supervised learning; another is to improve the natural gradient approximation via K-FAC (Martens & Grosse, 2015) and PPO (Schulman et al., 2017). Given the strong performance of our models, we also recommend them for other downstream domain adaptation applications.
We gratefully acknowledge funding from Adobe, NSF (grants #1651565, #1522054, #1733686), Toyota Research Institute, Future of Life Institute, and Intel. We also thank Daniel Levy, Shengjia Zhao, and Jiaming Song for insightful discussions, and the anonymous reviewers for their helpful comments and suggestions.
References
Appendix A Architectures
Appendix B Hyperparameters
We observed that extensive hyperparameter-tuning is not necessary to achieve state-of-the-art performance. To demonstrate this, we restrict our hyperparameter search for each task to , in all experiments with instance-normalized inputs. We fixed . Note that the decision to turn on or off that can often be determined a priori based on prior belief regarding the extent to covariate shift. In the absence of such prior belief, a reliable choice is .
When the target domain is MNIST/MNIST-M, the task is sufficiently simple that we only allocate iterations to each optimization problem in Eq. 14. In all other cases, we set the refinement interval . We apply Adam Optimizer (learning rate ) with Polyak averaging (more accurately, we apply an exponential moving average with momentum to the parameter trajectory). VADA was trained for iterations and DIRT-T takes VADA as initialization and was trained for iterations, with number of iterations chosen as hyperparameter.
Appendix C Replacing Gradient Reversal
We note from Goodfellow et al. (2014) that the gradient of is tends to have smaller norm than during initial training since the latter rescales the gradient by . Following this observation, we replace the gradient reversal procedure with alternating minimization of
The choice of using gradient reversal versus alternating minimization reflects a difference in choice of approximating the mini-max using saturating versus non-saturating optimization (Fedus et al., 2017). In some of our initial experiments, we observed the replacement of gradient reversal with alternating minimization stabilizes domain adversarial training. However, we encourage practitioners to try either optimization strategy when applying VADA.
Appendix D Instance Normalization for Domain Adaptation
Theorem 1 suggests that we should identify ways of constraining the hypothesis space without hurting the global optimal classifier for the joint task. We propose to further constrain our model by introducing instance normalization as an image pre-processing step for the input data. Instance normalization was proposed for style transfer Ulyanov et al. (2016) and applies the operation
For visual data the application of instance normalization to the input layer makes the classifier invariant to channel-wide shifts and scaling of the pixel intensities. For most visual tasks, sensitivity to channel-wide pixel intensity changes is not critical to the success of the classifier. As such, instance normalization of the input may help reduce without hurting the globally optimal classifier. Interestingly, Figure 3 shows that input instance normalization is not equivalent to gray-scaling, since color is partially preserved. To test the effect of instance normalization, we report results both with and without the use of instance-normalized inputs.
Appendix E Limitation of Domain Adversarial Training
For a joint distribution , we denote the generalization error of a classifier as
In a slight abuse of notation, we define the generalization error with respect to as
such that generalization error under the distribution is minimized.
Domain adversarial training seeks to find a single classifier used for both the source and target distributions. To do so, domain adversarial training sets up the objective
where and are the hypothesis spaces for the embedding function and embedding classifier. Intuitively, domain adversarial training operates under the hypothesis that good source generalization error in conjunction with source-target feature matching implies good target generalization error. We shall see, however, that if and is sufficiently complex, this implication does not necessarily hold.
Such a set of classifiers satisfies the feature-matching constraint while achieving source generalization error no worse than the optimal source-domain hard classifier. It suffices to show that includes hypotheses that perform poorly in the target domain.
We first show is not an empty set by constructing an element of this set. Choose a partitioning where
Let . It follows that the composite classifier is an element of .
Next, we show that a classifier does not necessarily achieve good target generalization error. Consider the partitioning which solves the following optimization problem
Such a partitioning is the worst-case partitioning subject to the probability mass constraint. It follows that worse case has generalization error
no better than the worst-case partitioning of the target domain.
E.2 Connection to Theorem 1
A justification for domain adversarial training is that the -divergence term is smaller than the -divergence, thus yielding a tighter upper bound for Theorem 1. However, we shall see that the -divergence term is in fact maximal.
Let . It follows that the composite classifiers and are elements of .
From the definition of , we see that
The -divergence thus achieves the maximum value of .
E.3 Implications
Our analysis assumes infinite capacity embedding functions and the ability to solve optimization problems exactly. The empirical success of domain adversarial training suggests that the use of finite-capacity convolutional neural networks combined with stochastic gradient-based optimization provides the necessary regularization for domain adversarial training to work. The theoretical characterization of domain adversarial training in the case finite-capacity convolutional neural networks and gradient-based learning remains a challenging but important open research problem.
Appendix F Non-Visual Domain Adaptation Task
To evaluate the performance of our models on a non-visual domain adaptation task, we applied VADA and DIRT-T to the Wi-Fi Activity Recognition Dataset (Yousefi et al., 2017). The Wi-Fi Activity Recognition Dataset is a classification task that takes the Wi-Fi Channel State Information (CSI) data stream as input to predict motion activity within an indoor area as output . The dataset collected the CSI data stream samples associated with seven activities, denoted as “bed”, “fall”, “walk”, “pick up”, “run”, “sit down”, and “stand up”.
However, the joint distribution over the CSI data stream and motion activity changes depending on the room in which the data was collected. Since the data was collected for multiple rooms, we selected two rooms (denoted here as Room A and Room B) and constructed the unsupervised domain adaptation task by using Room A as the source domain and Room B as the target domain. We compare the performance of DANN, VADA, and DIRT-T on the Wi-Fi domain adaptation task in Table 2, using the hyperparameters .
Table 2 shows that VADA significantly improves classification accuracy compared to Source-Only and DANN. However, DIRT-T does not lead to further improvements on this dataset. We believe this is attributable to VADA successfully pushing the decision boundary away from data-dense regions in the target domain. As a result, further application of DIRT-T would not lead to better decision boundaries. To validate this hypothesis, we visualize the t-SNE embeddings for VADA and DIRT-T in Figure 6 and show that VADA is already capable of yielding strong clustering in the target domain. To verify that the decision boundary indeed did not change significantly, we additionally provide the confusion matrix between the VADA and DIRT-T predictions in the target domain (Fig. 7).