Energy-constrained Self-training for Unsupervised Domain Adaptation
Xiaofeng Liu, Bo Hu, Xiongchang Liu, Jun Lu, Jane You, Lingsheng Kong
I Introduction
Deep neural networks are usually data-starved and rely on the assumption of training and testing data . However, in reality, the deployment target tasks are usually significantly diverse, and collecting labeled data in the target domain is expensive or even prohibitive. For instance, densely annotating a Cityscapes image on average takes about 90 minutes , which severely hinders the generalization of an autonomous driving system in different cities.
Therefore, the unsupervised domain adaptation (UDA) seeks to transfer knowledge from one labeled source domain to another target domain with the aid of unlabeled target data . Recently, one of the promising methods in UDA is the self-training , which iteratively generates a set of one-hot pseudo-labels in the target domain, and then retrains network based on these pseudo-labels with target data. Previous research of deep self-training usually adopts the one-hot pseudo-label, and evidence that it is essentially an entropy minimization process that pushing network output to be as sharp as hard pseudo-label.
However, the correctness of pseudo-labels cannot be guaranteed. Trusting all selected pseudo-labels as “ground truth” by encoding them as hard labels can lead to overconfident mistakes and propagated errors . Actually, the state-of-the-art accuracy of many UDA benchmarks are just around 50%. Especially at the first few epochs, it is hard to produce reliable pseudo-label. Besides, the labels of natural images can also be highly ambiguous. Taking a sample image from VisDA17 (see Fig. 1) as an example, both person and car dominate significant portions of this image. Enforcing a model to be very confident in only one of the classes during training can also hurt the learning behavior , particularly within the context of no ground truth label for the target samples.
The manually defined pseudo label smoothing or entropy minimization are proposed to make the network more conservative . Essentially, our previous work is modifying the inaccurate pseudo label histogram distribution to be more smooth following a fixed smoothing operation for any data. For instance, revising the three-classes one-hot pseudo label to or . However, these constraints for pseudo-label can not be adaptively adjusted for different inputs or network parameters.
The aforementioned issues motivate us to introduce a regularization signal for the target sample that depends on the input and the network parameters at the present training iteration, and not related to the inaccurate pseudo label.
Recent study points out that a standard discriminative classifier can be essentially reinterpreted as an energy-based model (EBM) for the joint distribution . Both the energy-based model and supervised discriminative model can be benefited from simultaneously optimizing both with cross-entropy (CE) loss and optimizing with EBM . Here, we demonstrate that optimizing can be even more promising on self-training-based UDA. Since we do not have reliable labels on target domain, and can be an idea regularization signal, which is correlated with the input and network parameter, and independent to pseudo label .
Therefore, we propose a simple and straightforward idea that configures the energy minimization of data point (i.e., ) as an additional regularization term. The target domain examples are optimized with pseudo-label CE loss and EBM objective. With the help of the energy-based model, our self-training is expected to be more controllable.
In this paper, we propose a novel and intuitive framework to incorporate the EBM into the self-training UDA as pseudo label-irrelevant adaptive regularization signal. We empirically validate the effectiveness and generality of the proposed method on multiple challenging benchmarks (classification and semantic segmentation) and achieve state-of-the-art performance.
II Related Works
Unsupervised domain adaptation (UDA) with deep networks targets to learn domain invariant embeddings by minimizing the cross-domain difference of feature distributions with certain criteria . Examples of these methods include maximum mean discrepancy (MMD), deep Correlation Alignment (CORAL), sliced Wasserstein discrepancy, adversarial learning at input-level, feature level, output space level, etc . Despite the underlying difference, there exists an interesting connection between some of these methods with conditional formsE.g., class-wise adversarial learning, or discriminators taking network predictions as input. and Self-training as they can be broadly considered as EM algorithms, and such conditional formulation has been widely proved to benefit the adaptation.
Self-training was initially investigated in semi-supervised learning . A subtle difference between self-training with fixed feature input and deep self-training is that the latter involves the learning of deep embedding which renders greater flexibility towards domain alignment than classifier-level adaptation only. Recently, there have been multiple deep self-training/pseudo-label based methods that are proposed for UDA .
Considering the noisy pseudo-label, our previous work proposes to construct a more conservative pseudo-label that smoothing the one-hot distribution or regularize it with the entropy. The solution proposes in this paper is orthogonal with , which resorts to the additional supervision signal of EBM that independent of pseudo-label. Compared with the manually defined label smoothing in , the energy-constraint can adaptively regularize the training w.r.t. the input and the present network parameters. We note that our EBM regularization can be simply added on state-of-the-art self-training methods following a plug-and-play fashion.
Energy-Based Models (EBMs) capture dependencies between variables by associating a scalar energy to each configuration of the variables . Recently, approximated the expectation of the log-likelihood for a single example with respect to network parameter using a sampler based on Stochastic Gradient Langevin Dynamics (SGLD) .
Targeting on the combination of classifier and EBMs, reinterpret the logits to define a class-conditional EBM , which require additional parameters to be learned to derive a classifier and an unconditional model. is similar as well but is trained using a GAN-like generator and is applied to different applications. Based on , scales the training of EBMs to high-dimensional data with Contrastive Divergence and SGLD.
Here, we regard EBMs as the pseudo label-irrelevant optimization constrains and explore its potentials on domain adaptation, especially UDA. This helps realize the potential of energy-based models on downstream supervised discriminative problems. Moreover, the thorough convergence analysis and theoretically connection with CEM are investigated.
III Methodology
The recent study reveals that a standard classifier can be interpreted as an energy based model of joint distribution , and optimizing the likelihood can be helpful for both the discrimination and generation task. Specifically, optimizing is simply achieved by using the standard cross-entropy loss as conventional classification. The additional optimization objective has been proven and evidenced that can improve the confidence calibration and robustness for conventional classification task .
Considering the target samples do not have ground truth label, the self-training methods utilize the inaccurate pseudo label to calculate the cross-entropy loss. Therefore, optimizing can potentially be more helpful for UDA setting. Actually, is adaptive w.r.t. the input and network parameter , and irrelevant to the inaccurate pseudo label, which can be an ideal regularizer of self-training based UDA.
Considering can be approximated with , it is possible to modeling the energy function instead of . Following , we can define an EBM of the joint distribution , by defining . By marginalizing out , we have . Considering , the energy function of can be
In this setting, . The normalization constant will be canceled out and yielding the standard softmax function, which bridges the EMB and conventional classifiers.
Therefore, we can simultaneously optimize with standard cross-entropy loss, and optimize with Stochastic Gradient Langevin Dynamics (SGLD), where gradients are taken with respect to .
Self-training for UDA is an iterative loss minimization framework , which regards pseudo-labels as learnable latent variables. The optimization objective for unlabeled target example is the cross-entropy with one-hot or smoothed pseudo-label. Therefore, a straightforward solution for adapting EBM to UDA is to incorporate Eq. 1 as a regularization term. For the target sample, it not only needs to achieve a good prediction of pseudo-label, but also minimize the energy of . Therefore, our energy-constraint is essentially depended on and , which is more flexible than pre-defined label smoothing or entropy regularization . Considering that the pseudo-label is usually noisy, the latter objective is expected to be even more important than the setting of supervised learning.
Following the formulation in our CRST , the self-training with EBM regularization (R-EBM) for target sample, i.e., , can be formulated as
For each class , is determined by the confidence value selecting the most confident portion of class predictions in the entire target set . If a sample’s predication is relatively confident with , it is selected and labeled as class . The less confident ones with are not selected. The same class-balanced strategy introduced in is adopted for all self-training methods in this work.
The feasible set is the union of and a probability simplex . is a balancing hyper-parameter of the regularization term, which does not directly relate to the label . The self-training can be solved by an alternating optimization scheme.
Step 1) Pseudo-label generation Fix and solve:
For solving step 1), there is a global optimizer for arbitrary as :
Despite a long period of little development, there has been recent work using the sampler based on SGLD to train the large-scale EBMs on high-dimensional data, parameterized by deep neural networks.
Given that the goal of our work is to incorporate EBM training into the standard classification setting, the classification part is the same as , and we only change the regularizer. Therefore, we can follow to train the network with both cross-entropy loss and SGLD to ensure this distribution is being optimized with an unbiased objective. Similar to we also adopt the contrastive divergence to estimate the expectation of derivative of the log-likelihood for a single example with respect to . Since it gives an order of magnitude savings in computation compared to seeding new chains at each iteration as in .
Step 2) Network retraining Fix and minimize
w.r.t. . Carrying out step 1) and 2) for one time is defined as one round in self-training.
IV Experiments
In this section, we provide comprehensive evaluations of the proposed EBM objective both on image classification and semantic segmentation UDA tasks. We implement our methods using the PyTorch toolbox .
The VisDA17 benchmark is a 12-class UDA classification problem. We follow the standard protocol in where the source domain is the training set including synthetic 2D images and the target domain is the validation set including real images from COCO dataset.
To make a fair comparison with other methods, we use the same backbone network (e.g., ResNet101 or ResNet152) for VisDA17. Both networks are pre-trained in ImageNet as previous works and fine-tuned in source domain by SGD with a fixed learning rate , weight decay , momentum and batch size . As shown in Fig. 2, the performance is not sensitive to when , we simply choose in all settings.
We present the results on VisDA 17 in Table I in terms of per-class accuracy and mean accuracy. For each proposed approach, we independently run 5 times and the average and standard deviation of the accuracy metrics are reported.
Class-balanced self-training (CBST) and Conservative regularized self-training (CRST) are the state-of-the-art self-training UDA methods. The CRST+ denote the CRST with EBM regularization.
Benefited by the additional EMB objective, CBST/CRST + can outperform CBST/CRST significantly, and outperforms the recent UDA methods other than self-training by a large margin. We note that the adversarial training can further boost the performance of self-training . Our energy constraint can adaptively regularize the training w.r.t. the input and the present network parameters, which is more flexible than the manually pre-defined label smoothing .
More appealingly, CRST+ achieves the better performance than the recent adversarial training , dropout and moment matching methods, which revokes the potential of self-training in UDA.
The more powerful backbones have also been applied and show better results . The EBM objective with ResNet152 backbone outperforms the other state-of-the-arts . It also demonstrate the flexibility of for different backbones.
Actually, our can be orthogonal with the recent progress of self-training based UDA. CBST is a pioneer of vanilla adversarial UDA, and its performance on VisDA17 is reported on . The label smoothing or the entropy regularization used in CRST or the vanilla self-training UDA can be simply add-on our to further improve the performance. CRST+ obtains better or competitive performances in all settings, even compared with a generative pixel-level domain adaptation method GTA, which is a very complex algorithm in both architecture and objectives.
IV-B UDA for Semantic Segmentation
The semantic segmentation is essentially making the pixel-wise classification. We consider the challenging segmentation adaptation settings in GTA5 to Cityscapes (19 classes are shared). GTA5 dataset includes 24,966 annotated images with size 1,0521,914, which rendered by the GTA5 game engine. Following the standard protocols , we use the full set of GTA5 and adapt the model to the Cityscapes train set with images. In testing, we evaluate on the Cityscapes validation set with images.
To make a fair comparison with other methods, we use the ResNet101 as the backbone network as . Noticing that in PSPNET , Wide ResNet38 is a stronger basenet than ResNet101. The basenet is pre-trained in ImageNet and fine-tuned in source domain by SGD with learning rate , weight decay , momentum , batch size , patch size and data augmentation of multi-scale training () and horizontal flipping. All results on this dataset in the main paper are unified to report the mIoU of the models at the end of the epoch.
We compare CBST/CRST+ with the other methods in Table II. Based on the previous self-training UDA methods CBST or CRST, our additional EBM objective outperforms CBST or CRST by about 2% w.r.t. the mean IoUs.
We achieve a new state-of-the-art, even compared with the generative pixel-level domain adaptation method GTA , which is a relatively complex algorithm in both architecture and objectives.
Introducing the EBM objective to self-training methods can significantly improve the segmentation performance. Consistently with the classification, can be better or comparable with adversarial learning, indicating the self-training can still be a powerful methodology in UDA. Note that moment matching-based methods are usually not well scaleable to segmentation.
V Conclusions
In this paper, we propose to introduce the energy based model’s objective into the self-training unsupervised domain adaptation. Considering the lack of target label, we resort to the pseudo-label which is usually noisy. The EBM optimization objective provides additional signal that is independent of pseudo-label. It can be even more promising in UDA setting than supervised learning. It can be a regularization term. Our solution is orthogonal with the recent progress of self-training and can be added on in a plug-and-play manner without introducing large computation and the change of network structure. Extensive experiments on both UDA classification and semantic segmentation evidenced its effectiveness and generality. The more advanced EBM training and self-training methods can also be adopted to further improve the results.
VI Acknowledgements
This work was supported by the Jangsu Youth Programme [BK20200238], National Natural Science Foundation of China, Younth Programme [grant number 61705221], NIH [NS061841, NS095986], Fanhan Technology, and Hong Kong Government General Research Fund GRF (Ref. No.152202/14E) are greatly appreciated.