Label Efficient Learning of Transferable Representations across Domains and Tasks

Zelun Luo, Yuliang Zou, Judy Hoffman, Li Fei-Fei

Introduction

Humans are exceptional visual learners capable of generalizing their learned knowledge to novel domains and concepts and capable of learning from few examples. In recent years, computational models based on end-to-end learnable convolutional networks have made significant improvements for visual recognition and have been shown to demonstrate some cross-task generalizations while enabling faster learning of subsequent tasks as most frequently evidenced through fine-tuning .

However, most efforts focus on the supervised learning scenario where a closed world assumption is made at training time about both the domain of interest and the tasks to be learned. Thus, any generalization ability of these models is only an observed byproduct. There has been a large push in the research community to address generalizing and adapting deep models across different domains , to learn tasks in a data efficient way through few shot learning , and to generically transfer information across tasks .

While most approaches consider each scenarios in isolation we aim to directly tackle the joint problem of adapting to a novel domain which has new tasks and few annotations. Given a large labeled source dataset with annotations for a task set, A, we seek to transfer knowledge to a sparsely labeled target domain with a possibly wholly new task set, B. This setting is in line with our intuition that we should be able to learn reusable and general purpose representations which enable faster learning of future tasks requiring less human intervention. In addition, this setting matches closely to the most common practical approach for training deep models which is to use a large labeled source dataset (often ImageNet ) to train an initial representation and then to continue supervised learning with a new set of data and often with new concepts.

In our approach, we jointly adapt a source representation for use in a distinct target domain using a new multilayer unsupervised domain adversarial formulation while introducing a novel cross-domain and within domain class similarity objective. This new objective can be applied even when the target domain has non-overlapping classes to the source domain.

We evaluate our approach in the challenging setting of joint transfer across domains and tasks and demonstrate our ability to successfully transfer, reducing the need for annotated data for the target domain and tasks. We present results transferring from a subset of Google Street View House Numbers (SVHN) containing only digits 0-4 to a subset of MNIST containing only digits 5-9. Secondly, we present results on the challenging setting of adapting from ImageNet object-centric images to UCF-101 videos for action recognition.

Related work

Domain adaptation. Domain adaptation seeks to learn from related source domains a well performing model on target data distribution . Existing work often assumes that both domains are defined on the same task and labeled data in target domain is sparse or non-existent . Several methods have tackled the problem with the Maximum Mean Discrepancy (MMD) loss between the source and target domain. Weight sharing of CNN parameters and minimizing the distribution discrepancy of network activations have also shown convincing results. Adversarial generative models aim at generating source-like data with target data by training a generator and a discriminator simultaneously, while adversarial discriminative models focus on aligning embedding feature representations of target domain to source domain. Inspired by adversarial discriminative models, we propose a method that aligns domain features with multi-layer information.

Transfer learning. Transfer learning aims to transfer knowledge by leveraging the existing labeled data of some related task or domain . In computer vision, examples of transfer learning include which try to overcome the deficit of training samples for some categories by adapting classifiers trained for other categories . With the power of deep supervised learning and the ImageNet dataset , learned knowledge can even transfer to a totally different task (i.e. image classification →\rightarrow object detection ; image classification →\rightarrow semantic segmentation ) and then achieve state-of-the-art performance. In this paper, we focus on the setting where source and target domains have differing label spaces but the label spaces share the same structure. Namely adapting between classifying different category sets but not transferring from classification to a localization plus classification task.

Few-shot learning. Few-shot learning seeks to learn new concepts with only a few annotated examples. Deep siamese networks are trained to rank similarity between examples. Matching networks learns a network that maps a small labeled support set and an unlabeled example to its label. Aside from these metric learning-based methods, meta-learning has also served as a essential part. Ravi et al. propose to learn a LSTM meta-learner to learn the update rule of a learner. Finn et al. tries to find a good initialization point that can be easily fine-tune with new examples from new tasks. When there exists a domain shift, the results of prior few-shot learning methods are often degraded.

Unsupervised learning. Many unsupervised learning algorithms have focused on modeling raw data using reconstruction objectives . Other probabilistic models include restricted Boltzmann machines , deep Boltzmann machines , GANs , and autoregressive models are also popular. An alternative approach, often terms “self-supervised learning” , defines a pretext task such as predicting patch ordering , frame ordering , motion dynamics , or colorization , as a form of indirect supervision. Compared to these approaches, our unsupervised learning method does not rely on exploiting the spatial or temporal structure of the data, and is therefore more generic.

Method

We introduce a semi-supervised learning algorithm which transfers information from a large labeled source domain, S\mathcal{S}, to a sparsely labeled target domain, T\mathcal{T}. The goal being to learn a strong target classifier without requiring the large annotation overhead required for standard supervised learning approaches.

In fact, this setting is very commonly explored for convolutional network (convnet) based recognition methods. When learning with convnets the usual learning procedure is to use a very large labeled dataset (e.g. ImageNet ) for initial training of the network parameters (termed pre-training). The learned weights are then used as initialization for continued learning on new data and for new tasks, called fine-tuning. Fine-tuning has been broadly applied to reduce the number of labeled examples needed for learning new tasks, such as recognizing new object categories after ImageNet pre-training , or learning new label structures such as detection after classficiation pre-training . Here we focus on transfer in the case of a shared label structure (e.g. classification of different category sets).

Unlike standard domain adaptation approaches which transfer knowledge from source to target domains assuming a marginal or conditional distribution shift under a shared label space (YS=YT\mathcal{Y}^{\mathcal{S}}=\mathcal{Y}^{\mathcal{T}}), we tackle joint image or feature space adaptation as well as transfer across semantic spaces. Namely, we consider the case where the source and target label spaces are not equal, YS≠YT\mathcal{Y}^{\mathcal{S}}\neq\mathcal{Y}^{\mathcal{T}}, and even the most challenging case where the sets are non-overlapping, YS∩YT=∅\mathcal{Y}^{\mathcal{S}}\cap\mathcal{Y}^{\mathcal{T}}=\emptyset.

Our approach consists of unsupervised feature alignment between source and target as well as semantic transfer to the unlabeled target data from either the labeled target or the labeled source data. We introduce a new multi-layer domain discriminator which can be used for domain alignment following the recent domain adversarial learning approaches . We next introduce a new semantic transfer learning objective which uses cross category similarity and can be tuned to account for varying size of label set overlap.

Our model jointly optimizes over a target supervised loss, Lsup\mathcal{L}_{\text{sup}}, a domain transfer objective, LDT\mathcal{L}_{\textit{DT}}, and finally a semantic transfer objective, LST\mathcal{L}_{\textit{ST}}. Thus, our total objective can be written as follows:

where the hyperparameters α\alpha and β\beta determine the influence of the domain transfer loss and the semantic transfer loss, respectively. In the following sections we elaborate on our domain and semantic transfer objectives.

2 Multi-layer domain adversarial loss

We define a novel domain alignment objective function called multi-layer domain adversarial loss. Recent efforts in deep domain adaptation have shown strong performance using feature space domain adversarial objectives . These methods learn a target representation such that the target distribution viewed under this model is aligned with the source distribution viewed under the source representation. This alignment is accomplished through an adversarial minimization across domain, analogous to the prevalent generative adversarial approaches . In particular, a domain discriminator, D(⋅)D(\cdot), is trained to classify whether a particular data point arises from the source or the target domain. Simultaneously, the target embedding function Et(xt)E^{t}(\mathbf{x}^{t}) (defined as the application of layers of the network is trained to generate the target representation that cannot be distinguished from the source domain representation by the domain discriminator. Similar to , we consider a representation to be domain invariant if the domain discriminator can not distinguish examples from the two domains.

Prior work considers alignment for a single layer of the embedding at a time and as such learns a domain discriminator which takes the output from the corresponding source and target layers as input. Separately, domain alignment methods which focus on first and second order statistics have shown improved performance through applying domain alignment independently at multiple layers of the network . Rather than learning independent discriminators for each layer of the network we propose a simultaneous alignment of multiple layers through a multi-layer discriminator.

At each layer of our multi-layer domain discriminator, information is accumulated from both the output from the previous discriminator layer as well as the source and target activations from the corresponding layer in their respective embeddings. Thus, the output of each discriminator layer is defined as:

Thus, the following loss functions are proposed to optimize the multi-layer domain discriminator and the embeddings, respectively, according to our domain transfer objective:

where dls,dlt\mathbf{d}_{l}^{s},\mathbf{d}_{l}^{t} are the outputs of the last layer of the source and target multi-layer domain discriminator. Note that these losses are placed after the final domain discriminator layer and the last embedding layer but then produce gradients which back-propagate throughout all relevant lower layer parameters. These two losses together comprise LDTL_{DT}, and there is no iterative optimization procedure involved.

This multi-layer discriminator (shown in Figure 1 - yellow) allows for deeper alignment of the source and target representations which we find empirically results in improved target classification performance as well as more stable adversarial learning.

3 Cross category similarity for semantic transfer

In the previous section, we introduced a method for transferring an embedding from the source to the target domain. However, this only enforces alignment of the global domain statistics with no class specific transfer. Here, we define a new semantic transfer objective, LST\mathcal{L}_{\textit{ST}}, which transfers information from a labeled set of data to an unlabeled set of data by minimizing the entropy of the softmax with temperature of the similarity vector between an unlabeled point and all labeled points. Thus, this loss may be applied either between the source and unlabeled target data or between the labeled and unlabeled target data.

where, H(⋅)H(\cdot) is the information entropy function, σ(⋅)\sigma(\cdot) is the softmax function and τ\tau is the temperature of the softmax. Note that the temperature can be used to directly control the percentage of source examples we expect the target example to be similar to (see Figure 2).

Entropy minimization has been widely used for unsupervised and semi-supervised learning by encouraging low density separation between clusters or classes. Recently this principle of entropy minimization has be applied for unsupervised adaptation . Here, the source and target domains are assumed to share a label space and each unlabeled target example is passed through the initial source classifier and the entropy of the softmax output scores is minimized.

In contrast, we do not assume a shared label space between the source and target domains and as such can not assume that each target image maps to a single source label. Instead, we compute pairwise similarities between target points and the source points (or per class averages of source points ) across the features spaces aligned by our multi-layer domain adversarial transfer. We then tune the softmax temperature based on the expected similarity between the source and target labeled set. For example, if the source and target label set overlap, then a small temperature will encourage each target point to be very similar to one source class, whereas a larger temperature will allow for target points to be similar to multiple source classes.

For semantic transfer within the target domain, we utilize the metric-based cross entropy loss between labeled target examples to stabilize and improve the learning. For a labeled target example, in addition to the traditional cross entropy loss, we also calculate a metric-based cross entropy loss We refer this as ”metric-based” to cue the reader that this is not a cross entropy within the label space.. Assume we have kk labeled examples from each class in the target domain. We compute the embedding for each example and then the centroid ciTc_{i}^{\mathcal{T}} of each class in the embedding space. Thus, we can compute the similarity vector for each labeled example, where the ithi^{th} element is the similarity between this labeled example and the centroid of each class: [vt(xt)]i=ψ(xt,ciT)[v_{t}(\mathbf{x}^{t})]_{i}=\psi(\mathbf{x}^{t},c_{i}^{\mathcal{T}}). We can then calculate the metric based cross entropy loss:

Similar to the source-to-target scenario, for target-to-target we also have the unsupervised part,

With the metric-based cross entropy loss, we introduce the constraint that the target domain data should be similar in the embedding space. Also, we find that this loss can provide a guidance for the unsupervised semantic transfer to learn in a more stable way. LST\mathcal{L}_{ST} is the combination of LST,unsupervised\mathcal{L}_{ST,\text{unsupervised}} from source-target (Equation 5), LST,supervised\mathcal{L}_{ST,\text{supervised}} from source-target (Equation 6), and LST,unsupervised\mathcal{L}_{ST,\text{unsupervised}} from target-target (Equation 7), i.e.,

Experiment

This section is structured as follows. In section 4.1, we show that our method outperform fine-tuning approach by a large margin, and all parts of our method are necessary. In section 4.2, we show that our method can be generalized to bigger datasets. In section 4.3, we show that our multi-layer domain adversarial method outperforms state-of-the-art domain adversarial approaches.

Datasets We perform adaptation experiments across two different paired data settings. First for adaptation across different digit domains we use MNIST and Google Street View House Numbers (SVHN) . The MNIST handwritten digits database has a training set of 60,000 examples, and a test set of 10,000 examples. The digits have been size-normalized and centered in fixed-size images. SVHN is a real-world image dataset for machine learning and object recognition algorithms with minimal requirement on data preprocessing and formatting. It has 73,257 digits for training, 26,032 digits for testing. As our second experimental setup, we consider adaptation from object centric images in ImageNet to action recognition in video using the UCF-101 dataset. ImageNet is a large benchmark for the object classification task. We use the task 1 split from ILSVRC2012. UCF-101 is an action recognition dataset collected on YouTube. With 13,320 videos from 101 action categories, UCF-101 provides a large diversity in terms of actions and with the presence of large variations in camera motion, object appearance and pose, object scale, viewpoint, cluttered background, illumination conditions, etc.

Implementation details We pre-train the source domain embedding function with cross-entropy loss. For domain adversarial loss, the discriminator takes the last three layer activations as input when the number of output classes are the same for source and target tasks, and takes the second last and third last layer activations when they are different. The similarity score is chosen as the dot product of the normalized support features and the unnormalized target feature. We use the temperature τ=2\tau=2 for source-target semantic transfer and τ=1\tau=1 for within target transfer as the label space is shared. We use α=0.1\alpha=0.1 and β=0.1\beta=0.1 in our objective function. The network is trained with Adam optimizer and with learning rate 10−310^{-3}. We conduct all the experiments with the PyTorch framework.

Experimental setting. In this experiment, we define three datasets: (i) labeled data in source domain D1\mathcal{D}_{1}; (ii) few labeled data in target domain D2\mathcal{D}_{2}; (iii) unlabeled data in target domain D3\mathcal{D}_{3}. We take the training split of SVHN dataset as dataset D1\mathcal{D}_{1}. To fairly compare with traditional learning paradigm and episodic training, we subsample kk examples from each class to construct dataset D2\mathcal{D}_{2} so that we can perform traditional training or episodic (k−1k-1)-shot learning. We experiment with k=2,3,4,5k=2,3,4,5, which corresponds to 10,15,20,2510,15,20,25 labeled examples, or 0.017%,0.025%,0.333%,0.043%0.017\%,0.025\%,0.333\%,0.043\% of the total training data respectively. Since our approach involves using annotations from a small subset of the data, we randomly subsample 1010 different subsets {D2i}i=110\{\mathcal{D}_{2}^{i}\}_{i=1}^{10} from the training split of MNIST dataset, and use the remaining data as {D3i}i=110\{\mathcal{D}_{3}^{i}\}_{i=1}^{10} for each kk. Note that source domain and target domain have non-overlapping classes: we only utilize digits -44 in SVHN, and digits 55-99 in MNIST.

Baselines and prior work. We compare against six different methods: (i) Target only: the model is trained on D2\mathcal{D}_{2} from scratch; (ii) Fine-tune: the model is pretrained on D1\mathcal{D}_{1} and fine-tuned on D2\mathcal{D}_{2}; (iii) Matching networks : we first pretrain the model on D3\mathcal{D}_{3}, then use D2\mathcal{D}_{2} as the support set in the matching networks; (iv) Fine-tuned matching networks: same as baseline iii, except that for each kk the model is fine-tuned on D2\mathcal{D}_{2} with 5-way (k−1k-1)-shot learning: k−1k-1 examples in each class are randomly selected as the support set, and the last example in each class is used as the query set; (v) Fine-tune + adversarial: in addition to baseline ii, the model is also trained on D1\mathcal{D}_{1} and D3\mathcal{D}_{3} with a domain adversarial loss; (vi.) Full model: fine-tune the model with the proposed multi-layer domain adversarial loss.

Results and analysis. We calculate the mean and standard error of the accuracies across 1010 sets of data, which is shown in Table 1. Due to domain shift, matching networks perform poorly without fine-tuning, and fine-tuning is only marginally better than training from scratch. Our method with multi-layer adversarial only improves the overall performance, but is more sensitive to the subsampled data. Our method achieves significant performance gain, especially when the number of labeled examples is small (k=2k=2). For reference, fine-tuning on full target dataset gives an accuracy of 99.65%99.65\%.

2 Image object recognition →→\rightarrow video action recognition

Problem analysis. Many recent works study the domain shift between images and video in the object detection settings. Compared to still images, videos provide several advantages: (i) motion provides information for foreground vs background segmentation ; (ii) videos often show multiple views and thus provide 3D information. On the other hand, video frames usually suffer from: (i) motion blur; (ii) compression artifacts; (iii) objects out-of-focus or out-of-frame.

Experimental setting. In this experiment, we focus on three dataset splits: (i) ImageNet training set as the labeled data in source domain D1\mathcal{D}_{1}; (ii) kk video clips per class randomly sampled from UCF-101 training as the few labeled data in target domain set D2\mathcal{D}_{2}; (iii) the remaining videos in UCF-101 training set as the unlabeled data in target domain D3\mathcal{D}_{3}. We experiment with k=3,5,10k=3,5,10, which corresponds 303,505,1010303,505,1010 video clips, or 2.27%,3.79%,7.59%2.27\%,3.79\%,7.59\% of the total training data respectively. Each experiment is run 33 times on D1\mathcal{D}_{1}, {D2i}i=13\{\mathcal{D}_{2}^{i}\}_{i=1}^{3}, and {D3i}i=13\{\mathcal{D}_{3}^{i}\}_{i=1}^{3}.

Baselines and prior work. We compare our method with two baseline methods: (i) Target only: the model is trained on D2\mathcal{D}_{2} from scratch; (ii) Fine-tune: the model is first pre-trained on D1\mathcal{D}_{1}, then fine-tuned on D2\mathcal{D}_{2}. For reference, we report the performance of a fully supervised method .

Results and analysis. The accuracy of each model is shown in Table 2. We also fine-tune a model with all the labeled data for comparison. Per-frame performance (img) and average-across-frame performance (vid) are both reported. Note that we calculate the average-across-frame performance by averaging the softmax score of each frame in a video. Our method achieves significant improvement on average-across-frame performance over standard fine-tuning for each value of kk. Note that compared to fine-tuning, our method has a bigger gap between per-frame and per-video accuracy. We believe that this is due to the semantic transfer: our entropy loss encourages a sharper softmax variance among per-frame softmax scores per video (if the variance is zero, then per-frame accuracy = per-video accuracy). By making more confident predictions among key frames, our method achieves a more significant gain with respective to per-video performance, even when there is little change in the per-frame prediction.

3 Ablation: unsupervised domain adaptation

To validate our multi-layer domain adversarial loss objective, we conduct an ablation experiment for unsupervised domain adaptation. We compare against multiple recent domain adversarial unsupervised adaptation methods. In this experiment, we first pretrain a source embedding CNN on the training split SVHN and then adapt the target embedding for MNIST by performing adversarial domain adaptation. We evaluate the classification performance on the test split of MNIST . We follow the same training strategy and model architecture for the embedding network as .

All the models here have a two-step training strategy and share the first stage. ADDA optimizes encoder and classifier simultaneously. We also propose a similar method, but optimize encoder only. Only we try a model with no classifier in the last layer (i.e. perform domain adversarial training in feature space). We choose γ=0.1\gamma=0.1 as the decay factor for this model.

The accuracy of each model is shown in Table 3. We find that our method achieve 6.5%6.5\% performance gain over the best competing domain adversarial approach indicating that our multilayer objective indeed contributes to our overall performance. In addition, in our experiments, we found that the multilayer approach improved overall optimization stability, as evidenced in our small standard error.

Conclusion

In this paper, we propose a method to learn a representation that is transferable across different domains and tasks in a data efficient manner. The framework is trained jointly to minimize the domain shift, to transfer knowledge to new task, and to learn from large amounts of unlabeled data. We show superior performance over the popular fine-tuning approach. We hope to keep improving the method in future work.

Acknowledgement

We would like to start by thanking our sponsors: Stanford Computer Science Department and Stanford Program in AI-assisted Care (PAC). Next, we specially thank De-An Huang, Kenji Hata, Serena Yeung, Ozan Sener and all the members of Stanford Vision and Learning Lab for their insightful discussion and feedback. Lastly, we thank all the anonymous reviewers for their valuable comments.

References

Network Architecture

(b) Image object recognition →→\rightarrow video action recognition

Embedding network structure: ResNet-18 We refer readers to the PyTorch implementation: https://github.com/pytorch/vision/blob/master/torchvision/models/resnet.py.

(c) Ablation: unsupervised domain adaptation