Learning Attributes Equals Multi-Source Domain Generalization

Chuang Gan, Tianbao Yang, Boqing Gong

Introduction

Visual attributes are middle-level concepts which humans use to describe objects, human faces, scenes, activities, and so on (e.g., four-legged, smiley, outdoor, and crowded). A major appeal of attributes is that they are not only human-nameable but also machine-detectable, making it possible to serve as the building blocks to describe instances , teach machines to recognize previously unseen classes by zero-shot learning , or offer a natural human-computer interactions channel for image/video search .

However, we contend that the long-standing pursuit after utilizing attributes for various computer vision problems has left the most basic problem—how to accurately and robustly detect attributes from images or videos—far from being solved. Especially, the existing work rarely explicitly tackles the need that attribute detectors should generalize well across different categories, including those previously unseen ones. For instance, the attribute detector “four-legged” is expected to correctly tell a giant panda is four-legged even if it is trained from the images of horses, cows, zebras, and pigs (i.e., no pandas).

Indeed, most of the existing attribute detectors are built using features engineered or learned for object recognition together with off-shelf machine learning classifiers—without tailoring them to reflect the idiosyncrasies of attributes. This is suboptimal; the successful techniques on object recognition do not necessarily apply to attributes learning mainly for two reasons. First, attributes are in a different semantic space as opposed to objects; they are in the middle of low-level visual cues and the high-level object labels. Second, attribute detection can even be considered as an orthogonal task to object recognition, in that attributes are shared by different objects (e.g., zebras, lions, and mice are all “furry”) and distinctive attributes are present in the same object (e.g., a car is boxy and has wheels). As shown in Figure 1, the boundaries between attributes and between object categories cross each other. Therefore, we do not expect that the features originally learned for separating elephant, sheep, and giraffe could also be optimal for detecting the attribute “bush”, which is shared by them.

In this paper, we propose to re-examine the fundamental attribute detection problem and aim to develop an attribute-oriented feature representation, such that one can conveniently apply off-shelf classifiers to obtain high-quality attribute detectors. We expect that the detectors learned from our new representation are capable of breaking the boundaries of object categories and generalizing well across both seen and unseen classes. To this end, we cast attribute detection as a multi-source domain generalization problem by noting that the desired properties from attributes are analogous to the objective of the latter.

Particularly, a domain refers to an underlying data distribution. Multi-source domain generalization aims to extract knowledge from several related source domains such that it is applicable to different domains, especially to those unseen at the training stage. This is in accordance with our objective for learning cross-category generalizable attributes detectors, if we consider each category as a distinctive domain.

Motivated by this observation, we employ the Unsupervised Domain-Invariant Component Analysis (UDICA) as the basic building block for our approach. The key principle of UDICA is that minimizing the distributional variance of different domains—categories in our context, can improve the cross-domain (cross-category) generalization capabilities of the classifiers. A supervised extension to UDICA was introduced in depending on the inverse of a covariance operator as well as some mild assumptions. However, the inverse operation is both computationally expensive and unstable in practice. We instead propose to use the alternative of centered kernel alignment to account for the attribute labeling information. We show that the centered kernel alignment can be seamlessly integrated with UDICA, enabling us to learn both category-invariant and attribute-discriminative feature representations.

Our approach takes as input the features of the training images, their class (domain) labels, as well as their attribute labels. It operates upon kernels derived from the input data and learns a kernel projection to “distill” category-invariant and attribute-discriminative signals embedded in the original features. The overall output is a new feature vector for each image, which can be readily used in traditional machine learning models like SVMs for training the cross-category generalizable attribute detectors.

The contributions of the paper are summarized below.

To the best of our knowledge, this work is the first attempt to tackle attribute detection from the multi-source domain generalization point of view. This enables us to explicitly model the need that the attribute detectors should transcend different categories and generalize to previously unseen ones.

We introduce the centered kernel alignment to UDICA and arrive at an integrated method to strengthen the discriminative power of the learned attributes on one hand, and eliminate the domain differences between categories on the other hand.

We test our approach to four datasets: Animal With Attributes , Caltech-UCSD Birds , aPascal-aYahoo , and UCF101 , and test the learned representations on three tasks: attribute detection itself, zero-shot learning, and image retrieval. Our results are significantly better than those of competitive baselines, verifying the effectiveness of the new perspective for solving attribute detection as domain generalization.

The rest of this paper is organized as follows. In Section 2, we review related work in attribute detection, domain generalization, and domain adaptation. Section 3 and section 4 present the attribute learning framework. The experimental settings and evaluation results are presented in Section 5. Section 6 concludes the paper.

Related work and background

Our approach is related to two separate research areas, attribute detection and domain adaptation/generalization. We unify them in this work.

Earlier work on attribute detection mainly focused on modeling the correlations among attributes , localizing some special part-related attributes (e.g., tails of mammals) , and the relationship between attributes and categories . Some recent work has applied deep models to attribute detection . None of these methods explicitly model the cross-category generalization of the attributes, except the one by Farhadi et al. where the authors select features within each category to down-weight category-specific cues. Likely due to the fact that the attribute and category cues are interplayed, their feature selection procedure only gives limited gain. We propose to overcome this challenge by investigating all categories together and employing nonlinear mapping functions.

Attributes possess versatile properties and benefit a wide range of challenging computer vision tasks. They serve as the basic building blocks for one to compose categories (e.g., different objects) and describe instances , enabling knowledge transfer between them. Attributes also reveal the rich structures underlying categories and are thus often employed to regulate machine learning models for visual recognition . Moreover, attributes offer a natural human-computer interaction channel for visual recognition with humans in the loop , relevance feedback in image retrieval , and active learning . In this paper, we test the proposed approach on both attribute detection and its applications to zero-shot learning and image retrieval.

Domain generalization and adaptation.

Domain generalization is still at its early developing stage. A feature projection-based algorithm, Domain-Invariant Component Analysis (DICA), was introduced in to learn by minimizing the variance of the source domains. Recently, domain generation has been introduce into computer vision community for object recognition and video recognition . We propose to gear multi-source domain generalization techniques for the purpose of learning cross-category generalizable attribute detectors. Multi-source domain adaptation is related to our approach if we consider a transductive setting (i.e., the learner has access to the test data). While it assumes a single target domain, in attribute detection the test data are often sampled from more than one unseen domain.

1 Background on distributional variance

Denote by H\mathcal{H} and k(⋅,⋅)k(\cdot,\cdot) respectively a Reproducing Kernel Hilbert Space and its associated kernel function. For an arbitrary distribution Py(x)P_{y}(\bm{x}) indexed by y∈Yy\in\mathcal{Y}, the following mapping,

is injective if kk is a characteristic kernel . In other words, the kernel mean map μy\mu_{y} in the RKHS H\mathcal{H} preserves all the statistical information of Py(x)P_{y}(\bm{x}).

The distributional variance follows naturally,

Attribute detection

This section formalizes attribute detection and shows its in-depth connection to domain generalization.

Attribute detection as domain generalization —A new perspective.

In this paper, we understand attribute detection as a domain generalization problem. A domain refers to an underlying data distribution. In our context, it refers to the distribution Py(x,a)P_{y}(\bm{x},\bm{a}) of a category y∈[C]y\in[\mathsf{C}] over the input x\bm{x} and attribute labels a\bm{a}. As shown in Figure 2, the domains/categories are assumed to be related and are sampled from a common distribution P\mathcal{P}. This is reasonable considering that images and categories can often be organized in a hierarchy. Thanks to the relationships between different categories, we expect to learn new image representations for attribute detection, such that the corresponding detectors will perform well on both seen and unseen classes.

Approach

Our key idea is to find a feature transformation of the input x\bm{x} to eliminate the mismatches between different domains/categories in terms of their marginal distributions over the input, whereas ideally we should consider the joint distributions Py(x,a)P_{y}(\bm{x},\bm{a}), y∈[C]y\in[\mathsf{C}]. In particular, we use Unsupervised Domain Invariant Component Analysis (UDICA) and centered kernel alignment for this purpose. Note that modeling the marginal distributions Py(x)P_{y}(\bm{x}) is a common practice in domain generalization and domain adaptation and performs well in many applications. We leave investigating the joint distributions Py(x,a)P_{y}(\bm{x},\bm{a}) for future work.

Next, we present how to integrate UDICA and centered kernel alignment. Jointly they give rise to new feature representations which account for both attribute discriminativeness and cross-category generalizability.

The empirical distributional variance (cf. Section 2.1) between different domains/categories becomes the following in our context,

Intuitively, the domains would be perfectly matched when the variance is 0. Since there are many seen categories, each as a domain, we expect the learned projection to be generalizable and work well for the unseen classes as well.

Maximizing data variance.

Starting from the empirical kernel map (KB)(KB), it is not difficult to see that the data covariance is (KB)T(KB)/M(KB)^{T}(KB)/\mathsf{M} and the variance is

Regularizing the transformation.

UDICA regularizes the transformation by minimizing

Alternatively, one can use the Frobenius norm ∥B∥F\|B\|_{F}, as did in , to constrain the complexity of BB.

Combining the above criteria, we arrive at the following problem,

where the nominator corresponds to the data variance and the denominator sums up the distributional variance and the regularization over BB.

By solving the above problem, we are essentially blurring the boundaries between different categories and match the classes with each other, due to the distributional variance term in the denominator. This thus eliminates the barrier for attribute detectors to generalize in various classes. Our experiments verify the effectiveness the learned new representations (KB)(KB). Nonetheless, we can further improve the performance by modeling the attribute labels using centered kernel alignment.

2 Centered kernel alignment

Note that our training data are in the form of (xm,am,ym),m∈[M](\bm{x}_{m},\bm{a}_{m},y_{m}),m\in[\mathsf{M}]. For each image there are multiple attribute labels which may be highly correlated. Besides, we would like to stick to kernel methods to be consistent with our choice of UDICA—indeed, the distributional variance is best implemented by kernel methods (cf. Section 2.1). These considerations lead to our decision on using kernel alignment to model the multi-attribute supervised information.

Let Lm,m′=⟨am,am′⟩L_{m,m^{\prime}}=\left\langle\bm{a}_{m},\bm{a}_{m^{\prime}}\right\rangle be the kernel matrix over the attributes. Since LL is computed directly from the attribute labels, it preserves the correlations among them and serves as the “perfect” target kernel for the transformed kernel K~=KBBTK\widetilde{K}=KBB^{T}K to align to. The centered kernel alignment is then computed by,

where we abuse the notation LL slightly to denote that it is centered .

We would like to integrate the kernel alignment with UDICA in a unified optimization problem. To this end, firstly it is safe to drop tr(LL)\text{tr}(LL) in eq. (7) since it has nothing to do with the projection BB we are learning. Moreover, note that the role of tr(K~K~){\text{tr}(\widetilde{K}\widetilde{K})} duplicates with the regularization in eq. (6) to some extent, as it is mainly to avoid trivial solutions for the kernel alignment. We thus only add tr(K~L)\text{tr}(\widetilde{K}L) to the nominator of UDICA,

where γ∈\gamma\in balances the data variance and the kernel alignment with the supervised attribute labeling information. We cross-validate γ\gamma in our experiments. We name this formulation KDICA, which couples the centered kernel alignment and UDICA. The former closely tracks the attribute discriminative information and the latter facilitates the cross-category generalization of the attribute detectors to be trained upon KDICA.

By writing out the Lagrangian of the formalized problem (eq. (8)) and then setting the derivative with respect to BB to 0, we arrive at a generalized eigen-decomposition problem,

where Γ\Gamma is a diagonal matrix containing all the eigenvalues (Lagrangian multipliers). We find the solution BB as the Leading eigen-vectors. The number of eigen-vectors is cross-validated in our experiments. Again, we remind that (KB)(KB) serves as the new feature representations of the images for training attribute detectors. The details of our proposed framework has been shown in algorithm 1.

Experiment

This section presents our experimental results on four benchmark datasets. We test our approach for both the immediate task of attribute detection and two other problems, zero-shot learning and image retrieval, which could benefit from high-quality attribute detectors.

Dataset. We use four datasets to validate the proposed approach; three of them contain images for object and scene recognition and the last one contains videos for action recognition. (a) The Animal with attribute (AWA) dataset comprises of 30,475 images belonging to 50 animal classes. Each class is annotated with 85 attributes. Following the standard split by the dataset, we divide the dataset into 40 classes (24,295 images) to be used for training and 10 classes (6,180 images) for testing. (b) Caltech-UCSD Birds 2011 (CUB) is a dataset with fine-grained objects. There are 11,788 images of 200 different bird classes in CUB. Each class is annotated with 312 binary attributes. We split the dataset as suggested in to facilitate direct comparison (150 classes for training and 50 classes for testing). (c) aPascal-aYahoo consists of two attribute datases: the a-PASCAL dataset, which contains 12,695 images (6,340 for training and 6,355 for testing) collected for the Pascal VOC 2008 challenge, and a-Yahoo including 2,644 test images. Each images is annotated with 64 attributes. There are 20 object classes in a-Pascal and 12 in a-Yahoo and they are disjoint. Following the settings of , we use the pre-defined training images in a-Pascal as the training set and test on a-Yahoo. (d) UCF101 dataset is a large dateset for video action recognition. It contains 13,320 videos of 101 action classes. Each action class comes with 115 attributes. The videos are collected from YouTube with large variations in camera motion, object appearance, viewpoint, cluttered background, and illumination conditions. We run 10 rounds of experiments each with a random split of 81/20 classes for the training/testing sets, and then report the averaged results.

Implementation details. We choose the Gaussian RBF kernel and fix the bandwidth parameter as 1 for our approach to learning new image representations. After that, to train the attribute detectors, we input the learned representations into standard linear Support Vector Machines (see the empirical kernel map in Setcion 4.1). There are two free hyper-parameters when we train the detectors using the representations learned through UDICA, the hyper-parameter CC in SVM and the number b\mathsf{b} of leading eigen-vectors in UDICA. We use five-fold cross-validation to choose the best values for CC and b\mathsf{b} respectively from {0.01,0.1,1,10,100}\{0.01,0.1,1,10,100\} and {30,50,70,90,110,130,150}\{30,50,70,90,110,130,150\}. We use the same CC and b\mathsf{b} for KDICA and only cross-validate γ\gamma in equation (9) from {0.2,0.5,0.8}\{0.2,0.5,0.8\} to learn the SVM based attribute detectors with KDICA.

Evaluation. We first test our approach to attribute detection on all the four datasets (AWA, CUB, aPascal-aYahoo, and UCF101). To see how much the other tasks, which involve attributes, can gain from higher-quality attribute detectors, we further conduct zero-shot learning experiments on AWA, CUB, and UCF101, and multi-attribute based image retrieval on AWA. We evaluate the results of attribute detection and image retrieval by the averaged Area Under ROC Curve (AUC), the higher the better, and the results of zero-shot learning by classification accuracy.

2 Attribute prediction

We include in Table 1 both the results of these methods reported in the original papers, when they are available, and those we obtained (marked by ‘*’) by running the source code provided by the authors. We use the same CNN features (for AWA, CUB, and aPascal-aYahoo) and C3D features (for UCF101) we extracted for the baselines and our approach.

Overall results. From Table 1, we can find that UDCIA and KDICA outperform all the baselines on all the four datasets. More specifically, the relative accuracy gains of UDCIA over DAP are 6.3% on the AWA dataset and 5.4% on the CUB dateset, respectively, under the same feature and experimental settings. These clearly validate our assumption that blurring the category boundaries improves the generalizabilities of attribute detectors to previously unseen categories. The KDICA with centered kernel alignment is slightly better than the UDICA approach by incorporating attribute discriminative signals into the new feature representations. Delving into the per-unseen-class attribute detection result, we find that our KDICA-based approach improves the results of DAP for 71 out of 85 attributes on AWA and 272 out of 312 on CUB.

When domain generalization helps. We give some qualitative analyses using Figure 3 and 4 here. For the attributes in Figure 3, the proposed KDICA significantly improves the performance of the DAP approach. Those attributes (“muscle”, “domestic”, etc.) appear in visually quite different object categories. It seems like breaking the category boundaries is necessary in this case in order to make the attribute detectors generalize to the unseen classes. On the other hand, Figure 4 shows the attributes for which our approach can hardly improve DAP’s performance. The attribute “yellow” is too trivial to detect with nearly 100% accuracy already by DAP. The attribute “swim” is actually shared by visually similar categories, leaving not much room for KDICA to play any role.

3 Zero-shot learning

As the intermediate representations of images and videos, attributes are often used in high-level computer vision applications. In this section, we conduct experiments on zero-shot learning to examine whether the improved attribute detectors could also benefit this task.

Given our UDICA and KDICA based attribute detection results, we simply input them to the second layer of the DAP model to solve the zero-shot learning problem. We then compare with several well-known zero-shot recognition systems as shown in Table 2. We run our own experiments for some of them whose source code are provided by the authors. The corresponding results are again marked by ‘*’.

We see that in Table 2 the proposed simple solution to zero-shot learning outperforms the other state-of-the-art methods on the AWA, CUB, and UCF101 datasets, especially its immediate rival DAP. In addition, we notice that our kernel alignment technique (KDICA) improves the zero-shot recognition results over UDICA significantly on AWA. The improvements over UDICA on the other two datasets are also more significant than the improvements for the attribute prediction task (see Section 5.2 and Table 1). This observation is interesting; it seems like implying that increasing the quality of the attribute detectors is rewarding, because the increase will be magnified to even larger improvement for the zero-shot learning. Similar observation applies if we compare the differences between DAP and UDICA/KDICA respectively in Table 2 and Table 1. Finally, we note that our main purpose is indeed to investigate how better attribute detectors can benefit zero-shot learning. We do not expect to have a thorough comparison of the existing zero-shot learning methods.

4 Multi-attribute based image retrieval

In this section, we do some experiments on the AWA dataset for the multi-attribute based image retrieval, whose performance depends on the reliabilities of the attribute predictions. We input our learned feature representations to two popular frameworks for multi-attribute based image retrieval: TagProp and the fusion of individual prediction scores . In TagProp, we use its σ\sigmaML variant to compute the ranking scores of the multi-attributes queries. For the fusion of individual classifiers, we directly sum up the confidence scores corresponding to the multiple attributes in a query. The results of the fusion and TagProp are respectively shown in Table 3 and Table 4. We can observe that our attribute-oriented representations improve the fusion technique for image retrieval on a variety of queries (single attribute, attribute pairs, and triplets). Under the TagProp framework, the improvement is marginal on querying by attribute pairs and triples and significant for single-attribute queries.

Conclusion

In this paper, we propose to re-examine the fundamental attribute detection problem and develop a novel attribute-oriented feature representation by casting the problem as multi-source domain generalization, such that one can conveniently apply off-shelf classifiers to obtain high-quality attribute detectors. The attribute detectors learned from our new representation are capable of breaking the boundaries of object categories and generalizing well to unseen classes. Extensive experiment on four datasets, and three tasks, validate that our attribute representation not only improves the quality of attributes, but also benefits succeeding applications, such as zero-shot recognition and image retrieval.

Acknowledgement. This work was supported in part by NSF IIS-1566511. Chuang Gan was partially supported by the National Basic Research Program of China Grant 2011CBA00300, 2011CBA00301, the National Natural Science Foundation of China Grant 61033001, 61361136003. Tianbao Yang was partially supported by NSF IIS-1463988 and NSF IIS-1545995.

References