Evaluation of Output Embeddings for Fine-Grained Image Classification

Zeynep Akata, Scott Reed, Daniel Walter, Honglak Lee, Bernt Schiele

Introduction

The image classification problem has been redefined by the emergence of large scale datasets such as ImageNet . Since deep learning methods dominated recent Large-Scale Visual Recognition Challenges (ILSVRC12-14), the attention of the computer vision community has been drawn to Convolutional Neural Networks (CNN) . Training CNNs requires massive amounts of labeled data; but, in fine-grained image collections, where the categories are visually very similar, the data population decreases significantly. We are interested in the most extreme case of learning with a limited amount of labeled data, zero-shot learning, in which no labeled data is available for some classes.

Without labels, we need alternative sources of information that relate object classes. Attributes , which describe well-known common characteristics of objects, are an appealing source of information, and they can be easily obtained through crowd-sourcing techniques . However, fine-grained concepts present a special challenge: due to the high degree of similarity among categories, a large number of attributes are required to effectively model these subtle differences. This increases the cost of attribute annotation. One aim of this work is to move towards eliminating the human labeling component from zero-shot learning, e.g. by using alternative sources of information.

On the other hand, large-margin support vector machines (SVM) operate with labeled training images, so a lack of labels limits their use for this task. Inspired by previous work on label embedding and structured SVMs , we propose to use a Structured Joint Embedding (SJE) framework (Fig. 1) that relates input embeddings (i.e. image features) and output embeddings (i.e. side information) through a compatibility function, therefore taking advantage of a structure in the output space. The SJE framework separates the subspace learning problem from the specific input and output features used in a given application. As a general framework, it can be applied to any learning problem where more than one modality is provided for an object.

Our contributions are: (1) We demonstrate that unsupervised class embeddings trained from large unlabeled text corpora are competitive to previously published results that use human supervision. (2) Using the most recent deep architectures as input embeddings, we significantly improve the state-of-the-art (SoA). (3) We extensively evaluate several unsupervised output embeddings for fine-grained classification in a zero-shot setting on three challenging datasets. (4) By combining different output embeddings we obtain best results, surpassing the SoA by a large margin. (5) We propose a novel weakly-supervised Word2Vec variant that improves the accuracy when combined with other output embeddings.

The rest of the paper is organized as follows. Section 2 provides a review of the relevant literature; Sec. 3 details the SJE method; Sec. 4 explains the output embeddings that we analyze; Sec. 5 presents our experimental evaluation; Sec. 6 presents the discussion and our conclusions.

Related Work

Learning to classify in the absence of labeled data (zero-shot learning) is a challenging problem, and achieving better-than-chance performance requires structure in the output space. Attributes provide one such space; they relate different classes through well-known and shared characteristics of objects.

Attributes, which are often collected manually , have shown promising results in various applications, i.e. caption generation , face recognition , image retrieval , action recognition and image classification . The main challenge of attribute-based zero-shot learning arises on more challenging fine-grained data collections , in which categories may visually differ only subtly. Therefore, generic attributes fail at modeling small intra-class variance between objects. Improved performance requires a large number of specific attributes which increases the cost of data gathering.

As an alternative to manual annotation, side information can be collected automatically from text corpora. Bag-of-words is an example where class embeddings correspond to histograms of vocabulary words extracted automatically from unlabeled text. Another example is using taxonomical order of classes as structured output embeddings. Such a taxonomy can be built automatically from a pre-defined ontology such as WordNet . In this case, the distance between nodes is measured using semantic similarity metrics . Finally, distributed text representations learned from large unsupervised text corpora can be employed as structured embeddings. We compare several representatives of these methods (and their combinations) in our evaluation.

Embedding labels in an Euclidean space is an effective tool to model latent relationships between classes . These relationships can be collected separately from the data , learned from the data or derived from side information . In order to collect relationships independently of data, compressed sensing uses random projections whereas Error Correcting Output Codes builds embeddings inspired from information theory. WSABIE uses images with their corresponding labels to learn an embedding of the labels, and CCA maximizes the correlation between two different data modalities. DeViSE employs a ranking formulation for zero-shot learning using images and distributed text representations. The ALE method employs an approximate ranking formulation for the same using images and attributes. ConSe uses the probabilities of a softmax-output layer to weigh the semantic vectors of all the classes. In this work, we use the multiclass objective to learn structured output embeddings obtained from various sources.

Among the closest related work, ALE uses Fisher Vectors (FV ) as input and binary attributes / hierarchies as output embeddings. Similarly, DeviSe uses CNN features as input and Word2Vec representations as output embeddings. In this work, we benefit from both ideas: (1) We use SoA image features, i.e. FV and CNN, (2) among others, we also use attributes and Word2Vec as output embeddings. Our work differs from w.r.t. two aspects: (1) We propose and evaluate several output embedding methods specifically built for fine-grained classification. (2) We show how some of these output embeddings complement each other for zero-shot learning on general and fine-grained datasets. The reader should be aware of .

Structured Joint Embeddings

In this work, we aim to leverage input and output embeddings in a joint framework by learning a compatibility between these embeddings. We are interested in the problem of zero-shot learning for image classification where training and test images belong to two disjoint sets of classes.

The parameter vector ww can be written as a D×ED\times E matrix WW with DD being the input embedding dimension and EE being the output embedding dimension. This leads to the bi-linear form of the compatibility function:

Here, the input embedding is denoted by θ(x)\theta(x) and the output embedding by φ(y)\varphi(y). The matrix WW is learned by enforcing the correct label to be ranked higher than any of the other labels (Sec. 3.2), i.e. multiclass objective. This formulation is closely related to . Within the label embedding framework, ALE and DeViSe use pairwise ranking objective, WSABIE learns both φ(y)\varphi(y) and WW through ranking, whereas we use multiclass objective. Similarly, use the regression objective and CCA maximizes the correlation of input and output embeddings.

2 Parameter Learning

According to the unregularized structured SVM formulation , the objective is:

For the zero-shot learning scenario, the training and test classes are disjoint. Therefore, we fix φ\varphi to the output embeddings of training classes and learn WW. For prediction, we project a test image onto the WW and search for the nearest output embedding vector (using the dot product similarity) that corresponds to one of the test classes.

where ηt\eta_{t} is the learning step-size used at iteration tt. We use a constant step size chosen by cross-validation and we perform regularization through early stopping.

3 Learning Combined Embeddings

For some classification tasks, there may be multiple output embeddings available, each capturing a different aspect of the structure of the output space. Each may also have a different signal-to-noise ratio. Since each output embedding possibly offers non-redundant information about the output space, as also shown in , we can learn a better joint embedding by combining them together. We model the resulting compatibility score as

where W1,...,WKW_{1},...,W_{K} are the joint embedding weight matrices corresponding to the KK output embeddings (φk\varphi_{k}). In our experiments, we first train each WkW_{k} independently, then perform a grid search over αk\alpha_{k} on a validation set. Interestingly, we found that the optimal αk\alpha_{k} for previously-seen classes is often different from the one for unseen classes. Therefore, it is critical to cross-validate αk\alpha_{k} on the zero-shot setting.

Note that if we take αk=1/K,∀k\alpha_{k}=1/K,\forall k, Equation 5 is equivalent to simply concatenating the φk\varphi_{k}. This corresponds to stacking the WkW_{k} into a single matrix WW and computing the standard compatibility as in Equation 1. However, such a stacking learns a large WW where a high dimensional φ\varphi biases the final prediction. In contrast, α\alpha eliminates the bias, leading to better predictions. Thus, αk\alpha_{k} can be thought of as the confidence associated with φk\varphi_{k} whose contribution we can control. We show in Sec. 5.2 that finding an appropriate αk\alpha_{k} can yield improved accuracy compared to any single φ\varphi.

Output Embeddings

In this section, we describe three types of output embeddings: human-annotated attributes, unsupervised word embeddings learned from large text corpora, and hierarchical embeddings derived from WordNet.

Annotating images with class labels is a laborious process when the objects represent fine-grained concepts that are not common in our daily lives. Attributes provide a means to describe such fine-grained concepts. They model shared characteristics of objects such as color and texture which are easily annotated by humans and converted to machine-readable vector format. The set of descriptive attributes may be determined by language experts or by fine-grained object experts . The association between an attribute and a category can be a binary value depicting the presence/absence of an attribute (φ0,1\varphi^{0,1} ) or a continuous value that defines the confidence level of an attribute (φA\varphi^{\cal A} ) for each class. We write per-class attributes as:

where ρy,i\rho_{y,i} can be {0,1}\{0,1\} or a real number that associates a class with an attribute, yy denotes the associated class and EE is the number of attributes. Potentially, φA\varphi^{\cal A} encodes more information than φ0,1\varphi^{0,1}. For instance, for classes rat, monkey, whale and the attribute big, φ0,1=\varphi^{0,1}= implies that in terms of size rat == monkey << whale, whereas φA=\varphi^{\cal A}= can be interpreted as rat << monkey <<<< whale which is more accurate. We empirically show the benefit of φA\varphi^{\cal A} over φ0,1\varphi^{0,1} in Sec. 5.2. In practice, our output embeddings use a per-class vector form, but they can vary in dimensionality (EE). For the rest of the section we denote the output embeddings as φ\varphi for brevity.

2 Learning Label Embeddings from Text

In this section, we describe unsupervised and weakly-supervised label embeddings mined from text. With these label embeddings, we can (1) avoid dependence on costly manual annotation of attributes and (2) combine the embeddings with attributes, where available, to achieve better performance.

Word2Vec (φW\varphi^{\cal W}). In Word2Vec , a two-layer neural network is trained to predict a set of target words from a set of context words. Words in the vocabulary are assigned with one-shot encoding so that the first layer acts as a look-up table to retrieve the embedding for any word in the vocabulary. The second layer predicts the target word(s) via hierarchical soft-max. Word2Vec has two main formulations for the target prediction: skip-gram (SG) and continuous bag-of-words (CBOW). In SG, words within a local context window are predicted from the centering word. In CBOW, the center word of a context window is predicted from the surrounding words. Embeddings are obtained by back-propagating the prediction error gradient over a training set of context windows sampled from the text corpus.

GloVe (φG\varphi^{\cal G}). GloVe incorporates co-occurrence statistics of words that frequently appear together within the document. Intuitively, the co-occurrence statistics encode meaning since semantically similar words such as “ice” and “water” occur together more frequently than semantically dissimilar words such as “ice” and “fashion.” The training objective is to learn word vectors such that their dot product equals the co-occurrence probability of these two words. This approach has recently been shown to outperform Word2Vec on the word analogy prediction task .

Weakly-supervised Word2Vec (φWws\varphi^{{\cal W}_{ws}}). The standard Word2Vec scans the entire document using each word within a sample window as the target for prediction. However, if we know the global context, i.e. the topic of the document, we can use that topic as our target. For instance, in Wikipedia, the entire article is related to the same topic. Therefore, we can sample our context windows from any location within the article rather than searching for context windows where the topic explicitly appears in the text. We consider this method as a weak form of supervision.

We achieve the best results in our experiments using our novel variant of the CBOW formulation. Here, we pre-train the first layer weights using standard Word2Vec on Wikipedia, and fine-tune the second layer weights using a negative-sampling objective only on the fine-grained text corpus. These weights correspond to the final output embedding. The negative sampling objective is formulated as follows:

where vwv_{w} and vw′v_{w^{\prime}} are the label embeddings we seek to learn, and vcv_{c} is the average of word embeddings viv_{i} within a context window around word ww. D+D_{+} consists of context vcv_{c} and matching targets vwv_{w}, and D−D_{-} consists of the same vcv_{c} and mismatching vw′v_{w^{\prime}}. To find the viv_{i} (which are the columns of the first-layer network weights), we take them from a standard unsupervised Word2Vec model trained on Wikipedia.

During SGD, the viv_{i} are fixed and we update each sampled vwv_{w} and vw′v_{w^{\prime}} at each iteration. Intuitively, we seek to maximize the similarity between context and target vectors for matching pairs, and minimize it for mismatching pairs.

Bag-of-Words (φB\varphi^{\cal B}). BoW builds a “bag” of word frequencies by counting the occurrence of each vocabulary word that appears within a document. It does not preserve the order in which words appear in a document, so it disregards the grammar. We collect Wikipedia articles that correspond to each object class and build a vocabulary of most frequently occurring words. We then build histograms of these words to vectorize our classes.

3 Hierarchical Embeddings

Semantic similarity measures how closely related two word senses are according to their meaning. Such a similarity can be estimated by measuring the distance between terms in an ontology. WordNethttp://wordnetweb.princeton.edu/, a large-scale hierarchical database of over 100,000 words for English, provides us a means of building our class hierarchy. To measure similarity, we use Jiang-Conrath (φjcn\varphi^{jcn}), Lin (φlin\varphi^{lin}) and path (φpath\varphi^{path}) similarities formulated in Table 1. We denote our whole family of hierarchical embeddings as φH\varphi^{\cal H}. For a more detailed survey, the reader may refer to .

Experiments

While our main contribution is a detailed analysis of output embeddings, good image representations are crucial to obtain good classification performance. In Sec. 5.1 we detail datasets, input and output embeddings used in our experiments and in Sec. 5.2 we present our results.

We evaluate SJE on three datasets: Caltech UCSD Birds (CUB) and Stanford Dogs (Dogs)We use 113 classes that appear in the Federation Cynologique Internationale (FCI) database of dog breeds. are fine-grained, and Animals With Attributes (AWA) is a standard attribute dataset for zero-shot classification. CUB contains 11,788 images of 200 bird species, Dogs contains 19,501 images of 113 dog breeds and AWA contains 30,475 images of 50 different animals. We use a truly zero-shot setting where the train, val, and test sets belong to mutually exclusive classes. We employ train and val, i.e. disjoint subsets of training set, for cross-validation. We report average per-class top-1 accuracy on the test set. For CUB, we use the same zero-shot split as with 150 classes for the train+val set and 50 disjoint classes for the test set. AWA has a predefined split for 40 train+val and 10 test classes. For Dogs, we use approximately the same ratio of classes for train+val/test as CUB, i.e. 85 classes for train+val and 28 classes for test. This is the first attempt to perform zero-shot learning on the Dogs dataset.

Input Embeddings. We use Fisher Vectors (FV) and Deep CNN Features (CNN). FV aggregates per image statistics computed from local image patches into a fixed-length local image descriptor. We extract 128-dim SIFT from regular grids at multiple scales, reduce them to 64-dim using PCA, build a visual vocabulary with 256 Gaussians and finally reduce the FVs to 4,096. As an alternative, we extract features from a deep convolutional network. Features that are typically obtained from the activations of the fully connected layers have been shown to induce semantic similarities. We resize each image to 224×\times224 and feed into the network which was pre-trained following the model architecture of either AlexNet or GoogLeNet . For AlexNet (denoted as CNN) we use the 4,096-dim top-layer hidden unit activations (“fc7”) as features, and for GoogLeNet (denoted as GOOG) we use the 1,024-dim top-layer pooling units. For both networks, we used the publicly-available BVLC implementations . We do not perform any task-specific pre-processing, such as cropping foreground objects or detecting parts.

Output Embeddings. AWA classes have 85 binary and continuous attributes. CUB classes have 312 continuous attributes and the continuous values are thresholded around the mean to obtain binary attributes. The Dogs dataset does not have human-annotated attributes available.

We train Word2Vec (φW\varphi^{\cal W}) and GloVe (φG\varphi^{\cal G}) on the English-language Wikipedia from 13.02.2014. We first pre-process it by replacing the class-names, i.e. black-footed albatross, with alternative unique names, i.e. scientific name, phoebastrianigripes. We cross-validate the skip-window size and embedding dimensions. For our proposed weakly-supervised Word2Vec (φWws\varphi^{{\cal W}_{ws}}), we use the same embedding dimensions as the plain Word2Vec (φW\varphi^{\cal W}). For BoW, we download the Wikipedia articles that correspond to each class and build the vocabulary by omitting least- and most-frequently occurring words. We cross-validate the vocabulary size. φB\varphi^{\cal B} is a histogram of the vocabulary words as they appear in the respective document.

For hierarchical embeddings (φH\varphi^{\cal H}), we use the WordNet hierarchy spanning our classes and their ancestors up to the root of the tree. We employ the widely used NLTK libraryhttp://www.nltk.org/ for building the hierarchy and measuring the similarity between nodes. Therefore, each φH\varphi^{\cal H} vector is populated with similarity measures of the class to all other classes.

Combination of output embeddings. We explore combinations of five types of output embeddings: supervised attributes φA\varphi^{\cal A}, unsupervised Word2Vec φW\varphi^{\cal W}, GloVe φG\varphi^{\cal G}, BoW φB\varphi^{\cal B} and WordNet-derived similarity embeddings φH\varphi^{\cal H}. We either concatenate (cnc) or combine (cmb) different embeddings. In cnc, for instance in AWA, 85-dim φA\varphi^{\cal A} and 400-dim φW\varphi^{\cal W} would be merged to 485-dim output embeddings. In this case, if we use 1,024-dim GOOG as input embeddings, we learn a single 1,024×\times485-dim WW. In cmb, we first learn 1,024×\times85-dim WAW_{\cal A} and 1,024×\times400-dim WWW_{\cal W} and then cross-validate the α\alpha coefficients to determine the amount each embedding contributes to the final score.

2 Experimental Results

In this section, we evaluate several output embeddings on the CUB, AWA and Dogs datasets.

Discrete vs Continuous Attributes. Attribute representations are defined as a vector per class, or a column of the (class ×\times attribute) matrix. These vectors (85-dim for AWA, 312-dim for CUB) can either model the presence/absence (φ0,1\varphi^{0,1}) or the confidence level (φA\varphi^{\cal A}) of each attribute. We show that continuous attributes indeed encode more semantics than binary attributes by observing a substantial improvement with φA\varphi^{\cal A} over φ0,1\varphi^{0,1} with deep features (Tab. 2). Overall, CNN outperforms FV, while GOOG gives the best performing results; therefore in the following, we comment only on our results obtained using GOOG.

On CUB, i.e. a fine-grained dataset, φ0,1\varphi^{0,1} obtains 37.8% accuracy, which is significantly above the SoA (26.9% ). Moreover, φA\varphi^{\cal A} achieves an impressive 50.1% accuracy; outperforming the SoA by a large margin. We observe the same trend for AWA, which is a benchmark dataset for zero-shot learning. On AWA, φ0,1\varphi^{0,1} obtains 52.0% accuracy and φA\varphi^{\cal A} improves the accuracy substantially to 66.7%, significantly outperforming the SoA (48.5% ). To summarize, we have shown that φA\varphi^{\cal A} improves the performance of φ0,1\varphi^{0,1} using deep features, which indicates that with φA\varphi^{\cal A}, the SJE method learns a matrix WW that better approximates the compatibility of images and side information than φ0,1\varphi^{0,1}.

Learned Embeddings from Text. As the visual similarity between objects in different classes increases, e.g. in fine-grained datasets, the cost of collecting attributes also increases. Therefore, we aim to extract class similarities automatically from unlabeled online textual resources. We evaluate three methods, Word2Vec (φW\varphi^{\cal W}), GloVe (φG\varphi^{\cal G}) and the historically most commonly-used method BoW (φB\varphi^{\cal B}). We build φW\varphi^{\cal W} and φG\varphi^{\cal G} on the entire English Wikipedia dump. Note that the plain Word2Vec was used in ; however, rather than using Word2Vec in an averaging mechanism, we pre-process the Wikipedia as described in Sec 4.2 so that our class names are directly present in the Word2Vec vocabulary. This leads to a significant accuracy improvement. For φB\varphi^{\cal B} we use a subset of Wikipedia populated only with articles that correspond to our classes. On CUB (Tab. 3), the best accuracy is observed with φW\varphi^{\cal W} (28.4%) improving the supervised SoA (26.9% , Tab. 2). This is promising and impressive since φW\varphi^{\cal W} does not use any human supervision. On AWA (Tab. 3), the best accuracy is observed with φG\varphi^{\cal G} (58.8%) followed by φW\varphi^{\cal W} (51.2%), improving the supervised SoA (48.5% ) significantly. On Dogs (Tab. 3), the best accuracy is obtained with φB\varphi^{\cal B} (33.0%). On the other hand, using φW\varphi^{\cal W} (19.6%) and φG\varphi^{\cal G} (17.8%) leads to significantly lower accuracies. Unlike birds, different dog breeds belong to the same species and thus they share a common scientific name. As a result, our method of cleanly pre-processing Wikipedia by replacing the occurrences of bird names with a unique scientific name was not possible for Dogs. This may lead to vectors obtained from Wikipedia for dogs that are vulnerable to variation in nomenclature. In summary, our results indicate no winner among φW\varphi^{\cal W}, φG\varphi^{\cal G} and φB\varphi^{\cal B}. These embeddings may be task specific and complement each other. We investigate the complementarity of embeddings in the following sections.

Effect of Text Corpus. For φW\varphi^{\cal W} and φG\varphi^{\cal G}, we analyze the effects of three text corpora (B, W, B+W) with varying size and specificity. We build our specialized bird corpus (B) by collecting bird-related information from various online resources, i.e. audubon.org, birdweb.org, allaboutbirds.org and BNAhttp://bna.birds.cornell.edu/bna/. In combination, this corresponds to 50MB of bird-related text. We use the English-language Wikipedia from 13.02.2014 as our large and general corpus (W) which is 40GB of text. Finally, we combine B and W to build a large-scale text corpus enriched with bird specific text (B+W). On W and B+W, a small window size (10 for φW\varphi^{\cal W} and 20 for φG\varphi^{\cal G}); on B, a large window size (35 for φW\varphi^{\cal W} and 50 for φG\varphi^{\cal G}) is required. We choose parameters after a grid search. Increased specificity of the text corpus implies semantic consistency throughout the text. Therefore, large context windows capture semantics well in our bird specific (B) corpus. On the other hand, W is organized alphabetically w.r.t. the document title; hence, a large sampling window can include content from another article that is adjacent to the target word alphabetically. Here, small windows capture semantics better by looking at the text locally. We report our results in Tab. 4.

Using φG\varphi^{\cal G}, B+W (26.1%) gives the highest accuracy, followed by W (24.2%). One possible reason is that when the semantic similarity is modeled with cooccurrence statistics, output embeddings become more informative with the increasing corpus size, since the probability of cooccurrence of similar concepts increases.

Using φW\varphi^{\cal W}, the accuracy obtained with B (22.5%) is already higher than the φ0,1\varphi^{0,1}-based SoA (22.3%), illustrating the benefit of using fine-grained text for fine-grained tasks. Another advantage of using B is that, since it is short, building φW\varphi^{\cal W} is efficient. Moreover, building φW\varphi^{\cal W} with B does not require any annotation effort. Building φW\varphi^{\cal W} using W (28.4%) gives the highest accuracy, followed by W + B (27.5%) which improves the supervised SoA (26.9%). We speculate that since Word2Vec is a variant of the Feedforward Neural Network Language Model (FNNLM) , a deep architecture, it may learn more from negative data than positives. This was also observed for CNN features learned with a large number of unlabeled surrogate classes .

Additionally, we propose a weakly-supervised alternative to Word2Vec framework (φWws\varphi^{{\cal W}_{ws}}, Sec. 4.2). The weak-supervision comes from using the specialized B corpus to fine-tune the weights of the network and model the bird-related information. With φWws\varphi^{{\cal W}_{ws}} alone, we obtain 21.0% accuracy. However, when it is combined with φW\varphi^{\cal W} (28.4%), the accuracy improves to 29.7%. Compared to the results in Tab. 4, 29.7% is the highest accuracy obtained using unsupervised embeddings. We regard these results as a very encouraging evidence that Word2Vec representations can indeed be made more discriminative for fine-grained zero-shot learning by integrating a fine-grained text corpus directly to the output embedding learning problem.

Hierarchical Embeddings. The hierarchical organization of concepts typically embodies a fair amount of hidden information about language, such as synonymy, semantic relations, etc. Therefore, semantic relatedness defined by hierarchical distance between classes can form numerical vectors to be used as output embeddings for zero-shot learning. We build ontological relationships between our classes using the WordNet taxonomy. Due to its large size, WordNet encapsulates all of our AWA and Dog classes. For CUB, the high level bird species, i.e. albatross, appear as synsets in WordNet, but the specific bird names, i.e. black-footed albatross, are not always present. Therefore we take the hierarchy up to high level bird species as-is and we assume the specific bird classes are all at the bottom of the hierarchy located with the same distance to their immediate ancestors. The WordNet hierarchy contains 319 nodes for CUB (200 classes), 104 nodes for AWA (50 classes) and 163 nodes for Dogs (113 classes). We measure the distance between classes using the similarity measures from Sec 4.1.

While as shown in Fig. 2 different hierarchical similarity measures have very different behaviors on each dataset. The best performing φH\varphi^{\cal H} obtains 51.2% (Tab. 3) accuracy on AWA which reaches our φ0,1\varphi^{0,1} (52.0%) and improves φB\varphi^{\cal B} (44.9%) significantly. On CUB, φH\varphi^{\cal H} obtains 20.6% (Tab. 3) which remain below our φ0,1\varphi^{0,1} (37.8%) and approaches φB\varphi^{\cal B} (22.1%). On the other hand, on Dogs φH\varphi^{\cal H} obtains 24.3% (Tab. 3) which is significantly higher than the unsupervised text embeddings φW\varphi^{\cal W} (19.6%) and φG\varphi^{\cal G} (17.8%).

Combining Output Embeddings. In this section, we combine output embeddings obtained through human annotation (φA\varphi^{\cal A}), from text (φW,G,B\varphi^{\cal W,G,B}) and from hierarchies (φH\varphi^{\cal H}).We empirically found that the hierarchical embeddings φH\varphi^{\cal H} consistently improved performance when combined or concatenated with other embeddings. Therefore, we report results using φH\varphi^{\cal H} by default. As a reference, Tab. 3 summarizes the results obtained using one output embedding at a time. Our intuition is that because the different embeddings attempt to encapsulate different information, accuracy should improve when multiple embeddings are combined. We can observe this complementarity either by simple concatenation (cnc) or systematically combining (cmb) output embeddings (Sec.3.3) also known as early/late fusion . For cnc, we perform full SJE training and cross-validation on the concatenated output embeddings. For cmb, we learn joint embeddings WkW_{k} for each output separately (which is trivially parallelized), and find ensemble weights αk\alpha_{k} via cross-validation. In contrast to the cnc method, no additional joint training is used, although it can improve performance in practice. We observe (Tab. 5) in almost all cases cmb outperforms cnc.

We analyze the combination of unsupervised embeddings (φW,G,B,H\varphi^{\cal W,G,B,H}). On AWA, φG\varphi^{\cal G} (58.8%, Tab. 3) combined with φH\varphi^{\cal H} (51.2%, Tab. 3), we achieve 60.1% (Tab. 5) which improves the SoA (48.5%, Tab. 2) by a large margin. On CUB, combining φG\varphi^{\cal G} (24.2%, Tab. 3) with φH\varphi^{\cal H} (20.6%, Tab. 3), we get 29.9% (Tab. 5) and improve the supervised-SoA (26.9%, Tab. 2). Supporting our initial claim, unsupervised output embeddings obtained from different sources, i.e. text vs hierarchy, seem to be complementary to each other. In some cases, cmb performs worse than cnc; e.g. 28.2% versus 35.1% when using φB\varphi^{\cal B} with φH\varphi^{\cal H} on Dogs. In most other cases cmb performs equivalent or better. Combining supervised (φA\varphi^{\cal A}) and unsupervised embeddings (φW,G,B,H\varphi^{\cal W,G,B,H}) shows a similar trend. On AWA, combining φA\varphi^{\cal A} (66.7%, Tab. 3) with φG\varphi^{\cal G} and φH\varphi^{\cal H} leads to 73.9% (Tab. 5) which significantly exceeds the SoA (48.5%, Tab. 2). On CUB, combining φA\varphi^{\cal A} with φG\varphi^{\cal G} and φH\varphi^{\cal H} leads to 51.7% (Tab. 5), improving both the results we obtained with φA\varphi^{\cal A} (50.1%, Tab. 3) and the supervised-SoA (26.9%, Tab. 2). We have shown with these experiments that output embeddings obtained through human annotation can also be complemented with unsupervised output embeddings using the SJE framework.

Qualitative Results. Fig. 3 shows top-5 highest ranked images for classes chimpanzee, leopard and seal that are selected from 10 test classes of AWA. We use GOOG as input embeddings and as output embeddings we use supervised φA\varphi^{\cal A}, the best performing unsupervised embedding on AWA (φG\varphi^{\cal G}), and the combination of the two (φG+A\varphi^{\cal G+A}). For the class chimpanzee, φA\varphi^{\cal A} emphasizes that chimpanzees live on trees, which is among the list of attributes. On the other hand, φG\varphi^{\cal G} models the social nature of the animal, ranking a group of chimpanzees interacting with each other at the highest. Indeed this information can easily be retrieved from Wikipedia. φG+A\varphi^{\cal G+A} synthesizes both aspects. Similarly, for leopard φA\varphi^{\cal A} puts an emphasis on the head where we can observe several of the attributes, i.e. color, spotted, whereas φG\varphi^{\cal G} seems to place the animal in the wild. φG+A\varphi^{\cal G+A} combines both aspects. In case of class seal, φA\varphi^{\cal A} retrieves images related to water and ranks whales and seals highest, whereas φG\varphi^{\cal G} adds more context by placing seals in the icy natural environment and within groups. Finally, φG+A\varphi^{\cal G+A} ranks seal-shaped animals on ice, close to water and within groups the highest. We find these qualitative results interesting as they depict how (1) unsupervised embeddings capture nameable semantics about objects and (2) different output embeddings are semantically complementary for zero-shot learning.

Conclusion

We evaluated the Structured Joint Embedding (SJE) framework on supervised attributes and unsupervised output embeddings obtained from hierarchies and unlabeled text corpora. We proposed a novel weakly-supervised label embedding technique. By combining multiple output embeddings (cmb), we established a new SoA on AWA (73.9%, Tab. 6) and CUB (51.7%, Tab. 6). Moreover, we showed that unsupervised zero-shot learning with SJE improves the SoA, to 60.1% on AWA and 29.9% on CUB, and obtains 35.1% on Dogs (Tab. 6).

We emphasize the following take-home points: (1) Unsupervised label embeddings learned from text corpora yield compelling zero-shot results, outperforming previous supervised SoA on AWA and CUB (Tab. 2 and 3). (2) Integrating specialized text corpora helps due to incorporating more fine-grained information to output embeddings (Tab. 4). (3) Combining unsupervised output embeddings improve the zero-shot performance, suggesting that they provide complementary information (Tab. 5). (4) There is still a large gap between the performance of unsupervised output embeddings and human-annotated attributes on AWA and CUB, suggesting that better methods are needed for learning discriminative output embeddings from text. (5) Finally, supporting , encoding continuous nature of attributes significantly improve upon binary attributes for zero-shot classification (Tab. 2).

As future work, we plan to investigate other methods to combine multiple output embeddings and to improve the discriminative power of unsupervised and weakly-supervised label embeddings for fine-grained classification.

This work was supported in part by ONR N00014-13-1-0762, NSF CMMI-1266184, Google Faculty Research Award, and NSF Graduate Fellowship.

References