Preserving Semantic Relations for Zero-Shot Learning
Yashas Annadani, Soma Biswas
Introduction
Novel categories of objects arise dynamically in nature. It is estimated that around 8000 species of animals and plants are discovered every year . However, current recognition models are quite incapable of handling this dynamic scenario when labeled examples of novel categories are not available. Obtaining labeled examples followed by retraining or transfer learning can be expensive and cumbersome. Zero-shot learning (ZSL) offers an elegant way to address this problem by utilizing the mid-level semantic descriptions of the categories. These descriptions are usually encoded in an attribute vector , sometimes referred to as side information or class embeddings. This paradigm is useful not only for emerging categories, but also for extending the recognition capabilities of a model beyond the categories it is trained on without requiring additional training data.
Some of the existing approaches treat zero-shot recognition as a ranking problem . In these approaches, a compatibility function between the image feature and the class embeddings is learned such that its score for the correct class is higher than that for an incorrect class by a fixed margin. Ranking can lead to loss of some of the semantic structure available from the attributes due to the fixed margin and the unbounded nature of the compatibility function. In some of the other approaches , zero-shot recognition is typically achieved by embedding either the image features or the attribute vectors, or both, to a predefined embedding space using a ridge regression or a mean squared error objective. Here, proper choice of the embedding space is essential. If the space spanned by the attributes (semantic space) is used as the embedding space, then the semantic structure is preserved, but the problem of hubness surfaces . To alleviate this problem, few recent approaches map the class embeddings to the space spanned by the image features (visual space). However, the visual space may not contain semantic properties, as it may be inherited from a model trained on a supervised classification task where labels are one-of-k coded.
We believe that two things are crucial for zero-shot recognition: (1) discriminative ability on the categories available during training and (2) inheriting the properties of semantic space for efficient classification on novel categories. Existing approaches focus on either of the two aspects. To this end, we propose a simple yet effective approach which ensures discriminative capability while retaining the structure of the semantic space in an encoder-decoder multilayer perceptron framework.
We decompose the structure of the semantic space to a set of relations between categories. Our aim is to preserve these relations in the embedding space so as to appropriately inherit the structure of the semantic space to the embedding space. Relation between categories is decomposed to three groups: identical, semantically similar and semantically dissimilar. We construct a semantic tuple which consists of samples belonging to categories of each of the relations with respect to a given category. Objective function specific to each relation is formulated so that the underlying semantic structure can be captured while still ensuring discriminative capability. The underlying principle is that the embeddings belonging to categories which are semantically similar in the attribute space must still be close in the embedding space, while the ones which are dissimilar should be far away.
Propose a simple and effective approach for zero-shot recognition which preserves the structure of the semantic space in the embedding space by utilizing semantic relations between categories.
Extensive experimental evaluation on multiple datasets, including the large scale ImageNet shows that the proposed method improves over the state-of-the art in multiple settings.
Demonstrate how the proposed approach can be useful for making approximate inferences about images belonging to novel categories even when the class embeddings corresponding to that category is not available.
The rest of the paper is organized is follows: Section 2 reviews the related work followed by Section 3 which describes the proposed approach in detail. Experimental evaluation is reported in Section 4 with some pertinent discussions in Section 5. We conclude the paper in Section 6.
Related Work
Apart from bilinear compatibility models, few other approaches map image features to semantic space by using a ridge regression objective. Kodirov et al. use an additional reconstruction constraint on the mapped features which enhanced zero-shot recognition performance. In , the parameters of a deep network is learned using side information like Word2Vec. This work uses binary cross entropy loss and hinge loss in addition to mean squared error loss. Recently, Zhang et al. proposed to reverse the direction of mapping from semantic space to visual space. However, mapping from semantic space to visual space may result in reduced semantic expressiveness of the model.
Few of the other approaches employ manifold learning to solve zero-shot recognition. Changpinyo et al. construct a weighted bipartite graph in a space where additional classes called phantom classes are introduced. They minimize a distortion objective which aligns this space with the class embeddings. use a transductive approach to learn projection function by matrix tri-factorization and preserving the underlying manifold structure of both visual space and semantic space. However, our approach differs from these approaches as our method is neither transductive nor involves manifold learning.
Our method relies on using semantic relations to learn the embeddings. Parikh and Grauman use partial ordering and ranking formulation to capture semantic relations across individual attributes. This is different from our approach wherein the semantic relations are defined on the categories themselves.
Proposed Approach
We explicitly define semantic relations between classes so that objective function specific to each relation can be formulated in order to preserve the relations. For a given set of classes, we wish to group them to three different categories with respect to a reference class - identical, semantically similar and semantically dissimilar. This particular grouping should be reflective of the underlying semanticity. There are many possible ways to define these class relations. For example, one can leverage prior knowledge about the classes specific to the task under study.
In this work, we employ class embeddings to define semantic relations. Let be a similarity measure between two class embeddings. We use cosine similarity as a semantic similarity measure .
and are any two vectors of same dimension. Given for any two embeddings and , we define class relations as follows:
Belonging to same class (identical) if .
Semantically similar if .
Semantically dissimilar if .
where is a threshold. Without loss of generality, can be fixed to zero, which is a reasonable estimate for a cosine similarity. can also be chosen over a validation set. In this paper, we choose based on the performance on the validation set. Since the semantic space of attributes is shared between seen and unseen classes, this particular definition provides a good generalization to novel classes.
2 Preserving Semantic Relations
Armed with the above definition, we wish to map the class embeddings to the visual space such that the semantic relation between the mapped class embeddings and the visual features reflects the relation between their corresponding classes. Our motivation to map to the visual space comes from the works of Shigeto et al. discusses hubness problem only for ridge regression based techniques. However, hubness problem can also arise in cosine similarity measure based approaches, as demonstrated in . and Zhang et al. , which showed that using visual space instead of semantic space or any other intermediate space as the embedding space alleviated the hubness problem , a problem in which a few points (called as hubs) arise which are in the -nearest neighbors of most of the other points. We employ nearest neighbor search for zero-shot recognition, and hence the problem of hubness can lower the performance if mapped to any space other than the visual space. Therefore, we use the visual space as the embedding space for our approach.
We use an encoder-decoder multilayer perceptron architecture to learn the embedding function. The encoder parameterized by learns the mapping and the decoder learns to reconstruct the input from the mapped class embeddings. Our model formulation is inspired from , though does not explicitly use a multilayer perceptron based encoder-decoder. Conventionally, mean squared error loss is used to reduce the discrepancy between the identical embeddings. However, this may not preserve the semantic structure. In this work, we explicitly formulate objective functions to preserve the semantic structure in the embedding space.
In order to facilitate the task of preserving semantic relations, we consider a tuple of visual features for every class embedding to be mapped. The elements of the tuple are sampled such that the semantic relationship between their corresponding class embeddings and satisfy the conditions in the definition. Specifically, if is the class embedding tuple corresponding to the visual tuple, then , and . The first feature corresponds to the same class (i.e. ), the second corresponds to a semantically similar class and the third feature corresponds to a semantically dissimilar class. We present a way to efficiently sample the tuples in Section 3.3. Note that this essentially forms a quadruplet . Though quadruplet based algorithms have been used in literature , the one explored here is fundamentally different. Objective for Identical and Dissimilar Classes. The mapped class embedding and the visual feature must have a high semantic similarity score as they belong to the same class. Ideally, it should be equal to one. Also and the visual feature must have a very low semantic similarity score as they belong to dissimilar classes, i.e. . The objective function which caters to the above needs is given by:
The first term caters to the identical class and aims to maximize the semantic similarity between and . The second term aims to minimize the semantic similarity between dissimilar entities and . Here acts as an adaptive scaling term, i.e. if the class embeddings are very dissimilar, we put a higher weight on the term to minimize it. Objective for Similar Classes. Since and are embeddings of semantically similar classes, and need to be close in order to preserve this relation. Explicitly, s\big{(}f(\mathbf{y_{r}};\theta_{f}),\,\mathbf{x_{j}}\big{)} must be greater than . In addition to this condition, we also want to ensure that the above enforced condition does not interfere with zero-shot recognition. Therefore, we restrict the semantic similarity score s\big{(}f(\mathbf{y_{r}};\theta_{f}),\,\mathbf{x_{j}}\big{)} to be less than . This ensures that semantic similarity is preserved without hindering the recognition task. Mathematically, the objective which reflects the above two conditions is as follows:
where . Note that only one of the two terms is triggered, corresponding to the either of the conditions s\big{(}f(\mathbf{y_{r}};\theta_{f}),\,\mathbf{x_{j}}\big{)}\geq\tau or s\big{(}f(\mathbf{y_{r}};\theta_{f}),\,\mathbf{x_{j}}\big{)}\leq\delta_{jr} they violate. The above constraints are enforced only on semantically similar classes. For semantically dissimilar classes, we aim to have a similarity score as small as possible regardless of the amount of dissimilarity because in most applications the amount of dissimilarity is of little concern. Reconstruction Loss. Since our setup involves a decoder which reconstructs the input , there is an accompanying reconstruction loss. We noted in our experiments that using this additional condition of reconstruction provided better updates to the encoder and enhanced zero-shot recognition performance. In addition, this is in spirit with the observation of Kodirov et al. that adding an additional reconstruction term is beneficial for zero-shot recognition.
where, is the output of the decoder . Overall Objective. With the above three objective functions, the overall objective is given by:
Here refers to the size of the mini-batch . and are hyper-parameters chosen based on the validation data. Given a testing sample , we infer its class as follows:
where refers to the class embeddings of only the unseen classes in the conventional zero-shot setting and to the class embeddings of both seen as well as unseen classes in the generalized zero-shot setting.
3 Mining the Tuples
The proposed algorithm relies on sampling the tuples for preserving the semantic relations. In the tuple , can be chosen at random such that it belongs to the same class as , which we choose sequentially from the dataset. There are many possible ways to choose and . Choosing the most informative tuples will help in faster convergence and provide useful updates for gradient descent based algorithms. In this work, we sample the tuples in an online fashion, wherein for each epoch a criterion is evaluated. Our method is similar to the hard negative mining approach for triplet based learning algorithms . For every we wish to embed, we randomly sample s which satisfy the condition . Among these s, we update the parameters of the model with that particular sample which gives the highest loss in objective . Similarly, we randomly sample different s which satisfy the condition that . Among these s, we update the parameters of the model with that particular sample which gives the highest value in the second term of objective .
It can be seen that we update the model from a set of randomly sampled points. This is much more efficient compared to updating the model using the hardest negative which involves computing the maximum over a much larger set of points. In addition, our method also circumvents the problem of potentially reaching a bad minima due to updation of the model with the hardest negative. We also tried updating with the semi-hard negatives as described in , but it did not lead to any significant impact on the results.
Experiments
We evaluate the proposed approach on four datasets for zero-shot learning : SUN , Animals with Attributes 2 (AWA2) , Caltech UCSD Birds 200-2011 (CUB) and Attribute Pascal and Yahoo dataset (aPY) . All these datasets are provided with annotated attributes.
The details of these datasets are listed in Table 1. It was observed in that some of the testing categories in the original split of the datasets are subset of the Imagenet categories. Hence, extracting features from Imagenet trained models will not result in a true zero-shot setting. In order to alleviate the problem, the authors proposed a new split such that none of the testing categories coincide with Imagenet categories. In addition, some samples from seen categories were held out for generalized zero-shot recognition. Hence, we employ the protocol and the splits as described in . We use continuous per-class attributes for all the datasets and average per-class top-1 accuracy to report the results.
We use the 2048-D Resnet-101 features provided by for all the datasets. The architecture details of the proposed approach are given in Table 1. ReLU activation is used for all the layers except for the output of the encoder and the decoder, which employ ELU activations. We use Adam optimizer with a learning rate of and a weight decay of . All the input features and attributes are normalized to have zero mean and unit standard deviation.
The comparisons with the state-of-the-art are made with the results reported in , as it is based on exactly the same protocol and use the same set of features. Besides, the algorithms on which the results are reported encompass a wide range of approaches in zero-shot learning. Baselines. We define three baseline settings which provide insights into the importance of each of the terms in the proposed objective function. For baseline , instead of and , mean squared error objective is used to learn the mapping. The same architecture as used by the proposed approach is used along with the reconstruction objective . This will align the mapped class embeddings with the structure of the visual space. Since this setup does not enforce the relations as described in Section 3.1, the embedding space may not maintain the structure of the semantic space. For baseline , we set and in our approach. This helps us to better understand the importance of objective . This results in maximizing the cosine similarity for identical classes and minimizing the same for all other classes. In addition to the above two baselines, we also demonstrate the importance of the objective . We employ just the encoder and set in our experiments for . This provides insight into the degree of enhancement in performance due to the reconstruction term.
2 Conventional Zero-Shot Learning Results
The results of the proposed approach on various datasets is listed in Table 2. We also provide comparisons with the state-of-the-art. With respect to the first baseline B1, we observe that the proposed approach consistently performs better on all the datasets. This supports our hypothesis that inheriting semantic properties to the embedding space is beneficial for zero-shot recognition. We also observe significant increase in performance when we include the objective in our approach. In this case, the difference in performance is pronounced in coarse-grained datasets AWA2 and aPY wherein the inter-class semantics are much different. This indicates that the objective , which is essentially the structure preserving term, is beneficial for zero-shot recognition. The reconstruction term also contributes to varying levels of gain in performance. Visualization of the embedding space is presented in Figure 1 for the ten unseen classes of AWA2 dataset. It can be seen that the semantic relations are preserved to a good extent.
The proposed approach also compares favorably with the existing approaches in literature, with our approach obtaining the state-of-the-art on SUN, AWA2 and CUB datasets. On aPY dataset, we obtain 38.4% which is slightly less than Deep Visual Semantic Embedding Model . However, on the generalized zero-shot setting, the proposed approach performs much better, as illustrated next. Effectiveness of the tuple mining approach. The graph showing the accuracy on the validation split against the number of epochs for the four datasets is shown in Figure 3. We observe that for all the datasets, around 80% of the maximum accuracy is reached in less than 5 epochs.
3 Generalized Zero-Shot Learning Results
In , some of the samples from seen classes are held out for testing. In this setting, the search space consists of both the seen classes as well as the unseen classes. This scenario is more realistic, as we cannot usually anticipate whether an incoming sample belongs to a seen class or an unseen class. Table 3 reports the result of generalized zero-shot learning on the four datasets under two different settings. The first setting (referred to as ts) involves comparison of samples from unseen classes against both seen and unseen classes. The second setting (referred to as tr) involves comparison of samples from unseen classes as well as held-out samples from seen classes against all the classes. High accuracy on tr and low accuracy on ts implies that the model performs well on the seen classes but fails to generalize to the unseen classes. The harmonic mean (denoted by H) of the two results is also reported, as this measure encourages accuracies for both the settings to be high .
With respect to B1, we can see that our method performs better in the first setting (ts) by a large margin. With respect to B2, there is a gain in accuracy on all the settings. The difference is pronounced in the first setting, which is concerned with samples belonging to novel classes. This indicates that employing the concept of semantic relations and preserving these relations in the embedding space is beneficial for classification on novel categories. In fact, the observations made for conventional zero-shot learning setting are also applicable here in a more realistic setting. With respect to the state-of-the-art, our approach gives a harmonic mean accuracy of 26.7% on SUN which is the best result among all the reported methods. In addition, we obtain 32.3% on the AWA2 dataset, better than the next best method by nearly 5%. On CUB, proposed approach obtains a best accuracy of 24.6% on the first setting and 33.9% overall. On aPY dataset, our approach achieves 13.5% on the first setting and 51.4% on the second setting, with an overall result of 21.4%. It can be seen that methods like CONSE perform very well on seen classes but do not generalize well for novel classes. Although our method does not match the accuracy of the methods like CONSE and CMT on the second setting, it outperforms them on the first setting by a large margin which is reflected in the harmonic mean of the two. In addition, our method performs better compared to other competitive methods like ALE and DEVISE on the first setting on AWA2, CUB and aPY.
4 Experiments on ImageNet
ImageNet is a large scale dataset consisting of nearly 14.1 million images belonging to 21,841 categories. The 1000 categories of ILSRVC are used as seen classesImages from one of the seen classes namely ‘teddy bear’ which belongs to the synset n04399382 was unavailable. Hence, we use only 999 categories for training. while the rest as unseen classes. There are no curated attributes available for this dataset. Distributed word embeddings like Word2Vec are employed instead as they have been shown to contain semantic properties which are suitable for zero-shot learning. We extract the Resnet-101 features from the pretrained model available in the pytorch model zoo. We use the Word2Vec provided by . We employ two hidden layers for the encoder and the decoder, similar to AWA2 and aPY datasets. For comparison, we implement the SYNC algorithm with the aforementioned features and settings. To the best of our knowledge, SYNC achieves the state-of-the-art performance on this very challenging dataset . We use the public code made available by the authors for evaluation.
Table 4 lists the results obtained using the proposed approach on different splits of test data. We observe that our approach consistently achieves better performance compared to SYNC, thus improving over the state-of-the-art. On the generalized zero-shot setting, we observe that the proposed approach outperforms SYNC by a large margin. This indicates that the proposed approach scales favorably to a realistic setting with large number of unseen classes.
5 Approximate Semantic Inference
It is reasonable to assume that attributes or other form of side information is available in most of the scenarios. However, there might be a situation in which side information for a few categories of interest might not be available. For example, in ImageNet, out of the 21,841 categories, Word2Vec for 497 categories is not present. This bottleneck might crop up in a very large scale setting as evident in the ImageNet example. In this scenario, though classification is not possible, we wish to infer the approximate semantic characteristics of the object in the image. Since our model has semantic characteristics in its embedding space, we can approximately infer the semantic properties of the image in relation to the existing categories (seen and unseen). Some of the empirical results on categories of ImageNet for which class embeddings is not available can be seen in Figure 4. In the first example, the bird in the image belongs to Aegypiidae, for which the class embedding is not available. The image feature is compared using cosine similarity with the existing class embeddings. It can be inferred that the given image is semantically similar to Prairie Chicken, Black Chicken and Vulture. Moreover, it can also be inferred that categories Poll Parrot, Trogon and Cannon are dissimilar to the bird in the image. This is very much in agreement to the actual semantic properties of the categories. In addition, the cosine similarity score is evocative of the degree of similarity between the actual category and the category with which it is compared. Similar observations can be made from other examples as well. This suggests that inheriting the structure of semantic space helps to make approximate inference about an image with respect to known entities, thus showing potential for tasks beyond just zero-shot classification.
Discussion
The cosine similarity function applied on the mapped class embeddings and the image features can be approximated to a normalized compatibility score function. Thus, the setup of baseline B2 is similar to the ranking based methods which employ compatibility functions wherein the embeddings which belong to the same class are pulled together while the rest are pushed apart. The results are also similar to the ones achieved using these methods. Thus, a particular instantiation of the proposed approach can be approximated to compatibility models with the ranking objective. In addition, the merits of approaches and are seamlessly incorporated in our model.
Although the advantages of the proposed approach is clear and encouraging, one of the limitations is its performance on the seen categories in the generalized zero-shot setting. Though it does not match the results on some of the previous approaches in literature, the proposed approach still gives encouraging performance. We believe that exploration of more intricate forms of relations between categories would help in furthering the state-of-the art in this setting.
Conclusion
In this work, we focus on efficiently utilizing the structure of the semantic space for improved classification on the unseen categories. We introduce the concept of relations between classes in terms of their semantic content. We devise objective functions which help in preserving semantic relations in the embedding space thereby inheriting the structure of the semantic space. Extensive evaluation of the proposed approach is carried out and state-of-the-art results are obtained on multiple settings including the tougher generalized zero-shot learning, thus proving its effectiveness for zero-shot learning. Acknowledgment. The authors would like to thank Devraj Mandal of IISc for helpful discussions.