Zero-shot Image Recognition Using Relational Matching, Adaptation and Calibration

Debasmit Das, C. S. George Lee

I Introduction

Recent work on visual recognition focuses on the importance of obtaining large labeled datasets such as ImageNet . Large-scale datasets when used for training deep neural network models tend to produce state-of-the-art results on visual recognition . However, in some cases, it may be difficult to obtain a large number of samples for certain rare or fine-grained categories. Hence, recognizing these rare categories become difficult. Humans, on the other hand, can easily recognize these rare categories by identifying the semantic description of the new category and how it is related to the seen categories. For example, a person can identify a new animal zebra by identifying the semantic description of a zebra having black and white stripes and looking like a horse. A similar approach is undertaken for learning models to recognize unseen and rare categories. This learning scenario is known as zero-shot learning (ZSL) because zero labeled samples of the unseen categories are available for the training stage. ZSL has promising ramifications in autonomous vehicles, medical imaging, robotics, etc., where it is difficult to annotate images of novel categories but high-level semantic descriptions of classes can be obtained easily.

To be able to recognize unseen categories, we usually train a learning model using a large collection of labeled samples from the seen categories and then adapt it to unseen categories. For zero-shot recognition, the seen and the unseen categories are related through a high-dimensional vector space known as semantic-descriptor space. Each category is assigned a unique semantic descriptor. Examples of semantic descriptor can be manually defined attributes or automatically extracted word vectors . Figure 1 depicts the ZSL problem in terms of how much information is available during training and testing.

Most ZSL methods involve mapping from the visual feature space to the semantic-descriptor space or vice versa . Sometimes, both the visual features and the semantic descriptors are mapped to a common feature space . Most of these mapping-based approaches learn an embedding function for samples and semantic descriptors. The embedding is learned by minimizing a similarity function between the embedded samples and the corresponding embedded semantic descriptors. Thus, most ZSL methods differ in the choice of the embedding and similarity functions. Lampert et. al used linear classifiers, identity function and Euclidean distance for the sample embedding, semantic embedding and similarity metric, respectively. Romera-Paredes et al. used linear projection, identity function and dot product. ALE , DEVISE , SJE all used a bilinear compatibility framework, where the projection was linear and the similarity metric was a dot product. They used different variations of pairwise ranking objective to train the model. LATEM was an extension of the above method, which used piecewise linear projections to account for the non-linearity. CMT used a neural network to map image features to semantic descriptors with an additional novelty detection stage to detect unseen categories. SAE used an auto-encoder-based approach, where the image feature is linearly mapped to a semantic descriptor as well as being reconstructed from the semantic-descriptor space. DEM used a neural network to map from a semantic-descriptor space to an image-feature space.

After the embedding is carried out, classification is performed using the nearest-neighbor search. An earlier study showed that the nearest-neighbor search in such a high-dimensional space suffers from the hubness phenomenon because only a certain number of data-points becomes nearest neighbor or hubs for almost all the query points, resulting in erroneous classification results. However, Shigeto et al. showed that mapping from a semantic-descriptor space to a visual-feature space does not aggravate the hubness problem. Thus, in this paper, we pursue a semantic-descriptor-space to a visual-feature-space mapping approach. We further introduce the concept of relative features that uses pairwise relations between data-points. This not only provides additional structural information about the data but also reduces the dimensionality of the feature space implicitly , thus alleviating the hubness problem.

Zero-shot learning further suffers from a projection-domain-shift problem because the mapping from the semantic-descriptor space to the visual-feature space is learned from the data belonging to only the seen categories. As a result, the projected semantic descriptors of the unseen categories are misplaced from the unseen test-data distribution. Fu et al. identified the domain-shift problem and used multiple semantic information sources and label propagation on unlabeled data from the unseen categories to counter the problem. Kodirov et al. cast ZSL as a dictionary-learning problem and constrained the dictionary of the seen and unseen data to be close to each other. This transductive approach is unrealistic as it assumes access to the unlabeled test data from unseen categories during the training stage. At the very least, we could carry out the test-time post-processing of the semantic descriptors. For test-time adaptation, we propose to find correspondences between the projected semantic descriptors and the unlabeled test data after which the descriptors are further mapped to the corresponding data-points. This is inspired by recent work on local correspondence-based approach to unsupervised domain adaptation , which produces better results than global domain-adaptation methods.

Another problem with ZSL is that models are generally evaluated only on unseen categories. In a real-world scenario, we expect the seen categories to appear more frequently compared to the unseen categories. As a result, it is appropriate to test our model on both seen and unseen categories. This evaluation setting is known as Generalized Zero-Shot Learning (GZSL) and was initially introduced by Chao, et al. . They found that the performance of unseen categories in the GZSL setting was poor and proposed a shifted-calibration mechanism to improve the performance. This shifted-calibration mechanism lowers the classification scores of the seen categories. We propose to develop a scaled-calibration mechanism to study the effect on recognition performance. This has an effect of changing the effective variance of a class and is therefore more interpretable.

Other methods for ZSL include hybrid and synthesized methods. Hybrid models expressed image features or semantic embeddings as a combination/mixture of existing seen features or semantic embeddings. Semantic Similarity Embedding (SSE) exploits class relationship at both the image-feature and semantic-descriptor spaces to map them into a common embedding space. Our proposed ZSL method also exploits pairwise relationships between classes by minimizing the discrepancy between the projected semantic descriptors and the corresponding class prototype obtained from the image features. CONSE learns the probability of a seen sample belonging to a seen class and uses the probability of an unseen sample belonging to seen classes to relate to the semantic-descriptor space. Synthetic Classifiers (SYNC) learn a mapping between the semantic-embedding space and the model-parameter space. The model parameters of the classes are represented as a combination of phantom classes, the relationship with which is encoded through a weighted bipartite graph. Synthesized methods generally convert ZSL into a standard supervised-learning problem by generating samples for the unseen categories. Some of these methods include . The limitations of these methods lie in not being able to generate samples very close to the true distribution. A more comprehensive overview of recent work on ZSL can be found in .

To summarize, we propose a three-step approach to zero-shot learning. Firstly, to prevent aggravating the hubness problem, a mapping is learned from the semantic-descriptor space to the image-feature space that minimizes both one-to-one and pairwise distances between semantic embeddings and the image features. Secondly, to alleviate the domain-shift problem at test time, we propose a domain-adaptation method that finds correspondences between the semantic descriptors and the image features of test data. Thirdly, to reduce biased-ness in the GZSL setting, we propose scaled calibration on the classification scores of the seen classes to balance the performance on the seen and unseen categories. Finally, we evaluated our proposed approach on four standard ZSL datasets and compared our approach against state-of-the art methods followed by further analyzing the contribution of each component of our approach.

II Methodology

II-B Relational Matching

Our goal is to learn a mapping f(⋅)\mathbf{f}(\cdot) that maps a semantic descriptor ai\mathbf{a}_{i} to its corresponding image feature ϕ(xi)\mathbf{\phi}(\mathbf{x}_{i}). Here, xi\mathbf{x}_{i} is an image and ϕ(⋅)\mathbf{\phi}(\cdot) represents a CNN architecture that extracts a high-dimensional feature map. The mapping f(⋅){\bf f}(\cdot) is a fully-connected neural network. Since our goal is to make the embedded semantic descriptor close to the corresponding image feature, we use a least square loss function to minimize the difference. We also need to regularize the parameters of f(⋅)\mathbf{f}(\cdot). Including these costs and averaging over all the instances, our initial objective function L1\mathcal{L}_{1} is as follows:

where g(⋅)g(\cdot) is the regularization loss for the mapping function. The loss function L1\mathcal{L}_{1} minimizes the point-to-point discrepancy between the semantic descriptors and the image features. To account for the structural matching between the semantic-descriptor space and the image-feature space, we try to minimize the inter-class pairwise relations in these two spaces. Thus, we construct relational matrices for both the semantic descriptors and image features. The semantic relational matrix Da\mathbf{D}_{a} is established such that each element, [Da]uv=∣∣f(au)−f(av)∣∣22[\mathbf{D}_{a}]_{uv}=||{\bf f}({\bf a}^{u})-{\bf f}({\bf a}^{v})||_{2}^{2}, where au{\bf a}^{u} and av{\bf a}^{v} are semantic descriptors of seen categories uu and vv, respectively. The image feature relational matrix Dϕ\mathbf{D}_{\phi} is constructed such that each element, [Dϕ]uv=∣∣ϕ‾u−ϕ‾v∣∣22[\mathbf{D}_{\phi}]_{uv}=||\overline{\mathbf{\phi}}^{u}-\overline{\mathbf{\phi}}^{v}||_{2}^{2}, where ϕ‾u\overline{\mathbf{\phi}}^{u} and ϕ‾v\overline{\mathbf{\phi}}^{v} are mean representations of the categories uu and vv, respectively. ϕ‾u\overline{\mathbf{\phi}}^{u} can be represented as

where the summation is over the representations of class uu, and ∣Ytru∣|\mathcal{Y}^{u}_{tr}| is the cardinality of the training set of class uu. A similar formula holds for class vv. For structural alignment, we want the two relational matrices, Da\mathbf{D}_{a} and Dϕ\mathbf{D}_{\phi}, to be close to one another. Hence, we want to minimize the structural alignment loss function L2\mathcal{L}_{2},

where ∣∣⋅∣∣F2||\cdot||_{F}^{2} stands for the Frobenius norm. Combining the loss functions L1\mathcal{L}_{1} and L2\mathcal{L}_{2}, we have the total loss Ltotal\mathcal{L}_{total},

where ρ≥0\rho\geq 0 weighs the loss contribution of L2\mathcal{L}_{2}. Ltotal\mathcal{L}_{total} is to be optimized with respect to the parameters of the semantic-descriptor-to-visual-feature-space mapping f(⋅){\bf f}(\cdot).

II-C Domain Adaptation

This loss function enforces that CU{\bf C}{\bf U} produces the adapted semantic descriptors. However, a problem may exist that an instance in U{\bf U} corresponds to more than one descriptor in A{\bf A}. This would essentially result in a test sample corresponding to more than one category. To avoid that, we use an additional group-based regularization function L4\mathcal{L}_{4} using Group-Lasso,

where IcI_{c} corresponds to the indices of those rows in A{\bf A} that belong to the unseen class cc. Therefore, [C]Icj[{\bf C}]_{I_{c}j} is the vector consisting of the row indices from IcI_{c} and the jthj^{th} column. Since C{\bf C} is a correspondence matrix, some constraints should be enforced such as C≥0{\bf C}\geq\mathbf{0}, C1ou=1nu{\bf C}\mathbf{1}_{o_{u}}=\mathbf{1}_{n_{u}} and CT1nu=nuou1ou{\bf C}^{T}\mathbf{1}_{n_{u}}=\frac{n_{u}}{o_{u}}\mathbf{1}_{o_{u}}, where 1n\mathbf{1}_{n} is an n ⁣× ⁣1n\!\times\!1 vector of one’s. The second equality constraint is scaled by the factor nuou\frac{n_{u}}{o_{u}} to account for the difference in the number of instances in the mapped semantic-descriptor space and the image-feature space for the unseen categories. Hence, the domain adaptation optimization problem becomes

where λg\lambda_{g} weighs the loss function L4\mathcal{L}_{4}.

The above optimization problem is convex and can be efficiently solved using the conditional gradient method . The conditional gradient method requires solving a linear program as an intermediate step over the constraints C∈D={C:C≥0,C1ou=1nu,CT1nu=nuou1ou}{\bf C}\in\mathcal{D}=\{{\bf C}:{\bf C}\geq{\bf 0},{\bf C}\mathbf{1}_{o_{u}}=\mathbf{1}_{n_{u}},{\bf C}^{T}\mathbf{1}_{n_{u}}=\frac{n_{u}}{o_{u}}\mathbf{1}_{o_{u}}\} as shown in Algorithm 1. The linear program of finding the intermediate variable Cd{\bf C}_{d} in Algorithm 1 can be easily solved using a network simplex formulation of the earth-mover’s distance problem .

Once the final solution of the correspondence matrix C0{\bf C}_{0} in Algorithm 1 is obtained, we inspect C0{\bf C}_{0}. For each test instance, we assign the class correspondence to the highest value of the correspondence variable. This is done for all the test instances. The new semantic descriptors are obtained by taking the mean of the feature instances belonging to the corresponding class. The adapted semantic descriptors are then stacked vertically in the matrix A′{\bf A}^{\prime}.

II-D Scaled Calibration

In the GZSL setting, it is known that the classification results are biased towards the seen categories . To counteract the bias, we propose the use of multiplicative calibration on the classification scores. In our case, we use 1-Nearest Neighbor (1-NN) with the Euclidean distance metric as the classifier. The classification score for a test point is given by the Euclidean distance of the test image feature to the mapped semantic descriptor of a category. For a test point x{\bf x}, we adjust the classification scores on the seen categories as follows

III Experimental Results

Following the previous experimental settings , we used the following four datasets for evaluation: AwA2 (Animal with Attributes) contains 37,322 images of 50 classes of animals. 40 classes of animals are considered to be the seen categories while 10 classes of animals are considered to be the unseen categories. Each class is associated with a 85-dimensional continuous semantic descriptor. aPY (attribute Pascal and Yahoo) consists of 20 seen categories and 12 unseen categories. Each category has an associated 64-dimensional semantic descriptor. CUB (Caltech-UCSD Birds-200-2011) is a fine-grained dataset consisting of 11,788 images of birds. For evaluation, all the bird categories are split into 150 seen classes and 50 unseen classes. Each class is associated with a 312-dimensional continuous semantic descriptor. SUN (Scene UNderstanding database) consists of 14340 scene images. Among these, 645 scene categories are selected as seen categories while 72 categories are selected as unseen categories and it consists of a 102-dimensional semantic descriptor.

For the purpose of evaluation, we used class-wise accuracy because it prevents dense-sampled classes from dominating the performance. Accordingly, class-wise accuracy is averaged as follows

where ∣Y∣|\mathcal{Y}| is the number of testing classes. In the GZSL case, class-wise accuracy of both seen and unseen classes are obtained separately and then averaged using harmonic mean HH . This is done so that the performance on seen classes does not dominate the overall accuracy,

where accsacc_{s} and accuacc_{u} are the class-wise accuracy on seen and unseen categories, respectively. In the GZSL classification setting, the search space of predicted categories consists of both seen and unseen categories. Based on and for fair comparison, a single trial of experimental results on a large batch of training and testing dataset is reported.

For the experiments, we used a two-layer feedforward neural network for the semantic embedding f(⋅){\bf f}(\cdot). The dimensionality of the hidden layer was chosen as 1600, 1600, 1200 and 1600 for the AwA2, aPY, CUB and SUN datasets, respectively. The activation used was ReLU. The image features used were the ResNet-101. We compared different variations of our proposed method with previous approaches. OURS-R variation is with the training stage including the structural loss L2\mathcal{L}_{2}. OURS-RA includes the structural loss as well as the domain adaptation stage including the loss functions L3\mathcal{L}_{3} and L4\mathcal{L}_{4}. OURS-RC includes the structural loss as well as the calibrated testing stage. OURS-RAC includes all the components of structural loss, domain adaptation and calibrated testing. Without all these components, the proposed method reduces to the Deep Embedding Model (DEM) baseline. The parameters (λr,ρ,λg,γ)(\lambda_{r},\rho,\lambda_{g},\gamma) for the AwA2, aPY, CUB and SUN datasets are set as (10−3,10−1,10−1,1.1)(10^{-3},10^{-1},10^{-1},1.1), (10−4,10−1,10−1,1.1)(10^{-4},10^{-1},10^{-1},1.1), (10−2,0,10−1,1.1)(10^{-2},0,10^{-1},1.1) and (10−5,10−1,10−1,1.1)(10^{-5},10^{-1},10^{-1},1.1), respectively. For the OURS-RAC variation, we used different calibration parameter values of 0.98,1.1,0.97,0.9990.98,1.1,0.97,0.999 for the AwA2, aPY, CUB and SUN datasets, respectively. ρ\rho was set to for the CUB dataset because it is a fine-grained dataset and since the categories are very close to each other in the feature space, structural matching does not provide additional information. In Table I, we reported class-wise accuracy results for the conventional unseen classes setting (tr), generalized unseen classes setting (u), generalized seen classes setting (s), and the Harmonic mean (H) of the generalized accuracies.

From the table, we observed that our proposed approach outperforms previous methods by a large margin in the generalized harmonic mean setting. To be more specific, our proposed method produces an improvement of around 20%, 23%, 10% and 16% harmonic mean accuracy over the previous best approach for the AwA2, aPY, CUB and SUN datasets, respectively. The large improvement in performance can be attributed to our three-step procedure for improvement. Using only the structural matching (OURS-R), we produced better results than previous approaches except for the CUB dataset, where it produces a harmonic mean accuracy of about 28%. This is because CUB requires minute fine-grained feature extraction. Additional usage of domain adaptation (OURS-RA) and calibrated testing (OURS-RC) produced much better results than OURS-R for all the datasets. However, domain adaptation produced better result than the calibration procedure. This is because our correspondence-based approach produced class-specific adaptation of the unseen class semantic embeddings. The scaled-calibration procedure is not class-specific and just differentiates between seen and unseen classes. It also does not adapt to the test data.

It is to be noted that the difference in performance between OURS-RA and OURS-RAC is negligible. This is because the domain adaptation step transforms the unseen semantic embeddings away from the seen categories towards the unseen categories, thus reducing the bias towards the seen categories and rendering further calibration ineffective. The effect of domain adaptation is visualized in Fig. 3 for the AwA2 dataset using t-SNE . In Fig. 3(a), the unseen class semantic embeddings (blue) remained very close to the seen class features (maroon). However, with the domain adaptation step, the unseen class semantic embeddings get transformed to near the centre of unseen class feature clusters (green) as shown in Fig. 3(b).

We also analyzed the effect of the structural matching by varying ρ∈{10−3,10−2,10−1,100,101,102}\rho\in\{10^{-3},10^{-2},10^{-1},10^{0},10^{1},10^{2}\} and observed how the class-wise accuracy changes. We carried out experiments using the AwA2 and SUN datasets, the results of which are reported in Fig. 4. We also reported the DEM baseline (ρ=0\rho=0) in dotted lines. From the plots, the Conventional Unseen and the Generalized Seen accuracies are better than or equal to the baseline for only a small range of ρ\rho. On the other hand, the Generalized Unseen accuracy is greater than the baseline over a large range of ρ\rho for the AwA2 dataset while it oscillated about the baseline for the SUN dataset. For the SUN dataset, we do not have a significant gain over the baseline because SUN is a fine-grained dataset where structural matching does not carry additional information. The goal of structural regularization is to exploit the pairwise relations among classes so as to generalize better to novel classes. Therefore, we did not see huge difference in performance from the baseline for the Generalized Seen accuracy. Surprisingly, there was a drop in conventional unseen accuracy as ρ\rho was increased. This might be probably because there was no overlap between the classes used for testing and the classes used for structural matching. This is not the case though in the generalized setting.

We also studied the effect of varying the calibration parameter γ\gamma on the generalized accuracy for the AwA2 and SUN datasets. The results are shown in Fig. 5. As expected, the generalized unseen accuracy increases and the generalized seen accuracy decreases with increasing γ\gamma. The peak of the harmonic mean accuracy was observed close to when the seen and unseen accuracies became equal. The maximum unseen accuracy is less than the maximum seen accuracy for the AwA2 dataset because the unseen classes are less separated and therefore more difficult to classify. The situation is reversed for the SUN dataset where the maximum unseen accuracy is more than the maximum seen accuracy.

We also reported convergence results of the test accuracy with respect to the number of epochs for both the AwA2 and the SUN datasets in Figs. 6 and 7, respectively. We used the OURS-R variation with ρ=0.1\rho=0.1 to compare with the DEM baseline. The convergence rate for the baseline and OURS-R variation seems to be similar in all the settings for both datasets. However, our steady-state values were higher for the generalized unseen and generalized harmonic mean setting. For the conventional unseen and generalized seen setting, our steady-state value was less than the baseline. The reason is explained previously while describing performance sensitivity to ρ\rho.

We also studied the effect of varying the number of test unseen samples per class on the generalized harmonic mean accuracy. We used OURS-RA variation of our model for this study. ρ=0.1\rho=0.1 was set for the experiments on the AwA2 (blue color) and the SUN (yellow color) datasets and the result was reported in Fig. 8. When the fraction is 0.01 for the SUN dataset, the number of samples in some classes becomes zero and therefore the performance is not reported. From the results, it is seen that the test accuracy was stable with change in the fraction of total number of samples used for testing. There is a slight increase in accuracy with decreasing number of samples, which is surprising because domain adaptation would perform poorly with less number of samples. However, this effect is nullified since the probability of including challenging examples is reduced and so we observed a slight improvement in performance.

We also studied how the test performance varies as the number of seen classes for training is reduced for the AwA2 dataset using OURS-R model. We set ρ=0.1\rho=0.1 and reported results over 5 trials in Fig. 9. We observed that the change in the seen-class accuracy is not much because the training and testing distributions are the same. The conventional unseen-class accuracy dips by a large amount as the number of training classes decreases because there is less representative information to be transferred to novel categories. However, we obtained a peak for the generalized unseen accuracy results at a fraction of 0.4 of the number of seen classes. This is because as the number of training classes decreases, the amount of representative information decreases, causing decrease in performance. On the other hand, less number of seen classes implies less bias towards seen categories and improvement of unseen-class accuracies. Also, there is large performance variation for unseen-class accuracy because training and testing distributions are different and the performance can vary depending on how related are the training classes to the unseen classes in a trial.

We also performed experiments to find whether the OURS-R variant reduces hubness compared to DEM. The hubness of a set of predictions is measured using the skewness of the 1-Nearest-Neighbor histogram (N1N_{1}). The N1N_{1} histogram is a frequency plot for N1[i]N_{1}[i] of the number of times a search solution ii (in our case a class attribute) is found as the Nearest Neighbor for the test samples. Less skewness of N1N_{1} histogram implies less hubness of the predictions. We used the test samples of the unseen classes in the generalized setting for both DEM and OURS-R on the AwA2 and the aPY datasets. We used ρ=0.1\rho=0.1 and reported results averaging over 5 trials in Table II. From the results, OURS-R method produced less skewness of the N1N_{1} histogram on both the datasets. This implies that using the additional structural term reduces hubness and therefore the curse of dimensionality is reduced.

IV Conclusion

This paper proposed a three-step approach to improve the performance of zero-shot learning for image classification. The three-step approach involved exploiting structural information in data, domain adaptation to unseen test samples and calibration of classification scores. When the proposed method was applied to standard datasets of zero-shot image classification, it outperformed previous methods by a large margin, where the most effective component was the domain adaptation step.

References