A Unified Semantic Embedding: Relating Taxonomies and Attributes
Sung Ju Hwang, Leonid Sigal
Introduction
Semantic approaches have gained a lot of attention recently for object categorization, as object categorization problems became more focused on large-scale and fine-grained recognition tasks and datasets. Attributes and semantic taxonomies are two of the popular semantic sources which impose certain relations between the category models. While many techniques have been introduced to utilize each of the individual semantic sources for object categorization, no unified model has been proposed to relate them.
We propose a unified semantic model where we can learn to place categories, supercategories, and attributes as points (or vectors) in a hypothetical common semantic space. Further, we propose a discriminative learning framework based on dictionary learning and large margin embedding, to learn each of these semantic entities to be well separated and pseudo-orthogonal, such that we can use them to improve visual recognition tasks such as category or attribute recognition.
However, having semantic entities embedded into a common space is not enough to utilize the vast number of relations that exist among them. Thus, we impose a graph-based regularization between the semantic embeddings, such that each semantic embedding is regularized by sparse combination of auxiliary semantic embeddings.
The observation we make to draw the relation between the categories and attributes, is that a category can be represented as the sum of its super category + the category-specific modifier, which in many cases can be represented by a combination of attributes. Further, we want the representation to be compact. Instead of describing a dalmatian as a domestic animal with a lean body, four legs, a long tail, and spots, it is more efficient to say it is a spotted dog (Figure 1). It is also more exact since the higher-level category dog contains all general properties of different dog breeds, including indescribable dog-specific properties, such as the shape of the head, and its posture. This exemplifies how a human would describe an object, to efficiently communicate and understand the concept.
This additional requirement imposed on the discriminative learning model would guide the learning such that we obtain not just the optimal model for class discrimination, but to learn a semantically plausible model which has a potential to be more robust and human-interpretable; we call this model Unified Semantic Embedding (USE).
Learning a unified semantic embedding space
To ensure that the projected instances have higher similarity to its own category embedding than to others, we add discriminate constraints, which are large-margin constraints on distance: . This translates to the following discriminative loss:
where is the columwise concatenation label embedding vectors, such that denotes column of . After replacing the generative loss in the ridge regression formula with the discriminative loss, we get the following discriminative learning problem:
where regularizes and from going to infinity. This is one of the most common objectives used for learning discriminative category embeddings for multi-class classification , while ranking loss-based models have been also explored for .
While our objective is to better categorize entry level categories, categories in general can appear in different semantic granularities. For example, a zebra could be both an equus, and an odd-toed ungulate. To learn the embeddings for the supercategories, we map each data instance to be closer to its correct supercategory embedding than to its siblings: where denotes the set of superclasses at all levels for class , and is the set of its siblings. The constraints can be translated into the following loss:
Attributes.
Attributes can be considered as a normalized basis vectors for the semantic space, whose combination represents a category. Basically, we want to maximize the correlation between the projected instance that possess the attribute, and its correct attribute embedding, as follows:
where is the set of all attributes for class , is the margin (we simply use a fixed value of ), is the label indicating presence/absence of each attribute for the training instance, and is the embedding vector for attribute .
Semantic regularization.
The previous multi-task formulation enables to implicitly associate the semantic entities, with the shared data embedding . However, we want to further explicitly impose structural regularization on the semantic embeddings , based on the intuition that an object class can be represented as its parent level class + a sparse combination of attribute as follows:
where is the aggregation of all attribute embeddings , is the set of children classes for class , is the sparsity parameter, and is the number of categories. is the matrix whose column vector is the reconstruction weight for class , is the set of all sibling classes for class , and is the parameters to enforce exclusivity. We require to be non-negative, since it makes more sense to describe an object with attributes that it has, rather than attributes it does not have.
The exclusive regularization term is used to prevent the semantic reconstruction for class from fitting to the same attributes fitted by its parents and siblings. Such regularization will enforce the categories to be ‘semantically’ discriminated as well. With the sparsity regularization enforced by , the simple sum of the two weights will prevent the two (super)categories from having high weight for a single attribute, which will let each category embedding to fit to exclusive attributes.
Unified semantic embeddings with semantic regularization.
After augmenting the categorization objective in Eq. 2 with the superclass and attributes loss and the sparse-coding based regularization in Eq. LABEL:comb2, we obtain the following multitask learning formulation:
where is the number of supercategories, is ’s column, and and are parameters to balance between the main and auxiliary tasks, and discriminative and generative objective.
Eq. 6 can also be used for knowledge transfer when learning a model for a novel set of categories, by replacing in with , learned on class set to transfer the knowledge from.
Numerical optimization.
Eq. 6 is not jointly convex, and has both discriminative and generative terms. The problem is similar to the problem in , and can be optimized using a similar alternating optimization, while alternating between the following two convex sub-problems: 1) Optimization of the data embedding and parameters , and 2) Optimization of the category embedding .
Results
We validate our method for multiclass categorization performance and knowledge transfer on the Animals with Attributes dataset , which consists of images on animal classes, with class-level attributes Attributes are defined on color (black, orange), texture (stripes, spots), parts (longneck, hooves), and other high-level behavioral properties (slow, hibernate, domestic) of the animals.. We use the Wordnet hierarchy to generate supercategories. Since there is no fixed training/test split, we use {30,30,30} random split for training/validation/test. For the features, we use the provided -D DeCAF features obtained from a deep convolutional neural network.
We implement multiple variants of our model to analyze the impact of each semantic entity and the proposed regularization. 1) LME-MTL-S: The multitask semantic embedding model learned with supercategories. 2) LME-MTL-A: The multitask embedding model learned with attributes. 3) USE-No Reg.: The unified semantic embedding model learned using both attributes and supercategories, without semantic regularization. 4) USE-Reg: USE with the sparse coding regularization. We find the optimal parameters for the USE model by cross-validation on the validation set.
We first evaluate the USE framework for categorization performance. We report the average classification performance and standard error over 5 random training/test splits in Table 1, using both flat hit@k, which is the accuracy at the top-k prediction made, and hierarchical precision@k from , which is a precision the given label is correct at , at all levels.
The implicit semantic baselines, ALE-variants, underperformed even the ridge regression baseline with regard to the top-1 classification accuracy We did extensive parameter search for the ALE variants., while they improve upon the top-2 and hierarchical precision. This shows that hard-encoding structures in the label space do not necessarily improve the discrimination performance, while it helps to learn a more semantic space.
Explicit embedding of semantic entities using our method improved both the top-1 accuracy and the hierarchical precision, with USE variants achieving the best performance in both. USE-Reg. made substantial improvements on flat hit and hierarchical precision @ 5, which shows the proposed regularization’s effectiveness in learning a semantic space that also discriminates well.
Qualitative analysis.
Besides learning a space that is both discriminative and generalizes well, our method’s main advantage is its ability to generate compact, semantic description of each category it has learned. This is a great caveat, since in most models, including the state-of-the art deep convolutional networks, humans cannot understand what has been learned; by generating human-understandable explanation, our model can communicate with the human, allowing understanding of rationale behind the categorization decision, and to possibly provide feedback for correction.
To show the effectiveness of using supercategory+attributes in the description, we report the learned reconstruction for our model, compared against the description generated by ground-truth attributes in Table 2. The results show that our method generates compact description of each category, focusing on its discriminative attributes. For example, our method selects flippers for otter, and stripes for skunk, instead of common nondescriminative attributes such as tail. Further, our method selects attributes for each supercategory, while there is no provided attribute label for supercategories.
One-shot/Few-shot learning.
Our method is expected to be especially useful for few-shot learning, by generating a richer description than existing methods that approximate the new input category using only trained categories, or attributes. For this experiment, we divide the categories into predefined training/test split. USE-Reg achieves the most improvement, improving two-shot result on AWA-DeCafe from 38.93% to 49.87%. Most learned reconstruction look reasonable, and fit to discriminative traits that help to discriminate between the test classes.