Knowledge-Embedded Representation Learning for Fine-Grained Image Recognition
Tianshui Chen, Liang Lin, Riquan Chen, Yang Wu, Xiaonan Luo
Introduction
Humans perform object recognition task based on not only the object appearance but also the knowledge acquired from daily lives or professions. Usually, this knowledge refers to a comprehensive visual concept organization including category labels and their attributes. It is extremely beneficial to fine-grained image classification as attributes are always key to distinguish different subordinate categories. For example, we might know from a book that a bird of category “Bohemian Waxwing” has a masked head with color black and white and wings with iridescent feathers. With this knowledge, to recognize the category “Bohemian Waxwing” given a bird image, we might first recall the knowledge, attend to the corresponding parts to see whether it possesses these attributes, and then perform reasoning. Figure 1 illustrates an example of how the professional knowledge aids fine-grained image recognition.
Conventional approaches for fine-grained image classification usually neglect this knowledge and merely rely on low-level image cues for recognition. These approaches either employ part-based models Zhang et al. (2014) or resort to visual attention networks Liu et al. (2016) to locate discriminative regions/parts to distinguish subtle differences among different subordinate categories. However, part-based models involve heavy annotations of object parts, preventing them from application to large-scale data, while visual attention networks can only locate the parts/regions roughly due to the lack of supervision or guidance. Recently, He and Peng (2017a) utilize natural language descriptions to help search the informative regions and combine with vision stream for final prediction. This method also integrates high-level information, but it directly models image-language pairs and requires detailed language descriptions for each image (e.g., ten sentences for each image in He and Peng (2017a)). Different from these methods, we organize knowledge about categories and part-based attributes in the form of knowledge graph and formulate a Knowledge-Embedded Representation Learning (KERL) framework to incorporate the knowledge graph into image feature learning to promote fine-grained image recognition.
To this end, our proposed KERL framework contains two crucial components: i) a Gated Graph Neural Network (GGNN) Li et al. (2015) that propagates node message through the graph to generate knowledge representation and ii) a novel gated mechanism that integrates this representation with image feature learning to learn attribute-aware features. Concretely, we first construct a large-scale knowledge graph that relates category labels with part-level attributes as shown in Figure 2. By initializing the graph node with information of a given image, our KERL framework might implicitly reason about the discriminative attributes for the image and associate these attributes with feature maps. In this way, our KERL framework can learn feature maps with a meaningful configuration that the highlighted regions finely associate with the relevant attributes in the graph. For example, the learned feature maps of samples from category “Bohemian Waxwing” always highlight the regions of head and wings, because these regions relate to attributes “head pattern: masked”, “head color: white & black” and “wing: iridescent” that are key to distinguish this category from others. This characteristic also provides insight into why the framework improves performance.
The major contributions of this work are summarized to three-fold: 1) This work formulates a novel Knowledge-Embedded Representation Learning framework that incorporates high-level knowledge graph as extra guidance for image representation learning. To the best of our knowledge, this is the first work to investigate this point. 2) With the guidance of knowledge, our framework can learn attribute-aware feature maps with a meaningful and interpretable configuration that the highlighted regions are finely related to the relevant attributes in the graph, which can also explain performance improvement. 3) We conduct extensive experiments on the widely used Caltech-UCSD bird dataset Wah et al. (2011) and demonstrate the superiority of the proposed KERL framework over the leading fine-grained image classification methods.
Related Work
We review the related work in term of two research streams: fine-grained image classification and knowledge representation.
With the advancement of deep learning He et al. (2016); Simonyan and Zisserman (2014), most works rely on deep Convolutional Neural Networks (CNNs) to learn discriminative features for fine-grained image recognition, which exhibit a notable improvement compared with conventional hand-crafted features He and Peng (2017b); Lin et al. (2015b). To better capture subtle visual difference for fine-grained classification, bilinear models Lin et al. (2015b); Gao et al. (2016); Kong and Fowlkes (2017) is proposed to compute high-order representation that can better model local pairwise feature interactions by two independent sub-networks. Another common approach for distinguishing subtle visual difference among sub-ordinate categories is first locating discriminative regions and then learning appearance model conditioned on these regions Zhang et al. (2014); Huang et al. (2016). However, these methods involve in heavy annotations of object parts, and moreover, manually defined parts may not be optimal for the final recognition. Instead, He et al. He and Peng (2017a) adopt salient region localization techniques Chen et al. (2016); Zhou et al. (2016) to automatically generate bounding box annotations of the discriminative regions. Recently, visual attention models Mnih et al. (2014); Wang et al. (2017); Chen et al. (2018); Liu et al. (2018) have been intensively proposed to automatically search the informative regions, and some works also apply this technique to fine-grained recognition task Liu et al. (2016); Zheng et al. (2017); Peng et al. (2018). Liu et al. (2016) introduce a reinforcement learning framework to adaptively glimpse local discriminative regions and propose a greedy reward strategy to train the framework with image-level annotations. Fu et al. (2017) further introduce a recurrent attention convolutional neural network to recursively learn the attentional regions at multiple scales and region-based feature representation. Liu et al. (2017) utilize part-level attribute to guide locating the attentional regions, which is related to ours. However, we organize the category-attribute relationships in the form of knowledge graph and implicitly reason discriminative attributes on the graph rather than using object-attribute pairs directly.
2 Knowledge Representation
Learning knowledge representation for visual reasoning increasingly receives attention as it benefits various tasks Malisiewicz and Efros (2009); Lao et al. (2011); Zhu et al. (2014); Lin et al. (2017) in vision community. For instance, Zhu et al. (2014) learn a knowledge base using a Markov Logic Network and employ first-order probabilistic inference to reason the object affordances. These approaches usually involve in hand-crafted features and manually-defined propagation rules, preventing them from end-to-end training. Most recently, a series of efforts are dedicated to adapt neural networks to process graph-structured data Duvenaud et al. (2015); Niepert et al. (2016). For example, Niepert et al. (2016) sort the nodes in the graph based on the graph edges to regular sequence and directly feed the node sequence to a standard CNN for feature learning. These methods are tried on small, clean graphs such as molecular datasets Duvenaud et al. (2015) or are used to encode the contextual dependencies for vision tasks Liang et al. (2016).
GGNN Li et al. (2015) is a fully differentiable recurrent neural network architecture for graph-structured data, which recursively propagates node message to its neighbors to learn node-level features or graph-level representation. Several works have developed a series of graph neural network variants and successfully apply them to various tasks, such as 3DGNN for RGBD semantic segmentation Qi et al. (2017), model-based GNN for situation recognition Li et al. (2017), and GSNN for multi-label image recognition Marino et al. (2017). Among these works, GSNN Marino et al. (2017) is mostly related to ours in the spirit of GGNN based knowledge graph encoding, but it simply concatenates image and knowledge features for image classification. In contrast, we develop a novel gated mechanism to embed the knowledge representation into image feature learning to enhances the feature representation. Besides, our learned feature maps exhibit insightful configurations that the highlighted regions finely accord with the semantic attributes in the graph, which also provide insight to explain performance improvement.
KERL Framework
In this section, we first briefly review the GGNN and present the construction of our knowledge graph that relates category labels with their part-level attributes. Then, we introduce our KERL framework in detail, which consists of a GGNN for knowledge representation learning and a gated mechanism to embed knowledge into discriminative image representation learning. An overall pipeline of the framework is illustrated in Figure 3.
We briefly introduce the GGNN Li et al. (2015) for completeness. GGNN is recurrent neural network architecture that can learn features for arbitrary graph-structured data by iteratively updating node features. For the propagation process, the input data is represented as a graph , in which is the node set and is the adjacency matrix that denotes the connections among nodes in the graph. For each node , it has a hidden state at time step , and the hidden state at is initialized by the input feature vector that depends on the problem in hand. Thus, the basic recurrent process is formulated as
2 Knowledge Graph Construction
The knowledge graph refers to an organization of a repository of visual concepts including category labels and part-level attributes, with nodes representing the visual concepts and edges representing their correlations. The graph is constructed based on the attribute annotations of the training samples. An example knowledge graph for the Caltech-UCSD bird dataset Wah et al. (2011) is presented in Figure 2.
Visual concepts. A visual concept refers to either a category label or an attribute. The attribute is an intermediate semantic representation of objects, and usually, it is key to distinguish two subordinate categories. Given a dataset that covers object categories and attributes, the graph has a node set with elements.
Correlation. The correlation between a category label and an attribute indicates whether this category possesses the corresponding attribute. However, for the fine-grained task, it is common that merely some instances of a category possess a specific attribute. For example, for a specific category, it is possible that one instance has a certain attribute, but another instance does not have. Thus, such category/attribute correlation is uncertain. Fortunately, we can assign an attribute/object instance pair with a score that denotes how likely this instance has the attribute. Then, we can sum up the scores of attribute/object instance pairs for all instances belonging to a specific category and obtain a score to denote the confidence that this category has the attribute. All the scores are linearly normalized to $C\times A\mathbf{S}$. Note that no connection exists between two object category nodes or between two attribute nodes; thus the complete adjacency matrix can be expressed as
where denotes a zero matrix of size . In this way, we can construct a knowledge graph .
3 Knowledge Representation Learning
After building the knowledge graph, we employ the GGNN to propagate node message through the graph and compute a feature vector for each node. All the feature vectors are then concatenated to generate the final representation for the knowledge graph.
We initialize the node referring to category label with a score that represents the confidence of this category being presented in the given image, and the node referring to each attribute with zero vector. The score vector for all categories is estimated by a pre-trained classier that will be introduced in detail in section 4.1. Thus, the input feature for each node can be represented as
where is a zero vector with dimension . As discussed above, messages of all nodes are propagated to each other during the propagation process. With the computational process of Equation 1, are used to propagate message from a certain node to its neighbors, and we use matrix for reverse message propagation. Thus, the adjacency matrix is .
For each node , its hidden state is initialized using , and at timestep , the hidden state is updated using the propagation process as Equation (1), expressed as
At each iteration, the hidden state of each node is determined by its history state and the messages sent by its neighbors. In this way, each node can aggregate information from its neighbors and simultaneously transfer its message to its neighbors. This process is shown in Figure 3. After iterations, the message of each node has propagated through the graph, and we can get the final hidden state for all nodes in the graph, i.e., . Similar to Li et al. (2015), the node-level feature is computed by
where is an output network that is implemented by a fully-connected layer. Finally, these features are concatenated to produce the final knowledge representation .
4 United Representation Learning
In this part, we introduce the gated mechanism that embeds the knowledge representation to enhance image representation learning.
Image feature extraction. We start by introducing the image feature extraction. As compact bilinear model Gao et al. (2016) works well on fine-grained image classification, we straightforwardly apply this model to extract image features. Specifically, given an image, we utilize a fully convolutional network (FCN) to extract feature maps with a size of , and a compact bilinear operator to produce feature maps . Note that we do not perform sum pooling like Gao et al. (2016) thus the size of is . For fair comparisons with existing works, we employ the convolutional layers of the VGG16-Net to implement the FCN and follow the default setting as Gao et al. (2016) to set as 8192.
Gao et al. (2016) treats all features equally important and simply performs sum pooling to obtain -dimensional features for prediction. In the context of fine-grained image classification, it is crucial to attend to the discriminative regions to capture the subtle difference between different subordinate categories. In addition, the knowledge representation encodes category-attribute correlation and it may capture the discriminative attributes. Thus, we embed this representation into image feature learning to learn feature corresponding to this attributes. Specifically, we introduce a gated mechanism that optionally allows the informative features through while suppressing the non-informative features under the guidance of knowledge, which can be formulated as
where is the feature vector at location . acts as a gated mechanism that decides which location is more important. is a neural network that takes the concatenation of and as input and outputs a -dimensional real-value vector. It is implemented by two stacked fully connected layers in which the first one is 10752 (8192 + 5125) to 4096 followed by the hyperbolic tangent function while the second one is 4096 to 8192. The feature vector is then fed into a simple fully-connected layer to compute the score vector for the given image.
Experiments
Datasets. We evaluate our KERL framework and the competing methods on the Caltech-UCSD bird dataset Wah et al. (2011) that is the most widely used benchmark for fine-grained image classification. The dataset covers 200 species of birds, which contains 5,994 images for training and 5,794 for test. Except for the category label, each image is further annotated with 1 bounding box, 15 part key-points, and 312 attributes. As shown in Figure 4, the dataset is extremely challenging because birds from similar species may share very similar visual appearance while birds within the same species undergo drastic changes owing to complex variations in scales, viewpoints, occlusion, and background. In this work, we evaluate the methods in two settings: 1) “bird in image”: the whole image is fed into the model at training and test stages, and 2) “bird in bbox”: the image region at the bounding box is fed into the model at training and test stages.
Implementation details. For the GGNN, we utilize the compact bilinear model released by work Gao et al. (2016) to produce the scores to initialize the hidden states. For fair comparisons, the model is implemented with VGG16-Net and trained on the training part of the Caltech-UCSD bird dataset. The dimension of the hidden state is set to 10 and that of the output feature is set to 5. The iteration time is set to 5. The KERL framework is jointly trained using the cross-entropy loss. All components of the framework are trained with SGD except GGNN that is trained with ADAM following Marino et al. (2017).
2 Comparison with State-of-the-Art Methods
In this subsection, we compare our KERL framework with 16 state-of-the-art methods, among which, some use merely image-level labels, and some also use bounding box/parts annotations; thus we also present this information for fair and direct comparisons. The methods are evaluated in both two settings, i.e., “bird in image” and “bird in bbox”, and the results are reported in Table 1. For the “bird in bbox” setting, the previous well-performing methods are PN-CNN and SPDA-CNN that achieve the accuracies of 85.4% and 85.1% respectively, but they require strong supervision of ground truth part annotations. The accuracy of B-CNN is also up to 85.1%, but it relies on a very high-dimensional feature representation (250k dimensions). In contrast, the KERL framework requires no ground truth part annotations and utilize a much lower-dimensional feature representation (i.e., 8,192 dimensions), but it achieves an accuracy of 86.6% that outperforms all previous state-of-the-art methods. For the “bird in image” setting, most existing methods explicitly search discriminative regions and aggregate deep features of these regions for classification. For example, RA-CNN recurrently discovers image regions over three scales and achieves an accuracy of 85.3%. Besides, CVL combines detailed human-annotated text description for each image with the visual features to further improve the accuracy to 85.6%. Different from them, the KERL framework learns knowledge representation that encodes category-attribute correlations and incorporates this representation for feature learning. In this way, our method can learn more discriminative attribute-related features, leading to improvement in performance, i.e., 86.3% in accuracy. Note that AGAL also employs part-level attributes for fine-grained classification, but it achieves accuracies of 85.5% and 85.4% in two settings, respectively, much worse than ours. These comparisons well demonstrate the effectiveness of the KERL framework method over existing algorithms.
Attention-based methods aggregate features of both image and located regions to promote fine-grained classification, and our results reported above merely use image features. Our KERL framework can learn feature maps that highlight the regions related to discriminative attributes, as discussed in section 4.4; thus, we also aggregate features of the highlighted regions to improve performance. Specifically, we sum up the feature values across channels to get a score at each location, and draw a region with a size of a centered at each location. We adopt non-maximum suppression to exclude the seriously overlapped regions and select top three ones. Three corresponding regions with a size of (16 mapping between the original image and feature map) in the image are cropped, resized to and fed to the VGG16 net to extract feature, respectively. The features are concatenated and fed to a fully-connected layer to compute the score vector, which is further averaged with the results of KERL to achieve the final results. It boosts the accuracies to 86.8% and 87.0% in two settings respectively.
3 Contribution of Knowledge Embedding
Note that our KERL framework employs CB-CNN Gao et al. (2016) as the baseline. Here, we emphasize the comparison with this baseline method to demonstrate the significance of knowledge embedding knowledge. As shown in Table 2, the CB-CNN achieves accuracies of 84.6% and 85.0% in “bird in bbox” and “bird in image” settings. By embedding the knowledge representation, the KERL framework boosts the accuracies to 86.6% and 86.3%, improving those of the CB-CNN by 2.0% and 1.3%, respectively.
To further clarify the contribution of knowledge guided feature selection, we implement two more baseline methods: self-guided feature learning and feature concatenation.
Comparison with self-guided feature learning. To better verify the benefit of embedding knowledge for feature learning, we conduct an experiment that removes the GGNN and only feeds the image features to the gated neural network, with other components left unchanged. The comparison results are presented in Table 2. It merely exhibits minor improvement over the baseline CB-CNN as it does not incur additional information but only increasing the complexity of the model. As expected, it performs much worse than ours.
Comparison with feature concatenation. To validate the benefit of our knowledge embedding method, we further conduct an experiment that incorporates knowledge by simply concatenating the image and graph feature vectors, followed by a fully-connected layer for classification. As shown in Table 2, directly concatenating image and graph features can achieve accuracies of 85.4% and 85.5% in the two settings, which is slightly better than the original CB-CNN but still much worse than ours. This indicates our knowledge incorporation method can make better use of knowledge to facilitate fine-grained image classification.
4 Representation Visualization
With knowledge embedding, our KERL framework can learn feature maps with an insightful configuration that the highlighted regions are always related to relevant attributes. Here, we visualize the feature maps before sum pooling to better evaluate this point in Figure 5. We sum up the feature values across channels at each location and normalize them to $$. At each row, we present the learned feature maps of several samples taken from a specific category and a sub-graph that shows the correlations of this category with its attributes. We find that the highlighted regions for samples of the same category refer to the same semantic parts, and these parts finely accord with the attributes that well distinguish this category from others. Taking the category of “Sayornis” as example, our KERL framework consistently highlights the regions of throats and wings for all samples, which correspond to two key attributes, i.e., “throat: buff & white” and “wing: buff & brown” (highlighted with orange circle in Figure 5). This suggests our KERL framework can learn attribute-aware features that can better capture subtle differences between different subordinate categories. Also, it can provide an explanation for the performance improvement of our framework.
To clearly verify that it is the knowledge embedding that brings about such appealing characteristic, we further visualize the feature maps generated by the CB-CNN model in Figure 6. We visualize the samples the same with those of the first two categories in Figure 5 for direct comparison. It is observed that some highlighted regions lie in the background and some scatter over the whole body of the birds.
Conclusion
In this paper, we propose a novel Knowledge-Embedded Representation Learning (KERL) framework to incorporate knowledge graph as extra guidance for image feature learning. Specifically, the KERL framework consists of a GGNN to learn the graph representation, and a gated neural network to integrate this representation into image feature learning to learn attribute-aware features. Besides, our framework can learn feature maps with an insightful configuration that the highlighted regions are always related to the relevant attributes in the graph, and this can well explain the performance improvement of our KERL framework. Experiments and evaluations conducted on the Caltech-UCSD bird dataset well demonstrate the superiority of our KERL framework over existing state-of-the-art methods. It is an early attempt to embed high-level knowledge into the modern deep network to improve fine-grained image classification, and we hope it can provide a step towards the integration of knowledge and traditional computer vision frameworks.