Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification

Renchun You, Zhiyao Guo, Lei Cui, Xiang Long, Yingze Bao, Shilei Wen

Introduction

Multi-label image classification (MLIC) and multi-label video classification (MLVC) are important tasks in computer vision, where the goal is to predict a set of categories present in an image or a video. Compared with single-label classification (e.g. assigns one label to an image or video), multi-label classification is more useful in many applications such as internet search, security surveillance, robotics, etc. Since MLIC and MLVC are very similar tasks, in the following technical discussion we will mainly focus on MLIC, whose conclusions can be migrated to MLVC naturally.

Recently, single-label image classification has achieved great success thanks to the evolution of deep Convolutional Neural Networks (CNN) (?; ?; ?; ?). Single-label image classification can be naively extended to MLIC tasks by treating the problem as a series of single-label classification tasks. However, such naive extension usually provides poor performance, since the semantic dependencies among multiple labels are ignored, which are especially important for multi-label classification. Therefore, a number of prior works aim to capture the label relations by Recurrent Neural Networks (RNN). However, these methods do not model the explicit relationships between semantic labels and image regions, thus they lack the capacity of sufficient exploitation of the spatial dependency in images.

An alternative solution for MLIC is to introduce object detection techniques. Some methods (?; ?; ?) extract region proposals using extra bounding box annotations, which are much more expensive to label than simple image level annotations. Many other methods (?; ?) apply attention mechanism to automatically focus on the regions of interest. However, the attentional regions are learned only with image-level supervision, which lacks explicit semantic guidance.

To address above issues, we argue that an effective model for multi-label classification should reach two capacities: (1) capturing semantic dependencies among multiple labels in terms of spatial context; (2) locating regions of interest with more semantic guidance.

In this paper, we propose a novel cross-modality attention network associated with graph embedding, so as to simultaneously search for discriminative regions and label spatial semantic dependencies. Firstly, we introduce a novel Adjacency-based Similarity Graph Embedding (ASGE) method which captures the rich semantic relations between labels. Secondly, the learned label embedding will guide the generation of attentional regions in terms of cross-modality guidance, which is referred to as Cross-modality Attention (CMA) in this paper. Compared with traditional self-attention methods, our attention explicitly introduces the rich label semantic relations. Benefiting from the CMA mechanism, our attentional regions are more meaningful and discriminative. Therefore they capture more useful information while suppressing the noise or background information for classification. Furthermore, the spatial context dependencies of labels will be captured, which further improve the performance in MLIC.

The major contributions of this paper are briefly summarized as follows:

We propose an ASGE method to learn semantic label embedding and exploit label correlations explicitly.

We propose a novel attention paradigm, namely cross-modality attention, where the attention maps are generated by leveraging more prior semantic information, resulting in more meaningful attention maps.

A general framework combining CMA and ASGE module, as shown in Fig.1 and Fig.2, is proposed for multi-label classification, which can capture dependencies between spatial and semantic space and discover the location of discriminative feature effectively. We evaluate our framework on MS-COCO dataset and NUS-WIDE dataset for MLIC task, and new state-of-the-art performances are achieved on both of them. We also evaluate our proposed method on YouTube-8M dataset for MLVC, which also achieves remarkable performances.

Related Works

The task of MLIC has attracted an increasing interest recently. The easiest way to address this problem is to treat each category independently, then the task can be directly converted into a series of binary classification tasks (?). However, such techniques are limited by without considering the relationships between labels.

Several approaches have been applied to model the correlations between labels. ? (?) extends the multi-label classification by training the chain binary-classifiers and introducing the correlations between labels by inputting the previously predicted labels. Some other works (?; ?; ?; ?) formulate the task as a structural inference problem based on probabilistic graphical models. Besides, the latest work (?) explores the label dependencies by graph convolutional network. However, none of aforementioned methods consider the associations between semantic labels and image contents, and the spatial contexts of images have not been sufficiently exploited.

In MLIC task, visual concepts are highly related with local image regions. To explore information in local regions better, some works (?; ?) introduce region proposal techniques to focus on informative regions. ? (?) extracts an arbitrary number of object hypotheses, then inputs them into the shared CNN and aggregates the output with max pooling to obtain the ultimate multi-label predictions. ? (?) introduces local information provided by generated proposals to boost the discriminative power of feature extraction. Although above methods have used region proposals to enhance feature representation, they are still limited by requiring extra object-level annotations and without considering the dependencies between objects.

Alternatively, ? (?) discovers the attentional regions corresponding to multiple semantic labels by spatial transformer network and captures the spatial dependencies of the regions by Long Short-Term Memory (LSTM). Analogously, ? (?) proposes the spatial regularization network to generate label-related attention maps and capture the latent relationships by attention maps implicitly. The advantage of above attention approaches is that no additional step of obtaining region proposal is needed. Nevertheless, the attentional regions are learned only with image-level supervision, which lacks of explicit semantic guidance. While in this paper, the semantic guidance is introduced to the generation of attention maps by leveraging label semantic embeddings, which improves the prediction performance significantly.

In this paper, the label semantic embeddings are learned by graph embedding, which is a technique aiming to learn representation of graph-structured data. The approaches of graph embedding mainly contain matrix factorization-based (?), random walk-based (?) and neural network-based methods (?; ?). A main assumption of these approaches is the embeddings of adjacent nodes on the graph are similar, while in our task, we also require embeddings of non-adjacent nodes are mutually exclusive from each other. Therefore, we propose an ASGE method, which can further separate the embeddings of non-adjacent nodes.

MLVC is similar to MLIC, but it involves additional temporal relationships (?; ?; ?; ?; ?; ?). It has been applied to many applications such as emotion recognition (?), human activity understanding (?), and event detection (?). In this paper, we validate our proposed method both in MLIC and MLVC task, and achieve remarkable performance.

Approach

The overall frameworks of our approach for MLIC and MLVC are shown in Fig.1 and Fig.2 respectively. The pipeline includes several stages: Firstly, the label graph is taken as the input of ASGE module to learn label embeddings which encode the semantic relationships between labels. Secondly, the learned label embeddings and visual features will be fed together into the CMA module to obtain category-wise attention maps. Finally, the category-wise attention maps are used to weightedly average the visual features for each category. We will describe our two key components ASGE and CMA in detail.

The relationships between labels play a crucial role in multi-label classification task as discussed in section 1. However, how to express such relationships is an open issue to be solved. Our intuition is that the co-occurrence properties between labels can be described as joint probability, which is suitable for modeling the label relationships. Nevertheless, the joint probability is easy to suffer from the influence of class imbalance. Instead, we utilize the conditional probability between labels to solve this issue, which is obtained by normalizing the joint probability through dividing by marginal probability. Based on this, it is possible to construct a label graph where the labels are nodes and the conditional probability between the labels is edge weight. Inspired by the popular applications of graph embedding method in natural language processing (NLP) tasks, where the learned label embeddings are entered into the network as additional information, we propose a novel ASGE method to encode the label relationships.

We formally define the graph as G=(V,C)\mathcal{G}=(\mathbf{V},\mathbf{C}), where V={v1,v2,...,vN}\mathbf{V}=\{v_{1},v_{2},...,v_{N}\} represents the set of NN nodes and C\mathbf{C} represents the edges. The adjacency matrix A={Aij}i,j=1N\mathbf{A}=\{A_{ij}\}_{i,j=1}^{N} of graph G\mathcal{G} contains non-negative weights associated with each edge. Specifically, V\mathbf{V} is the set of labels and C\mathbf{C} is the set of connections between any two labels, and the adjacency matrix A\mathbf{A} is the conditional probability matrix by setting Aij=P(vi/vj)A_{ij}=P(v_{i}/v_{j}), where PP is calculated through training set. Since P(vi∣vj)≠P(vj∣vi)P(v_{i}|v_{j})\not=P(v_{j}|v_{i}), namely Aij≠AjiA_{ij}\not=A_{ji}, in order to facilitate a better optimization, we symmetrize A\mathbf{A} by

To capture the label correlations defined by the graph structure, we apply a neural network to map the one-hot embedding of each label oi\mathbf{o}_{i} to semantic embedding space and produce the label embedding

where Lge\mathcal{L}_{ge} denotes the loss of our graph embedding.

In order to optimize Eq.3, the cosine similarity cos⁡(ei,ej)\cos(\mathbf{e}_{i},\mathbf{e}_{j}) are required to be close to the corresponding edge weight AijA_{ij} for all i,ji,j. However, it is hard to satisfy this strict constraint, especially when the graph is large and sparse. In order to address this problem, a hyperparameter α\alpha is introduced to the Eq.3 to relax the optimization. The new objective function is as follows:

where σij\sigma_{ij} is an indicator function:

By adding this relaxation, it only needs to make the embedding pairs (ei,ej)(e_{i},e_{j}) be away instead of strictly enforcing cos⁡(ei,ej)\cos(\mathbf{e}_{i},\mathbf{e}_{j}) to be AijA_{ij} when Aij<αA_{ij}<\alpha, thus focusing more on the strong relationships between labels and reducing the difficulty of the optimization.

2 CMA for Multi-label Classification

We formally define the multi-label classification task as a mapping function F:x→yF:\mathbf{x}\rightarrow\mathbf{y}, where x\mathbf{x} denotes a input image or video, y=[y1,y2,...,yN]\mathbf{y}=[y_{1},y_{2},...,y_{N}] denotes corresponding labels, NN is the total number of categories and yn∈{0,1}y_{n}\in\{0,1\} denotes whether the label is assigned to the image or video.

For multi-label classification, we propose an novel attention mechanism, named cross-modality attention, which uses semantic embeddings to guide spatial or temporal integration of visual features. The semantic embeddings here are the label embedding set E={ei}i=0N\mathbf{E}=\{\mathbf{e}_{i}\}_{i=0}^{N} achieved by ASGE and the visual features I=ψ(x)\mathbf{I}=\psi(\mathbf{x}) are extracted by backbone neural networks ψ\psi. Note that for different tasks, we only need to apply different backbones to extract visual features, and the rest part of the framework is completely generic for both tasks.

Cross-Modality Attention.

The learned label embeddings by ASGE compose a semantic embedding space, while the extracted features from CNN Backbone define a visual feature space. Our goal is to let semantic embeddings guide the generation of attention maps. However, semantic embedding space and visual feature space exist a semantic gap because of modality difference. In order to measure the compatibility between different modalities, we first learn a mapping function from the visual feature space to the semantic embedding space, then the compatibility can be measured by a cosine similarity between projected visual feature and semantic embedding, namely cross-modality attention. Formal definition is introduced as follows.

Firstly, we project the visual feature to semantic space by a Cross-Modality Transformer (CMT) module, which is built with several 1×11\times 1 convolution layers followed by a BN and a ReLU activation.

The category-specific attention map zkiz_{k}^{i} is then normalized to:

For each location ii, if the CMA mechanism generates a high positive value, it can be interpreted as the location ii is highly semantic related to label embedding kk or relative more important than other locations, thus the model needs to focus on location ii when considering category kk . Then the category-specific cross-modality attention map is used to weightedly average the visual feature vectors for each category:

where hk\mathbf{h}_{k} is the final feature vector for label kk. Then hk\mathbf{h}_{k} is fed into the fully-connected layers for estimating probability of category kk:

Compared with general single attention map method, where the attention map is shared by all categories, our CMA module benefits in two ways: Firstly, our category-wise attention map is related to image regions corresponding to category kk, thus better learn category-related regions. Secondly, with the guidance of label semantic embeddings, the discovered attentional regions can be better match with the annotated semantic labels.

The potential advantage of our framework is to capture the latent spatial dependency, which is helpful for visual ambiguous labels. As shown in Fig.3, we consider frisbee as an example to explain the spatial dependency. Firstly, The ASGE module learns label embeddings through the label graph, which encodes the label relationships. Since the dog and frisbee are often co-exist while glasses not, therefore the label embeddings of dog and frisbee are close to each other and far away from glasses’s, namely ed≈ef≠eg\mathbf{e}_{d}\approx{\mathbf{e}_{f}}\not=\mathbf{e}_{g}. The optimization procedure during training will enforce the cosine similarity between visual feature and the corresponding label embedding become high, in other words, the cos⁡(ed,vd′),cos⁡(eg,vg′)\cos(\mathbf{e}_{d},\mathbf{v}_{d}^{\prime}),\cos(\mathbf{e}_{g},\mathbf{v}_{g}^{\prime}) and cos⁡(ef,vf′)\cos(\mathbf{e}_{f},\mathbf{v}_{f}^{\prime}) will be large. Since ed≈ef≠eg\mathbf{e}_{d}\approx\mathbf{e}_{f}\not=\mathbf{e}_{g}, cos⁡(ef,vd′)\cos(\mathbf{e}_{f},\mathbf{v}_{d}^{\prime}) will also be large, while cos⁡(ef,vg′)\cos(\mathbf{e}_{f},\mathbf{v}_{g}^{\prime}) will be small. And the final feature representation of frisbee is hf=β1vg+β2vf+β3vd\mathbf{h}_{f}=\beta_{1}\mathbf{v}_{g}+\beta_{2}\mathbf{v}_{f}+\beta_{3}\mathbf{v}_{d},where β1=cos⁡(ef,vg′)\beta_{1}=\cos(\mathbf{e}_{f},\mathbf{v}_{g}^{\prime}), β2=cos⁡(ef,vf′)\beta_{2}=\cos(\mathbf{e}_{f},\mathbf{v}_{f}^{\prime}), β3=cos⁡(ef,vd′)\beta_{3}=\cos(\mathbf{e}_{f},\mathbf{v}_{d}^{\prime}). Thus, the recognition of frisbee is depending on the semantic related label dog and not related to label glasses, indicating that our model is capable of capturing spatial dependencies. Specially, considering that the frisbee is a hard case to be recognized, β2\beta_{2} will be small. Fortunately, the β3\beta_{3} may still be large, so the visual information of dog will be a helpful context to aid in the recognition of label frisbee.

Multi-Scale CMA.

Single-scale feature representation may not be sufficient for multiple objects from different scales. It is noteworthy that the calculation of attention involves the label embedding slides densely over all locations of the feature map, in other words, the spatial resolution of feature map may effect on attention result. Our intuition is that the low-resolution feature maps have more representational capacity for small objects while high-resolution is opposite. The design of CMA mechanism makes it can be naturally applied to multi-scale feature maps via a score fusion strategy. Specially, we extract a set of feature maps {I1,I2,...,IL}\{I_{1},I_{2},...,I_{L}\} and the final predicted probability of multi-scale CMA is

Training Loss.

Finally we define our object function for multi-label classification as follows

Where wkw_{k} is used to alleviate the class imbalance, β\beta is a hyperparameter and pkp_{k} is the ratio of label kk in the training set.

Experiments

To assess our model, we perform experiments on two benchmark multi-label image recognition datasets (MS-COCO (?) and NUS-WIDE (?)) . We also validate the effectiveness of our model on one multi-label video recognition dataset (YouTube-8M Segments) , and the results demonstrate the extensibility of our method. In this section, we will introduce the results on MLIC and MLVC respectively.

Evaluation Metrics.

We use the same evaluation metrics as other works (?), which are the per-category and overall metrics: precision (CP and OP), recall (CR and OR) and F1 (CF1 and OF1). In addition, we also calculate the mean average precision (mAP), which is relatively more important than other metrics, and we mainly focus on the performance of mAP.

Results on MS-COCO Dataset.

The MS-COCO dataset is widely used in MLIC task. It contains 122,218 images with 80 labels and almost 2.9 labels per image. We divide the dataset into two parts: 82,081 images for training and 40,137 images for testing, according to the officially provided division criteria.

We compare with the currently published state-of-the-art methods, including CNN-RNN(?), RNN-Attention(?), Order-Free RNN(?), ML-ZSL (?), SRN (?) and Multi-Evidence (?). Besides, we run the source code released by ML-GCN(?) to train and get the results for comparison. The quantitative results of CMA and MS-CMA model are shown in Table 1. Our two models both perform better than the state-of-the-art methods over almost all metrics. Specially, our MS-CMA model achieves better performance than the CMA model, demonstrating the multi-scale attentions yield performance improvement.

Results on NUS-WIDE Dataset.

The NUS-WIDE is a web dataset including 269,648 images and 5018 labels from the Flickr. After removing the noise and the rare labels, there are 1000 categories left. The images are further manually annotated into 81 concepts with 2.4 concepts per image on average. We follow the split used in (?), i.e. 150,000 images for training and 59,347 for testing after removing the images without any labels.

In this dataset, we compare with the current state-of-the-art models, including CNN-RNN (?) , CNN-SREL-RNN (?) , CNN-LSEP (?), Order-Free RNN(?), ML-ZSL (?), S-CLs(?), Attention transfer(?) and FitsNet(?).

The quantitative results are shown in the Table 2. The comparison results are similar to MS-COCO’s. Our CMA and MS-CMA perform better than state-of-the-art methods on most metrics. The mAP metric of our MS-CMA, which is mostly concerned, exceeds the previous state-of-the-art result by 1.3%1.3\%. We observe that the average edge weight per label of NUS-WIDE is 3.3, while that of MS-COCO is 3.9. This shows that the label graph of MS-COCO is denser than NUS-WIDE. And the performance gain of NUS-WIDE is slightly less obvious than that of MS-COCO. These observations indicate that richer label relationships may bring performance improvement.

Ablation Study.

In this section, we expect to answer the following questions:

Comparing with the backbone (ResNet-101) model, does our CMA model improve significantly?

Does our proposed CMA mechanism have an advantage over the general self-attention methods?

Can the CMA extend to multi-scale and bring performance improvement?

Is our ASGE more advantageous than other embedding methods, e.g. Word2vec?

To answer these questions, we conduct some ablation studies on the MS-COCO dataset, as shown in table 3. Firstly, we investigate how CMA contributes to mAP. It is obvious to see that the vanilla ResNet-101 achieves 79.9%79.9\% mAP, while increases to 83.4%83.4\% when CMA module is added. This result shows the significant effectiveness of the CMA mechanism. Secondly, we implement a general self-attention method by replacing Eq.7 with zi=σ⁡(fconv(Ii))z^{i}=\operatorname{\sigma}(f_{conv}(I^{i})), where fconvf_{conv} denotes the map function of 1×11\times 1 convolution layer. Our CMA mechanism performs better than general self-attention mechanism by achieving 2.3%2.3\% mAP improvement, which indicates that the label semantic embeddings guided attention mechanism is superior to the general self-attention due to introducing much more prior information. Thirdly, expanding our CMA mechanism to multiple scales could obtain about 0.4%0.4\% improvement. This result demonstrates that our attention mechanism is well adapted for multi-scale features. Finally, we compared our ASGE with other embedding methods, in this paper, we take Word2vec as an exmaple, which is a group of related models used to produce word embeddings. Specially, we view the label set in each image as a single sentence, and the window size in Word2vec is set as the length of longest sentence to eliminate the influence of label order. The experiment results show our ASGE based MS-CMA performs better than Word2vec based MS-CMA (represented as W2V-MS-CMA) by 1.3%1.3\% improvement. In our ASGE, label relationships are explicitly represented by adjacency matrix which is treated as a direct optimization target. Instead, Word2vec implicitly encodes label relationships in a data-driven manner without directly optimizing label relationships. Therefore, our ASGE will capture label relationships much better.

Visualization and Analysis.

In this section, we visualize the learned attention maps to illustrate the ability of exploiting discriminative or meaningful regions and capturing the spatial semantic dependencies.

We show the attention visualization examples in Fig.4. The three rows show the category-wise attention maps generated by CMA model and general self-attention respectively. It is observed that the CMA model concentrates more on semantic regions and has stronger response than general self-attention, thus it is capable of exploiting more discriminative and meaningful information. Besides, our CMA mechanism has the ability of capturing the spatial semantic dependencies, especially for the indiscernible or small objects occur in the image, e.g. attention of sports ball also pays attention to tennis racket due to their semantic similarity. It’s quite helpful because these objects need richer contextual cues to help recognition.

2 Multi-label Video Classification

For the training of ASGE, we apply optimization relaxation and set α=0.1\alpha=0.1. Other settings are same as described in MLIC task. For the training of classification, the initial learning rate is 0.0002 and decay each 2×1062\times 10^{6} samples with momentum 0.8 . The hyperparameter β\beta in the Eq.12 is 0. The optimizer is SGD with momentum 0.9. The batch size is 256. In this task, we only implemented a single scale CMA model.

Evaluation Metrics.

In the MLVC task, we use several metrics to evaluate our model, including Global Average Precision (GAP)(?), Average Hit Rate (Avg Hit@1), Precision at Equal Recall Rate (PERR) and Mean Average Precision (mAP)(?).

Results on YouTube-8M Segments Dataset.

In the MLVC task, we verify the effectiveness of our model on the YouTube-8M Segments dataset, which is an extension of the YouTube-8M dataset(?).

In our experiment, we only use frame-level image features, while the state-of-the-art methods use additional audio features and most are built on model ensemble, which is unfair to compare with. For this reason, we compare our CMA model with general self-attention model to validate the effectiveness in MLVC task. Besides, in order to explore the impact of backbone network on CMA mechanism, we implement the SNet-based and FC-based (two fully connected layers) models. The quantitative results are shown in the Table 4. It can be found that all of our metrics are better than the self-attention model. It seems that the improvements compared with self-attention model are not so significant as that of MLIC, but considering the input of our model is fixed pre-extracted features and the models only differ in attention mechanism, the performance gains are quite remarkable.

Visualization and Analysis.

We also present the visualization results of attention scores for a video in Fig.5. The label of the video is skiing, and our CMA model pays more attention to skiing-related frames while partly ignores the redundant frames, suggesting that our attention mechanism is capable of locating attentional frames and demonstrating the effectiveness of our model more intuitively.

Conclusion

In this paper, we propose a novel cross-modality attention mechanism with semantic graph embedding for both MLIC and MLVC task. The proposed method can effectively discover semantic location with rich discriminative features and capture the spatial or temporal dependencies between labels. The extensive evaluations on two MLIC datasets MS-COCO and NUS-WIDE show our method outperforms state-of-the-arts. In addition, we conducted expriments on MLVC datasset YouTube-8M Segments and achieve excellent performance, which validate the strong generalization of our method.

Acknowledgement

This work was supported by the National Natural Science Foundation of China(61671397). We thank all anonymous reviewers for their constructive comments.

References