Label2Label: A Language Modeling Framework for Multi-Attribute Learning

Wanhua Li, Zhexuan Cao, Jianjiang Feng, Jie Zhou, Jiwen Lu

Introduction

Attributes are mid-level semantic properties for objects which are shared across categories . We can describe objects with a wide variety of attributes. For example, human beings easily perceive gender, hairstyle, expression, and so on from a facial image . Multi-attribute learning, which aims to predict the attributes of an object accurately, is essentially a multi-label classification task . As multi-attribute learning involves many important tasks, including facial attribute recognition , pedestrian attribute recognition , and cloth attribute prediction , it plays a central role in a wide range of applications, such as face identification , scene understanding , person retrieval , and fashion search .

For a given sample, many of its attributes are correlated. For example, if we observe that a person has blond hair and heavy makeup, the probability of that person being attractive is high. Another example is that the attributes of beard and woman are almost impossible to appear on a person at the same time. Modeling complex inter-attribute associations is an important challenge for multi-attribute learning. To address this challenge, most existing approaches adopt a multi-task learning framework, which formulates multi-attribute recognition as a multi-label classification task and simultaneously learns multiple binary classifiers. To boost the performance, many methods further incorporate domain-specific prior knowledge. For example, PS-MCNN divides all attributes into four groups and presents highly customized network architectures to learn shared and group-specific representations for face attributes. In addition, some methods attempt to introduce additional domain-specific guidance or annotations . However, these methods struggle to model sample-wise attribute relationships with a simple multi-task learning framework.

Recent years have witnessed great progress in the large-scale pre-training language models . As a representative work, BERT utilizes a masked language model (MLM) to capture the word co-occurrence and language structure. Inspired by these methods, we propose a language modeling framework named Label2Label to model the complex instance-wise attribute relations. Specifically, we regard an attribute label as a “word”, which describes the current state of the sample from a certain point of view. For example, we treat the labels “attractive” and “no eyeglasses” as two “words”, which give us a sketch of the sample from different perspectives. As multiple attribute labels of each sample are used to depict the same object, these “words” can be organized as an unordered yet meaningful “sentence”. For example, we can describe the human face in Fig. 1 with the sentence “attractive, not bald, brown hair, no eyeglasses, not male, wearing lipstick, …”. Although this “sentence” has no grammatical structure, it can convey some contextual semantic information. By treating multiple attribute labels as a “sentence”, we exploit the correlation between attributes with a language modeling framework.

Our proposed Label2Label consists of an attribute query network (AQN) and an image-conditioned masked language model (IC-MLM). The attribute query network first generates the initial attribute predictions. Then these predictions are treated as pseudo label “sentences” and sent to the IC-MLM. Instead of simply adopting the masked language modeling framework, our IC-MLM randomly masks some “word” tokens from the pseudo label “sentence” and predicts the masked “words” conditioned on the masked “sentence” and image features. The proposed image-conditioned masked language model provides partial attribute prompts during the precise mapping from images to attribute categories, thereby facilitating the model to learn complex sample-level attribute correlations. We take facial attribute recognition as an example and show the key differences between our method and existing methods in Fig. 1.

We summarize the contributions of this paper as follows:

We propose Label2Label to model the complex attribute relations from the perspective of language modeling. As far as we know, Label2Label is the first language modeling framework for multi-attribute learning.

Our Label2Label proposes an image-conditioned masked language model to learn complex sample-level attribute correlations, which recovers a “sentence” from the masked one conditioned on image features.

As a simple and generic framework, Label2Label achieves very competitive results across three multi-attribute learning tasks, compared to highly tailored task-specific approaches.

Related Work

Multi-Attribute Recognition: Multi-attribute learning has attracted increasing interest due to its broad applications . It involves many different visual tasks according to the object of interest. Many works focus on domain-specific network architectures. Cao et al. proposed a partially shared multi-task convolutional neural network (PS-MCNN) for face attribute recognition. The PS-MCNN consists of four task-specific networks and one shared network to learn shared and task-specific representations. Zhang et al. proposed Two-Stream Networks for clothing classification and attribute recognition. Since some attributes are located in the local area of the image, many methods resort to the attention mechanism. Guo et al. presented a two-branch network and constrained the consistency between two attention heatmaps. A multi-scale visual attention and aggregation method was introduced in , which extracted visual attention masks with only attribute-level supervision. Tang et al. proposed a flexible attribute localization module to learn attribute-specific regional features. Some other methods further attempt to use additional domain-specific guidance. Semantic segmentation was employed in to guide the attention of the attribute prediction. Liu et al. learned clothing attributes with additional landmark labels. There are also some methods to study multi-attribute recognition with insufficient data, but this is beyond the scope of this paper.

Language Modeling: Pre-training language models is a foundational problem for NLP. ELMo was proposed to learn deep contextualized word representations. It was trained with a bidirectional language model objective, which combined both a forward and backward language model. ELMo representations significantly improve the performance across six NLP tasks. GPT employed a standard language model objective to pre-train a language model on large unlabeled text corpora. The Transformer was used as the model architecture. The pre-trained model was fine-tuned on downstream tasks and achieved excellent results in 9 of 12 tasks. BERT used a masked language model pre-training objective, which enabled BERT to learn bidirectional representations conditioned on the left and right context. BERT employed a multi-layer bidirectional Transformer encoder and advanced the state-of-the-art performance. Our work is inspired by the recent success of these methods and is the first attempt to model multi-attribute learning from the perspective of language modeling.

Transformer for Computer Vision: Transformer was first proposed for sequence modeling in NLP. Recently, Transformer-based methods have been deployed in many computer vision tasks . ViT demonstrated that a pure transformer architecture achieved very competitive results on image classification tasks. DETR formulated the object detection as a set prediction problem and employed a transformer encoder-decoder architecture. Pix2Seq regarded object detection as a language modeling task and obtained competitive results. Zheng et al. replaced the encoder of FCN with a pure transformer for semantic segmentation. Liu et al. utilized the Transformer decoder architecture for multi-label classification. Temporal query networks were introduced in for fine-grained video understanding with a query-response mechanism. There are also some efforts to apply Transformer to the task of multi-label image classification. Note that the main contribution of this paper is not the use of Transformer, but modeling multi-attribute recognition from the perspective of language modeling.

Approach

In this section, we first give an overview of our framework. Then we present the details of the proposed attribute query network and image-conditioned masked language model. Lastly, we introduce the training objective function and inference process of our method.

Given a sample x\bm{x} from a dataset D\mathcal{D} with MM attribute types, we aim to predict the multiple attributes y\bm{y} to the image x\bm{x}. We let A={a1,a2,...,aM}\mathcal{A}=\{\bm{a}_{1},\bm{a}_{2},...,\bm{a}_{M}\} denote the attribute set, where aj(1≤j≤M)\bm{a}_{j}(1\leq j\leq M) represents the jj-th attribute type. For simplicity, we assume that the values of all attribute types are binary. In other words, the value of aj\bm{a}_{j} is or 11, where 11 means that the sample has this attribute and means not. However, our method can be easily extended to the case where each attribute type is multi-valued. With this assumption, we have y∈{0,1}M\bm{y}\in\{0,1\}^{M}. Existing methods usually employ a multi-tasking learning framework, which uses MM binary classifiers to predict MM attributes respectively. Binary cross-entropy loss is used as the objective.

This paper proposes a language modeling framework. We show the pipeline of our framework in Fig. 2. The key idea of this paper is to treat attribute labels as unordered “sentences” and use an image-conditioned masked language model to exploit the relationships between attributes. Although we can directly use the real attribute labels as the input of the IC-MLM during training, we cannot access these labels for inference. To address this issue, our Label2Label introduces an attribute query network to generate the initial attribute predictions. These predictions are then treated as pseudo-labels and used as input to the IC-MLM in the training and testing phases.

2 Attribute Query Network

Our attribute query network learns a set of permutation-invariant query vectors Q={q1,q2,...,qM}\bm{Q}=\{\bm{q}_{1},\bm{q}_{2},...,\bm{q}_{M}\}, where each query qj\bm{q}_{j} corresponds to an attribute type aj\bm{a}_{j}. Then each query vector qj\bm{q}_{j} pools the attribute-related features from the image features with Transformer decoder layers and generates the corresponding response vector rj\bm{r}_{j}. Finally, we learn a binary classifier for each response vector to generate the initial attribute predictions.

It is worth noting that the predictions from the attribute query network are not 100%100\% correct, resulting in some wrong “words” in the generated label “sentence”. However, we can treat the wrong “words” as another form of masks, because the wrong predictions account for only a small proportion. In fact, the masking strategy of the wrong word is artificially performed in some language models, such as BERT .

3 Image-Conditioned Masked Language Model

In existing multi-attribute databases, images are annotated with a variety of attribute labels. This paper is dedicated to modeling sample-wise complex attribute correlations. Instead of treating attribute labels as numbers, we regard them as “words”. Since different attribute labels describe the object in an image from different perspectives, we can group them as a sequence of “words”. Although the sequence is essentially an unordered “sentence” without any grammatical structure, it still conveys meaningful contextual information. In this way, we treat y\bm{y} as an unordered yet meaningful “sentence”, where yjy_{j} is a “word”.

By treating the labels as sentences, we resort to language modeling methods to mine the instance-level attribute relations effectively. In recent years, pre-training large-scale task-agnostic language models have substantially advanced the development of NLP, among which representative works include ELMo , GPT-3 , BERT , and so on. Inspired by the success of these methods, we consider a masked language model to learn the relationship between “words”. We mask some percentage of the attribute label “sentence” y\bm{y} at random, and then reconstruct the entire label “sentence”. Specifically, for a binary label sequence, we replace those masked “words” with a special work token [mask] to obtain the masked sentence. Then we input the masked sentence to a masked language model, which aims to recover the entire label sequence. While the MLM has proven to be an effective tool in NLP, directly using it for multi-attribute learning is not feasible. Therefore, we propose several important improvements.

Instance-wise Attribute Relations: MLM essentially constructs the task P(y1,y2,...,yM∣M(y1),M(y2),...,M(yM))P(y_{1},y_{2},...,y_{M}|\mathcal{M}(y_{1}),\mathcal{M}(y_{2}),...,\mathcal{M}(y_{M})) to capture the “word” co-occurrence and learn the joint probability of “word” sequences P(y1,y2,...,yM)P(y_{1},y_{2},...,y_{M}), where M()\mathcal{M}() denotes the random masking operation. Such a naive approach leads to two problems. The first problem is that MLM only captures statistical attribute correlations. A diverse dataset means that the mapping {M(y1),M(y2),...,M(yM)}↦{y1,y2,...,yM}\{\mathcal{M}(y_{1}),\mathcal{M}(y_{2}),...,\mathcal{M}(y_{M})\}\mapsto\{y_{1},y_{2},...,y_{M}\} is a one-to-many mapping. Therefore MLM only learns how different attributes are statistically related to each other. Meanwhile, our experiments find that this prior can be easily modeled by the attribute query network P(y1,y2,...,yM∣x)P(y_{1},y_{2},...,y_{M}|\bm{x}). The second problem is that MLM and attribute query network cannot be jointly trained. Since MLM uses only the hard prediction of the attribute query network, the gradient from MLM cannot influence the training of the attribute query network. In this way, the method becomes a two-stage label refinement process, which significantly reduces the optimization efficiency.

To address these issues, we propose an image-conditioned masked language model to learn instance-wise attribute relations. Our IC-MLM captures the relations by constructing a task P(y1,y2,...,yM∣x,M(y1),M(y2),...,M(yM))P(y_{1},y_{2},...,y_{M}|\bm{x},\mathcal{M}(y_{1}),\mathcal{M}(y_{2}),...,\mathcal{M}(y_{M})). Introducing an extra image condition is not trivial, as this fundamentally changes the behavior of MLM. With the conditions of image x\bm{x}, the transformation {x,M(y1),M(y2),...,M(yM)}↦{y1,y2,...,yM}\{\bm{x},\mathcal{M}(y_{1}),\mathcal{M}(y_{2}),...,\mathcal{M}(y_{M})\}\mapsto\{y_{1},y_{2},...,y_{M}\} is an accurate one-to-one mapping. Our IC-MLM infers other attribute values by combining some attribute label prompts and image contexts in the precise image-to-label mapping, which facilitates the model to learn sample-level attribute relations. In addition, IC-MLM and the attribute query network can use shared image features, which enables them to be jointly optimized with a one-stage framework.

Word Embeddings: It is known that the word id is not a good word representation in NLP. Therefore, we need to map the word id to a token embedding. Instead of utilizing existing word embeddings with a large token vocabulary like BERT , we directly learn attribute-related word embeddings E\bm{E} from scratch. We use the word embedding module to map the “word” in the masked sentence to the corresponding token embedding. Since all attributes are binary, we need to build a token vocabulary with a size of 2M2M to model all possible attribute words. Also, we need to include the token embedding for the special word [mask]. This paper considers three different strategies for the [mask] token embedding. The first strategy believes the [mask] words for different attributes have different meanings, so MM attribute-specific learnable token embeddings are learned, where one [mask] token embedding corresponds to one attribute. The second strategy treats the [mask] words for different attributes as the same word. Only one attribute-agnostic learnable token embedding is learned and shared by all attributes. The third strategy is based on the second strategy, which simply replaces the learnable token embedding with a fixed 0\bm{0} vector. Our experiments find all three strategies work well while the first strategy performs best.

Positional Embeddings: In BERT, the positional embedding of each word is added to its corresponding token embeddings to obtain the position information. Since our “sentences” are unordered, there is no need to introduce positional embeddings to “word” representations. We conducted experiments with positional embeddings by randomly defining some word order and found no improvement. Therefore we do not use positional embeddings for “word” representations and the learned model is permutation invariant for “words”.

Architecture: In NLP, Transformer encoder layers are usually used to implement MLM, while we use multi-layer Transformer decoders to implement IC-MLM due to additional image input conditions. Following the design philosophy similar to the attribute query network, token embeddings E\bm{E} pool features from the local visual features X′\bm{X}^{\prime} with a cross-attention mechanism. We update the token features Ei−1\bm{E}_{i-1} in the ii-th Transformer decoder layer as follows:

We set E\bm{E} to E0\bm{E}_{0} and the number of Transformer decoder layers in IC-MLM to DD. Then we denote ED\bm{E}_{D} as R′={r1′,r2′,...,rM′}\bm{R}^{\prime}=\{\bm{r}^{\prime}_{1},\bm{r}^{\prime}_{2},...,\bm{r}^{\prime}_{M}\}, where rj′\bm{r}^{\prime}_{j} corresponds to the updated feature of token Ej\bm{E}_{j}. In the end, we perform the final multi-attribute classification with linear projection layers. Formally, we have:

4 Objective and Inference

As commonly used in most existing methods , we adopt the binary cross-entropy loss to train the IC-MLM. On the other hand, since most of the datasets for multi-attribute recognition are highly imbalanced, different tasks usually use different weighting strategies. The loss function for the IC-MLM is formulated as Lmlm(x) ⁣ ⁣= ⁣ ⁣∑j=1M ⁣wj(yj ⁣log⁡(pj) ⁣+ ⁣(1 ⁣− ⁣yj) ⁣log⁡(1 ⁣− ⁣pj))\mathcal{L}_{mlm}(\bm{x})\!\!=\!\!\sum_{j=1}^{M}\!w_{j}(y_{j}\!\log(p_{j})\!+\!(1\!-\!y_{j})\!\log(1\!-\!p_{j})), where wjw_{j} is the weighting coefficient. According to different tasks, we choose different weighting strategies and always follow the most commonly used strategy for a fair comparison. Meanwhile, to ensure the quality of the generated pseudo label sequences, we also supervise the attribute query network with the same loss function Laqn(x) ⁣ ⁣= ⁣ ⁣∑j=1M ⁣wj(yj ⁣log⁡(lj) ⁣+ ⁣(1 ⁣− ⁣yj) ⁣log⁡(1 ⁣− ⁣lj))\mathcal{L}_{aqn}(\bm{x})\!\!=\!\!\sum_{j=1}^{M}\!w_{j}(y_{j}\!\log(l_{j})\!+\!(1\!-\!y_{j})\!\log(1\!-\!l_{j})). The final loss function Ltotal\mathcal{L}_{total} is a combination of the two loss functions above:

where λ\lambda is used to balance these two losses. At inference time, we ignore the masking step and directly input the pseudo label “sentence” to the IC-MLM. Then the output of the IC-MLM is used as the final attribute prediction.

Experiments

In this section, we conducted extensive experiments on three multi-attribute learning tasks to validate the effectiveness of the proposed framework.

Dataset: LFWA is a popular unconstrained facial attribute dataset, which consists of 13,143 facial images of 5,749 identities. Each facial image has 40 attribute annotations. Following the same evaluation protocol in , we partition the LFWA dataset into two sets, with 6,263 images for training and 6,880 for testing. All images are pre-cropped to a size of 250×250250\times 250. We adopt the classification error for evaluation following .

Experimental Settings: We trained our model for 57 epochs with a batch size of 16. For optimization, we used an SGD optimizer with a base learning rate of 0.01 and cosine learning rate decay. The weight decay was set to 0.001. To augment the dataset, Rand-Augment and Random horizontal flipping were performed. We also adopted Mixup for regularization.

Parameters Analysis: We first analyze the influence of the number of Transformer decoder layers in the attribute query network and IC-MLM. The results are shown in Tables 2 and 2. We see that the best performance is achieved when L=1L=1 and D=2D=2. We further conduct experiments with different mask ratios α\alpha and list the results in Table 4. As we mentioned above, the wrong “words” in the pseudo label sequences also provide some form of masks. Therefore, our method performs well when α=0\alpha=0. We observe that our method attains the best performance when α=0.1\alpha=0.1. Table 4 shows the results with different λ\lambda, and we see that λ=1\lambda=1 gives the best trade-off in (4). We consider three different strategies for [MASK] token embedding and list the results in Table 7. We see that the attribute-specific strategy achieves the best performance among them, as it better models the differences between the attributes. Unless explicitly mentioned, we adopt these optimal parameters in all subsequent experiments.

Ablation Study: To validate the effectiveness of our Label2Label, we also conduct experiments on the LFWA dataset with two baseline methods. We first consider the Attribute Query Network (AQN) method, which ignores the IC-MLM and treats the outputs of AQN in Fig. 2 as final predictions. FC Head method further replaces the Transformer decoder layers in AQN with a linear classification layer. To further verify the generalization of our method, we use different feature extraction backbone networks for ablation experiments. To better demonstrate the significance of the results, we also report the standard deviation. The results are presented in Table 5. In addition, we report the computation cost (MACs) of each method in Table 5. We observe that our method significantly outperforms FC Head and AQN across various backbones with marginal computational overhead, which illustrates the effectiveness of our method.

We then conducted experiments to show how image-conditioned MLM improves performance. The results are listed in Table 7. As we analyzed above, MLM leads to a two-stage label refinement process. We consider two network architectures to implement MLM: Transformer encoder and multilayer perceptron (MLP). The results show that none of them improve the performance of AQN (13.36%). The reason is that MLM only learns statistical attribute relations, and this prior is easily captured by AQN. Meanwhile, our IC-MLM learns instance-wise attribute relations. To see the benefits of the additional image conditions, we still adopt the two-stage label refinement process, and train Transformer decoder layers with fixed image features. We see that performance is boosted to 13.01%, which demonstrates the effectiveness of modeling instance-wise attribute relations. We further jointly train the IC-MLM and attribute query network, which achieves significant performance improvement. These results illustrate the superiority of the proposed IC-MLM.

Comparison with State-of-the-art Methods: Following , we employ ResNet50 as the backbone. We present the performance comparison on the LFWA dataset in Table 8. We observe that our method attains the best performance with a simple framework compared to highly tailored domain-specific methods. Label2Label even exceeds the methods of using additional annotations, which further illustrates the effectiveness of our framework.

Visualization: As the Transformer decoder architecture is used to model the instance-level relations, our method can give better interpretable predictions. We visualize the attention scores in the IC-MLM with DODRIO . As shown in Fig. 3, we see that related attributes tend to have higher attention scores.

2 Pedestrian Attribute Prediction

Dataset: The PA-100K dataset is the largest pedestrian attribute dataset so far . It contains 100,000 pedestrian images from 598 scenes, which are collected from real outdoor surveillance videos. All pedestrians in each image are annotated with 26 attributes including gender, handbag, and upper clothing. The dataset is randomly split into three subsets: 80% for training, 10% for validation, and 10% for testing. Following SSC , we merge the training set and the validation set for model training. We use five metrics: one label-based and four instance-based. For the label-based metric, we adopt the mean accuracy (mA) metric. For instance-based metrics, we employ accuracy, precision, recall, and F1 score. As mentioned in , mA and F1 score are more appropriate and convincing criteria for class-imbalanced pedestrian attribute datasets.

Experimental Settings: Following the state-of-the-art methods , we adopted ResNet50 as the backbone network to extract image features. We first resize all images into 256×\times192 pixels. Then random flipping and random cropping were used for data augmentation. SGD optimizer was utilized with the weight decay of 0.0005. We set the initial learning rate of the backbone to 0.01. For fast convergence, we set the initial learning rate of the attribute query network and IC-MLM to 0.1. The batch size was equal to 64. We trained our model for 25 epochs using a plateau learning rate scheduler. We reduced the learning rate by a factor of 10 once learning stagnates and the patience was 4.

Results and Analysis: We report the results in Table 9. We observe that Label2Label achieves the best performance in mA, Accuracy, and F1 score. Compared to the previous state-of-the-art method SSC , which designs complex SPAC and SEMC modules to extract discriminative semantic features, our method achieves 0.37% performance improvements in mA. In addition, we report the re-implemented results of the MsVAA, VAC, and ALM methods in the same setting as did in . Our method consistently outperforms these methods. We further show the results of the FC Head and Attribute Query Network. We see that the performance is improved by replacing the FC head with Transformer decoder layers, which shows the superiority of our attribute query network. Our Label2Label outperforms the attribute query network method by 1.35% for mA, which shows the effectiveness of the language modeling framework.

3 Clothing Attribute Recognition

Dataset: Clothing Attributes Dataset consists of 1,856 images that contain clothed people. Each image is annotated with 26 clothing attributes, such as colors and patterns. We use 1,500 images for training and the rest for testing. For a fair comparison, we only use 23 binary attributes and ignore the remaining three multi-class value attributes as in . We adopt accuracy as the metric and also report the accuracy of four clothing attribute groups following .

Experimental Settings: For a fair comparison, we utilized AlexNet to extract image features following . We trained our model for 22 epochs using a cosine decay learning rate scheduler. We utilized an SGD optimizer with an initial learning rate of 0.05. The batch size was set to 32. For the attribute query network, we employed a 2-layer Transformer decoder (L=2L=2).

Results and Analysis: Table 10 shows the results. We observe that our Label2Label attains a total accuracy of 92.87%, which outperforms other methods with a simple framework. MG-CNN learns one CNN for each attribute, resulting in more training parameters and longer training time. Compared with the attribute query network method, our method achieves better performance on all attribute groups, which illustrates the superiority of our framework.

Conclusions

In this paper, we have presented Label2Label, which is a simple and generic framework for multi-attribute learning. Different from the existing multi-task learning framework, we proposed a language modeling framework, which regards each attribute label as a “word”. Our model learns instance-level attribute relations by the proposed image-conditioned masked language model, which randomly masks some “words” and restores them based on the remaining “sentence” and image context. Compared to well-optimized domain-specific methods, Label2Label attains competitive results on three multi-attribute learning tasks.

Acknowledgments. This work was supported in part by the National Key Research and Development Program of China under Grant 2017YFA0700802, in part by the National Natural Science Foundation of China under Grant 62125603 and Grant U1813218, in part by a grant from the Beijing Academy of Artificial Intelligence (BAAI). The authors would sincerely thank Yongming Rao and Zhiheng Li for their generous helps.

Appendix 0.A Evaluation Metrics

For pedestrian attribute prediction, we adopted five evaluation metrics. We present the details of these metrics. The only label-based metric is the mean accuracy (mA) metric, which is the mean of positive accuracy and negative accuracy for each attribute. Mathematically, the mA is calculated by:

where MM is the number of attributes, PjP_{j} and TPjTP_{j} represent the numbers of positive samples and correctly predicted positive samples of the jj-th attribute respectively, NjN_{j} and TNjTN_{j} are the numbers of negative samples and correctly predicted negative samples of the jj-th attribute respectively.

We also consider four example-based metrics: accuracy, precision, recall, and F1 score:

where NN denotes the number of samples, Yi\bm{Y}_{i} is the positive labels of the ii-th sample and Yi′\bm{Y}^{\prime}_{i} is the predicted positive values for the ii-th sample.

Appendix 0.B Weighting Strategy

For facial attribute recognition and clothing attribute recognition, we follow the common practice which does not utilize the weighting strategy for loss functions. Therefore, we have:

For pedestrian attribute recognition, we follow the widely used weighted binary-entropy strategy in . In this way, we have:

where γj\gamma_{j} is the positive example ratio of the jj-th attribute.

Appendix 0.C More Ablation Studies

We conducted ablation experiments on the position embeddings of word representations. Since we are dealing with unordered “sentences”, we randomly define three different label sequences and use the corresponding position embeddings respectively. We report the average performance of three different label sequences on the LFWA database in Table 12. We found no additional performance gain from the position embeddings of the word representations. The reason is that our “sentences” are essentially made up of unordered “words”.

C.2 Position Embeddings for Visual Features

In our paper, we add 2D-aware position embeddings to visual feature vectors to retain positional information. We conduct experiments to verify their effectiveness and show the results on the LFWA database in Table 12. We observe that introducing position embeddings in visual features is beneficial for performance.

C.3 Comparisons with Transformer-based Multi-label Classification Methods

Many Transformer-based multi-label classification methods have been proposed in recent years. To further verify the effectiveness of the proposed method, we conducted experiments on the three datasets used in our paper. Table 13 shows the results. We see our method consistently outperforms C-Tran and Q2L , which shows the superiority of our method.

C.4 The Need of Masking

To verify the effectiveness of masking, we construct three pure reconstruction (without masking) baselines. 1) Feature Reconstruction: direct reconstruction of the word features r1,r2,...,rM\bm{r}_{1},\bm{r}_{2},...,\bm{r}_{M}. 2) Score Reconstruction: direct reconstruction of the predicted scores l1,l2,...,lMl_{1},l_{2},...,l_{M}. 3) Label Reconstruction: direct reconstruction of the labels: y1,y2,...yMy_{1},y_{2},...y_{M}. Table 14 shows the results on the LFWA dataset. Although the Label Reconstruction works competitively, it is still inferior to our method with masking. Just as found in , although Autoencoder (reconstruction) works well, the Masked Autoencoder (masking) is the key factor to learning better features. In BERT, the masked word is replaced with the [mask] token or a random word. So the MLM has two tasks: mask-recovering and error-correcting. Both increase the training difficulty. In our method, the wrong predictions are like random words in BERT. See Table 14, Ours (α=0\alpha=0) outperforms Label Reconstruction (12.55 vs 12.70). The only difference is that the input of our IC-MLM contains wrong predictions while Label Reconstruction does not, which proves that our proposed IC-MLM also benefits from handling this special “mask”.

Appendix 0.D Network Structure Configuration

We show the default hyper-parameters for the Transformer decoder layer of our method in Table 15.

Appendix 0.E More Visualization Results

We provide more visualization results of the attention scores in Figure 4. We conducted the experiments on the LFWA database. We read out the attention from the self-attention layer of our label decoder. The DODRIO is used for visualization. We show the attention scores of the first head at layer 1 with four examples.

For the first example, the attribute “Wearing Earrings” is strongly related to the existence of “Wearing Lipstick”, “No Beard”, and “Female” and the absence of “5 o’clock shadow”. For the second sample, the attributes “Oval Face”, “Pointy Nose”, “Sideburns”, “Wearing Necktie” and “Male” imply the existence of “Attractive”. For the third sample, the attributes “Wearing Earrings” and “Wearing Lipstick” indicate the gender “Female”. For the last example, the attribute “Wearing Lipstick” assigns more attention to the existence of “Wearing Earrings”, “Wearing Necklace”, “Heavy Makeup” and the absence of “Mustache”, “Male”. We see our method can learn the instance-level attribute relations even if a sample has some wrong labels.

Appendix 0.F Detailed Results

For facial attribute recognition, some methods report the pre-class recognition accuracy. We report the pre-attribute classification error on the LFWA database in Table 16 for a comprehensive comparison.

We observe our method attains very competitive results with a simple framework compared to highly tailored domain-specific methods, which demonstrates the effectiveness of our method.

References