Spatial and Semantic Consistency Regularizations for Pedestrian Attribute Recognition

Jian Jia, Xiaotang Chen, Kaiqi Huang

Introduction

Pedestrian attribute recognition aims to predict multiple human attributes, such as age, gender, and clothing, as semantic descriptions for a pedestrian image. Due to the ubiquitous application in surveillance scenarios , scene understanding , and human perception , numerous methods have been proposed and significant progress has been made in the last decade.

Existing methods mainly utilize the complicated network, such as Feature Pyramid Network (FPN), to enrich attribute representation from multi-level feature maps, and combine the attention mechanisms to precisely locate attribute-related regions. Recently, VAC utilizes a human prior, that attention regions of random augmentations of the same image are consistent, to improve model robustness. The above methods mainly emphasize learning discriminative attribute features from an individual image, instead of exploiting the relation between different pedestrian images of the same attribute. In contrast, our methods show that mining the inter-image relations between different images of the same attribute can significantly help the model locate attribute-related regions and extract inherent semantic features. We exploit inter-image relations from the perspective of spatial relation and semantic relation.

For the inter-image spatial relation, we hypothesize that the spatial location of the same attribute is basically consistent between different pedestrian images, which is called SPAtial Consistency (SPAC) in this work. For example, the “hat” attribute and the “boots” attribute mostly appears at the top and bottom of the picture, respectively, which is shown in the first row of Figure 1(a). However, we observe that Class Activation Maps (CAMs) of the same attribute of the baseline method have significant location variations. Some examples are shown in the second row of Figure 1(a).

These CAMs of the same attribute between different pedestrians are inconsistent, some of which (with red boundary) deviate seriously from attribute-related areas, no matter for the “short sleeve”, “boots”, or “hat” attributes. This phenomenon contradicts our spatial consistency hypothesis, and indicates that the baseline model easily inclines to focus on the background, irrelevant foreground, or a small part of attribute-related regions, which is called the “spatial attention deviation problem” in this work.

For the inter-image semantic relation, inherent semantic features of the same attribute between different images should be consistent, which is called SEMantic Consistency (SEMC) in this work. For example, as illustrated in Figure 1(b), regardless of the difference in shape, size, and color between various samples, the intrinsic semantic features of the “hat” attribute should remain basically unchanged. This property is also indispensable for learning discriminative features and obtaining a robust model.

To achieve the spatial and semantic consistency between pedestrian images of the same attribute, we propose a novel framework composed of the SPAC and SEMC module. Specifically, the SPAC module generates reliable spatial locations for each attribute and maintains a stable spatial memory to suppressing location shift, which is caused by overfitting or label noise. Based on precise spatial locations, the SEMC module extracts intrinsic semantic features and maintains a stable semantic memory to suppress the influence of irrelevant characteristics, such as shape, color, and size for the “hat” attribute.

We make the following three contributions in this work:

We establish an effective consistency framework for pedestrian attribute recognition, which makes full use of inter-image spatial and semantic relations between images of the same attribute.

We design spatial and semantic consistency modules to generate precise spatial attention regions and extract discriminative semantic features for each attribute.

We confirm the efficacy of the proposed method by achieving state-of-the-art performance on three popular datasets including PA100K, PETA, and RAP.

Related Work

Pedestrian attribute recognition has witnessed a fast-growing development recently. Li et al. first formulated pedestrian attribute recognition as a multi-label classification task and proposed the weighted sigmoid cross-entropy loss to alleviate the serious imbalance between positive samples and negative samples. To explore attribute context and correlation, the JRL network adopted Long-Shot-Term-Memory to take the pedestrian attribute recognition task as a sequence prediction problem.

Attention mechanism has been widely used in pedestrian attribute recognition to locate attribute-related regions and learn discriminative feature representations. HydraPlus-Net with multi-directional attention modules was introduced to extract pixel-level features and semantic-level features, which were beneficial to locate fine-grained attributes. Based on CAM and EdgeBox , Liu et al. proposed a Localization Guided Network to extract attribute-related local features. PGDM framework utilized a pre-trained human hose estimator and Spatial Transformer Networks (STNs) to generate reliable attribute-related regions.

Considering the discrimination of multi-scale feature maps and the effectiveness of deep supervisions, WPAL , MsVAA , and ALM networks are proposed. Yu et al. proposed the WPAL network, which introduced a weakly-supervised object detection technique into pedestrian attribute recognition. Sarafianos et al. integrated attention mechanism into multi-scale feature maps and adopted a variant of focal loss to solve the imbalance between positive and negative samples of the attribute. ALM module , which was composed of a Squeeze-and-Excitation (SE) block and a STN , was applied to each layer of the Feature Pyramid Network (FPN) to enhance attribute localization. Considering the visual attention regions were consistent between multiple augmentations of the same image, Guo et al. proposed an attention consistency loss to get robust attribute locations. In addition, a hierarchical feature embedding (HFE) framework was proposed to learn fine-grained feature embeddings by combining attribute and ID information. Different from previous methods, person ID information was utilized in the HFE framework, which was not provided on the pedestrian attribute recognition task.

Previous methods mainly concentrated on generating precise attribute-related regions and learning to classify attributes from a single image individually. They neither considered the prior spatial structure knowledge of pedestrian attribute, nor exploited the inter-image relation between different pedestrian images of the same attribute. Whereas, both aspects are considered in our proposed method and introduced in Section 3.2 and 3.3.

From the perspective of using inter-image information, the most related method is the JRL network . Based on global feature similarities, the JRL network utilized inter-image information by aggregating several similar pedestrian features to get the final prediction. Different from JRL, our method utilizes spatial and semantic local features of each attribute, and exploits the inter-image relation to construct consistency regularizations as supervision signals of the training process. From the perspective of consistency constraints, the most related method is the VAC model , which aimed to make the global attention regions consistent between random augmentations of the same image. However, our method aligns the local attention regions between different pedestrian images of the same attribute. In addition, we also introduce the semantic consistency module to extract discriminative attribute features.

Methods

In this section, we first introduce the baseline method. Then, we present the proposed consistency framework, which consists of a classification branch and a consistency branch. The classification branch is completely the same as the baseline network. The consistency branch is divided into spatial consistency module and semantic consistency module, which are introduced separately. The overview of the proposed framework is illustrated in Figure 2(a), and intuitive elaborations of two consistency modules are shown in Figure 2(b). Compared with the baseline method, the proposed method does not introduce extra learnable parameters.

Following , we formulate pedestrian attribute recognition as a multi-label classification task, and multiple binary classifiers with sigmoid functions are adopted. Binary cross-entropy loss is used as the optimization target:

where pi,j=σ(zi,j)p_{i,j}=\sigma(z_{i,j}) is the prediction probability of the classifier output logits zi,jz_{i,j}, and σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}) is the sigmoid function.

2 Spatial Consistency Module

In this section, we propose the SPAtial Consistency (SPAC) module combined with spatial consistency regularization to tackle the spatial attention region deviation problem.

where Wm,c\bm{W}_{m,c} denotes the cc-th element of mm-th classifier weight, and Fi,c(x,y)\bm{F}_{i,c}(x,y) indicates the spatial location (x,y)(x,y) of the cc-th channel in the feature map Fi\bm{F}_{i}. After getting the spatial attention regions of each attribute for every image in a random batch btb_{t}, we adopt the selector — an indicator function takes logits zi,mz_{i,m} and ground truth label yi,my_{i,m} as inputs — to aggregate the attention maps of qualified positive samples We use “positive samples” to represent images that contain target attribute, and “negative samples” to represent images that do not contain target attribute. of the mm-th attribute by:

where Mˉmspa=Mmspa/ ∥Mmspa∥2\bm{\bar{M}}^{spa}_{m}=\bm{M}^{spa}_{m}/\ \|\bm{M}^{spa}_{m}\|_{2}, Aˉmq=Amq/ ∥Amq∥2\bm{\bar{A}}^{q}_{m}=\bm{A}^{q}_{m}/\ \|\bm{A}^{q}_{m}\|_{2}, and α∈(0,1]\alpha\in(0,1] is a momentum coefficient. The effect of momentum coefficient α\alpha is demonstrated in Figure 4.

As shown in Figure 2(b), due to overfitting and label noise, spatial attention regions of the “hat” attribute deviate from attribute-related regions severely. Model inclines to focus on the background, irrelevant foreground, and a small part of the attribute-related areas. Thus, spatial memory Mspa\bm{M}^{spa}, which retains reliable and stable spatial location regions of each attribute, can be taken as the supervision of attribute-related regions to correct the spatial attention deviation. Therefore, based on SPAC module, we propose a spatial consistency regularization LspacL_{spac} by calculating the l1l_{1}-distance between the spatial memory Mmspa\bm{M}^{spa}_{m} and the spatial attention map Amp\bm{A}^{p}_{m} :

where Aˉmp=Amp/ ∥Amp∥2\bm{\bar{A}}^{p}_{m}=\bm{A}^{p}_{m}/\ \|\bm{A}^{p}_{m}\|_{2}, nmp=∑i=1bt\mathds1{ yi,m=1}n_{m}^{p}=\sum_{i=1}^{bt}\mathds{1}_{\{\ y_{i,m}=1\}}, and btb_{t} indicates the batch size. To take all positive samples into consideration, Amp\bm{A}^{p}_{m} is formulated by averaging spatial attention regions of all positive samples of the mm-th attribute in a random batch. Please note the difference between indicator functions of Amq\bm{A}^{q}_{m} and Amp\bm{A}^{p}_{m}.

Overall, to fully utilize the inter-image spatial relation and address the spatial attention region deviation problem, the SPAC module is proposed to extract reliable attribute attention regions Amq\bm{A}^{q}_{m} to update spatial memory and adopt the l1l_{1}-distance to align spatial attention regions Amp\bm{A}^{p}_{m} with spatial memory Mmspa\bm{M}^{spa}_{m}. Considering the soft weights used in Amp\bm{A}^{p}_{m} and Mmspa\bm{M}^{spa}_{m}, we name this method as SSCsoftSSC_{soft}.

3 Semantic Consistency Module

Although the SPAC module considers the inter-image spatial relation that attention regions of different images of the same attribute are consistent, inter-image semantic relation has not been utilized, i.e., intrinsic semantic features of the same attribute are consistent between different images. For example, whether the sample is a beret, helmet, bucket hat, or baseball cap, intrinsic semantic features of the “hat” attribute should be consistent. Thus, based on the SPAC module, we propose SEMantic Consistency (SEMC) module to extract intrinsic and discriminative semantic features for each attribute.

where Vˉmq=Vmq/ ∥Vˉmq∥2\bm{\bar{V}}^{q}_{m}=\bm{V}^{q}_{m}/\ \|\bm{\bar{V}}^{q}_{m}\|_{2} , nmq=∑i=1bt\mathds1{σ(zi,m)>τ, yi,m=1}n_{m}^{q}=\sum_{i=1}^{b_{t}}\mathds{1}_{\{\sigma(z_{i,m})>\tau,\ y_{i,m}=1\}}, and α\alpha is momentum coefficient as same as that of the SPAC module in Equation 4 .

Finally, we design a semantic consistency regularization by computing the l1l_{1}-distance between the semantic memory Mmsem\bm{M}^{sem}_{m} and attribute semantic feature Vmp\bm{V}^{p}_{m} of all positive samples, which is defined as:

where Vˉmp=Vmp/ ∥Vmp∥2\bm{\bar{V}}^{p}_{m}=\bm{V}^{p}_{m}/\ \|\bm{V}^{p}_{m}\|_{2}, Mˉmsem=Mmsem/ ∥Mmsem∥2\bm{\bar{M}}^{sem}_{m}=\bm{M}^{sem}_{m}/\ \|\bm{M}^{sem}_{m}\|_{2}, nmq=∑i=1bt\mathds1{yi,m=1}n_{m}^{q}=\sum_{i=1}^{b_{t}}\mathds{1}_{\{y_{i,m}=1\}}, and the semantic consistency regularization is imposed on semantic features of all positive samples of the mm-th attribute.

By bridging the gap between semantic features of different samples of the same attribute, the SEMC module can extract intrinsic and discriminative semantic features for each attribute and eliminate the interference of attribute-irrelevant characteristics (such as shape, size, and color in the “hat” attribute).

4 Loss Function

As commonly adopted in most existing methods , the weighted binary cross-entropy loss is also utilized in the classification branch of the proposed method as classification loss, which is formulated as :

where rjr_{j} is the positive sample ratio of jj-thth attribute in the training set.

The final loss function LL is a weighted summation of the classification loss, SPAC regularization, and SEMC regularization:

where λ1=1, λ2=0.1\lambda_{1}=1,\ \lambda_{2}=0.1 is set as default in all experiments if not specially specified. Current epoch number in the training stage is indicated by e∈{0,⋯30}e\in\{0,\cdots 30\}, and initial epoch iei_{e} is used to ensure reliable consistency memory and effective consistency regularization.

Experiments

Datasets. We perform experiments on the PETA , RAP , and PA100K . The PEdesTrian Attribute (PETA) dataset is collected from 10 small-scale person datasets and consists of 19,000 person images, which is divided into 9500 images for the training set, 1900 for the validation set, and 7600 for the test set. Each image is labeled with 61 binary attributes and 4 multi-class attributes. We follow the common experimental protocol , and only 35 attributes whose positive ratios are higher than 5% are used for evaluation. The Richly Annotated Pedestrian (RAP) attribute dataset consists of 33,268 images for training and 8,317 images for testing, a total of 41,585 images extracted from 26 indoor surveillance cameras. Each image is labeled with 69 binary attributes and 3 multi-class attributes. Following the official protocol , 51 binary attributes are adopted to evaluate the recognition performance. The PA100K dataset consists of 100,000 pedestrian images and is split into training, validation, and test sets with a ratio of 8:1:1. Each image is described with 26 commonly used attributes. Considering the identical pedestrian identities between training set and test set on the RAP and PETA, performance on the largest dataset PA100K is more convincible.

Evaluation Protocal. Two types of metrics, i.e., a label-based metric and four instance-based metrics, are adopted to evaluate attribute recognition performance . For the label-based metric, we compute the mean value of classification accuracy of positive samples and negative samples as the metric for each attribute. Then we take an average over all attributes as mean accuracy. For instance-based metrics, accuracy, precision, recall, and F1-score are used.

2 Implementation Details

The proposed method is implemented with PyTorch and trained in an end-to-end manner. We adopt ResNet50 as the backbone network to extract pedestrian image features for a fair comparison. Pedestrian images are resized to 256×192256\times 192 as inputs. Random horizontal mirroring, padding, and random crop are used as augmentations. Adam is employed for training with the weight decay of 0.0005. The initial learning rate equals 0.0001, and the batch size is set to 64. Plateau learning rate scheduler is used with reduction factor 0.1 and loss patience 4. The total epoch number of the training stage is 30. Momentum coefficient α=0.9\alpha=0.9, confidence threshold τ=0.9\tau=0.9 by default. To obtain the stable and reliable spatial memory Mmspa\bm{M}^{spa}_{m} and semantic memory Mmsem\bm{M}^{sem}_{m}, consistency regularizations are added to the classification loss after epoch 4, i.e., ie=4i_{e}=4 in Equation 14 .

3 Comparison to the State of the Arts

In Table 1, we compare the performance of the proposed methods with several existing algorithms on the PETA, RAP, and PA100K. For a fair comparison, besides the performance reported by the papers , we also report the performance of our reimplements based on the same setting described in Section 4.2.

Compared with the performance reported by the paper of MsVAA , VAC , and ALM methods, the SSCsoftSSC_{soft} model achieves better performance on the PETA, PA100K, and RAP without increasing learnable parameters. Compared to the MsVAA model adopted ResNet101 as the backbone network, we achieve 1.93% and 0.53% performance improvements in mA and F1 on the PETA dataset. Compared to the ALM model, which utilizes the complicated combination of FPN, STN and SE modules introducing extra 17% parameters, the SSCsoftSSC_{soft} method achieves 0.22%, 1.19%, and 0.9% performance improvements in mA on three popular datasets. Besides, compared with the performance achieved by our reimplementation of the MsVAA, VAC, and ALM methods, the performance of the SSCsoftSSC_{soft} method has significant improvements from 1.02% to 4.30% in mA on the PETA, PA100K, and RAP, which fully demonstrates the effectiveness of our method.

It can be noticed that the proposed spatial and semantic consistency method substantially outperforms the visual attention consistency (VAC) method . The VAC method hypothesis that global attention regions of random augmentations of the same image are consistent. However, the VAC method focuses on global attention regions of an individual image and cannot generate precise local attention regions for each fine-grained attribute. In addition, for a pair of augmentations of the same image, if global attention regions of one augmentation are precise, the VAC method can improve the performance by aligning the global attention regions of another augmentation with these of the current one. However, if attention regions of both augmentations are inaccurate, the VAC method cannot solve the attention region deviation problem, which can be addressed by our proposed method.

4 Ablation Study and Discussion

In this section, we first investigate the effect of the SPAC and SEMC module by conducting analytical experiments on all three datasets. We then introduce two variants of our methods to demonstrate the effectiveness of spatial and semantic consistency regularizations. Quantitative performance improvements of each attribute on three datasets are presented in the supplementary material.

As shown in Table 2, compared to the baseline method, we have the following observations. First, adopting the SEMC module alone can hardly bring performance improvement. The results prove that, without correct attention regions, attribute semantic features lack discrimination and contain more noise, which is in line with the intuitive hypothesis. Second, adopting the SPAC module can directly bring 2.93%, 0.62%, 2.46% performance improvements in mA on the PETA, PA100K, and RAP, respectively. This improved performance demonstrates that spatial consistency regularization is beneficial for locating the attribute-related regions. Third, when the SPAC module and SEMC module are jointly adopted, our method improves the performance over the baseline model by 3.75%, 1.56%, 4.17% in mA on the PETA, PA100K, and RAP.

To further validate the reasonableness of proposed spatial and semantic consistency regularizations, we implement our method with two variants SSChardSSC_{hard} and SSCfixSSC_{fix}. For the SSChardSSC_{hard} method, we change the Aq\bm{A}^{q} in Equation 3 and Mspa\bm{M}^{spa} in Equation 4 of SPAC module from soft attention maps to binary (hard) attention maps based on a threshold thhard=0th_{hard}=0. For the SSCfixSSC_{fix} method, we first train a baseline model and obtain the qualified CAMs Aq\bm{A}^{q} of positive samples for each attribute according to Equation 3. Then, we fix Mspa\bm{M}^{spa} as Aq\bm{A}^{q} instead of momentum updating to training a new model SSCfixSSC_{fix}. The experimental results of two variants are listed in Table 1. Although method SSChardSSC_{hard} assigns the same weight to each pixel of the region of interest, which is not as flexible as SSCsoftSSC_{soft} and achieves slightly reduced performance, it still achieves competitive performance on PA100k and RAP. Since the SSCsoftSSC_{soft} and SSCfixSSC_{fix} method can get reliable and accurate spatial attention regions Mspa\bm{M}^{spa}, they both achieve the state-of-the-art performance. However, compared to the SSCfixSSC_{fix} method, SSCsoftSSC_{soft} with the momentum updated memory can avoid a two-stage training process and is more suitable for industry application.

5 Effects of SPAC and SEMC Module

Spatial and semantic consistency regularizations are two complementary and indispensable parts of a powerful model. The SPAC module can enhance the localization capability of the backbone network without being disturbed by overfitting and label noise. Based on the precise spatial attention regions of attributes, the backbone network further benefits from the SEMC module to extract intrinsic and discriminative semantic features.

To validate the effectiveness of the proposed SPAC module and SEMC module, we visualize the spatial attention regions and the similarity distributions of spatial and semantic features in Figure 3. The similarities are computed between each pair of images of the same attribute on PA100K. The higher the similarity, the more consistent the attention regions and semantic features of the two images with the same attribute. Compared to the baseline method, as shown in Figure 3(b) and Figure 3(c), we observe that plenty of similarities concentrate on 1, making the probability curve rise rapidly near 1. The same phenomenon can also be observed in other attributes of the PA100K, RAP, and PETA as shown in supplementary material.

6 Hyperparameter Evaluation

There are mainly three key hyperparameters in our method, which are confidence threshold τ\tau, initial epoch iei_{e}, and momentum coefficient α\alpha. We set τ=0.9\tau=0.9, ie=4i_{e}=4, α=0.9\alpha=0.9 if not specially specified. To fully demonstrate the effect of hyperparameters, the following experiments are all conducted on the largest pedestrian attribute dataset PA100K.

Confidence threshold τ\tau is used in Equation 3 and Equation 8 to select reliable spatial attention feature maps Aq\bm{A}^{q} and semantic feature vectors Vq\bm{V}^{q}, which are aggregated to spatial memory Mspa\bm{M}^{spa} and semantic memory Msem\bm{M}^{sem}. As shown in Table 3, with the increase of the confidence threshold τ\tau, there is an obvious performance improvement in mA from 79.63 to 81.28. It is easy to infer that higher confidence threshold τ\tau can select more precise spatial attention feature maps and more discriminative semantic feature vectors. Little performance fluctuation in the other four metrics shows the robustness of the threshold τ\tau.

Momentum coefficient α\alpha is adopted in Equation 4 and Equation 9 to determine the degree of integration of historical features and current batch features. The larger α\alpha is, the fewer historical features are retained. As shown in Table 4, more historical features can bring a few performance improvements.

Conclusion

This paper proposes the consistency framework for pedestrian attribute recognition, which makes full use of the inter-image relation of the same attribute and tackles the spatial attention region deviation problem. Specifically, we propose the SPAC module to pay attention to specific attribute-related spatial regions. We also propose the SEMC module to extract intrinsic and discriminative semantic features for each attribute. Moreover, we implement two variants of our method to demonstrate the efficacy of consistency regularizations. The ablation experiments show that two consistency modules can both bring performance improvements. Our proposed method achieves outstanding performance consistently on the PA100K, RAP, and PETA.

Acknowledgments

This work was supported in part by the National Natural Science Foundation of China (Grant No.61721004 and Grant No.61876181), the Projects of Chinese Academy of Science (Grant QYZDB-SSW-JSC006), the Strategic Priority Research Program of Chinese Academy of Sciences (Grant No. XDA27000000), and the Youth Innovation Promotion Association CAS.

References