ABD-Net: Attentive but Diverse Person Re-Identification
Tianlong Chen, Shaojin Ding, Jingyi Xie, Ye Yuan, Wuyang Chen, Yang Yang, Zhou Ren, Zhangyang Wang
Introduction
Person Re-Identification (Re-ID) aims to associate individual identities across different time and locations. It embraces many applications in intelligent video surveillance. Given a query image and a large set of gallery images, person Re-ID represents each image with a feature embedding, and then ranks the gallery images in terms of feature embeddings’ similarities to the query. Despite the exciting progress in recent years, person Re-ID remains to be extremely challenging in practical unconstrained scenarios. Common challenges arise from body misalignment, occlusion, background perturbance, view point changes, pose variations and noisy labels, among many others .
Substantial efforts have been devoted to addressing those various challenges. Among them, incorporating body part information has empirically proven to be effective in enhancing the feature robustness against body misalignment, incomplete parts, and occlusions. Motivated by such observations, the attention mechanism was introduced to enforce the features to mainly capture the discriminative appearances of human bodies (or certain body parts). Since then, the attention-based models have boosted person Re-ID performance much.
On a separate note, the feature embeddings are used to compute similarities between images, typically based on the Euclidean distance, to return the closest matches. Sun et al. pointed out that correlations among feature embeddings would significantly compromise the matching performance. The low feature correlation property is, however, not naturally guaranteed by attention-based models. Our observation is that those attention-based models are often more prone to higher feature correlations, because intuitively, the attention mechanism tends to have features focus on a more compact subspace (such as foreground instead of the full image, see Fig.~1 for examples).
In view of the above, we argue that a more desirable feature embedding for person Re-ID should be both attentive and diverse: the former aims to correct misalignment, eliminate background perturbance, and focus on discriminative local parts of body appearances; the latter aims to encourage lower correlation between features, and therefore better matching, and potentially make feature space more comprehensive. We propose an Attentive but Diverse Network (ABD-Net), that strives to integrate attention modules and diversity regularization and enforces them throughout the entire network. The main contributions of ABD-Net are outlined as below:
We incorporate a compound attention mechanism into ABD-Net, consisting of Channel Attention Module (CAM) and Position Attention Module (PAM). CAM facilitates channel-wise, feature-level information aggregation, while PAM captures the spatial awareness of body and part positions. They are found to be complementary and altogether benefit Re-ID.
We introduce a novel regularization term called spectral value difference orthogonality (SVDO) that directly constrains the conditional number of the weight Gram matrix. SVDO, efficiently implemented, is applied to both activations and weights, and is shown to effectively reduce learned feature correlations.
We perform extensive experiments on Market-1501 , DukeMTMC-Re-ID , and MSMT17 . ABD-Net significantly outperforms existing methods, achieving new state-of-the-art on all three popular benchmarks. We also verify that the attentive and diverse terms each contributes to a performance gain, through rigorous ablation studies and visualizations.
Related Work
Person Re-ID has two key steps: obtaining a feature embedding and performing matching under some distance metric . We mainly review the former where both handcrafted features and learned features were studied. In recent years, the prevailing success of convolutional neural networks (CNNs) in computer vision has made person Re-ID no exception. Due to many problem-specific challenges such as occlusion/misalignment, incomplete body parts, as well as background perturbance/view point changes, naively applying CNN backbones to feature extraction may not yield ideal Re-ID performance. Both image-level features and local features extracted from body parts prove to enhance the robustness. Many part-based methods have achieved superior performance . We refer readers to for a more comprehensive review.
2 Attention Mechanisms in Person Re-ID
Several studies proposed to integrate attention mechanism into deep models to address the misalignment issue in person Re-ID. Zhao et al. proposed a part-aligned representation based on a part map detector for each predefined body part. Yao et al. proposed a Part Loss Network which defined a loss for each average pooled body part and jointly optimized the summation losses. Si et al. proposed a dual attention matching network based on an inter-class and an intra-class attention module to capture the context information of video sequences for person Re-ID. Li et al. proposed a multi-task learning model that learns hard region-level and soft pixel-level attention jointly to produce more discriminative feature representations. Xu et al. used pose information to learn attention masks for rigid and non-rigid parts, and then combined the global and part features as the final feature embedding.
Our proposed attention mechanism differs from previous methods in several aspects. First, previous methods only use attention mechanisms to extract part-based spatial patterns from person images, which are usually focus in the foregrounds. In contrast, ABD-Net combines spatial and channel clues; besides, our added diversity constraint will avoid the overly correlated and redundant attentive features. Second, our attention masks are directly learned from the data and context, without relying on manually-defined parts, part region proposals, nor pose estimation . Our two attention modules are embedded within a single backbone, making our model lighter-weight than the multi-task learning alternatives .
3 Diversity via Orthogonality
Orthogonality has been widely explored in deep learning to encourage the learning of informative and diverse features. In CNNs, several studies perform regularization using “hard orthogonality constraints”, which typically depends on singular value decomposition (SVD) to strictly constrain their solutions on a Stiefel manifold. The similar idea was first exploited by for person Re-ID, where the authors performed SVD on the weight matrix of the last layer, in an effort to reduce feature correlations. Despite the effectiveness, SVD-based hard orthogonality constraints are computationally expensive, and sometimes appear to limit the learning flexibility.
Recent studies also investigated “softer” orthogonality regularizations by enforcing the Gram matrix of each weight matrix to be close to an identity matrix, under Frobenius norm or spectral norm . We propose a novel spectral value difference orthogonality (SVDO) regularization that directly constrains the conditional number of the Gram matrix. Also contrasting from that apply orthogonality only to CNN weights, we enforce the new regularization on both hidden activations and weights.
Attentive but Diverse Network
In this section, we first introduce the two attention modules, followed by the new diversity (orthogonality) regularization. We then wrap them up and describe the overall architecture of ABD-Net.
The goal of attention for Re-ID is to focus on person-related features while eliminating irrelevant backgrounds. Inspired by the successful idea in segmentation , we integrate two complementary attention mechanisms: Channel Attention Module (CAM) and Positional Attention Module (PAM). The full configurations for CAM and PAM can be found in the supplementary.
The high-level convolutional channels in a trained CNN classifier are well-known to be semantic-related and often category-selective. In the person Re-ID case, we hypothesize that the high-level channels in the person Re-ID case are also “grouped”, i.e., some channels share similar semantic contexts (such as foreground person, occlusions, or background) and are more correlated with each other. CAM is designed to group and aggregate those semantically similar channels.
where represents the impact of channel on channel . The final output feature map E is calculated by equation (2):
is a hyperparameter to adjust the impact of CAM.
1.2 Position Attention Module
2 Diversity: Orthogonality Regularization
Following , we enforce diversity via orthogonality, yet derive a novel orthogonality regularizer term. It is applied to both hidden features and weights, of both convolutional and fully-connected layers. Orthogonality regularizer on feature space (short for O.F. hereinafter) is to reduce feature correlations that can directly benefit matching. The orthogonal regularizer on weight (O.W.) encourages filter diversity and enhances the learning capacity.
Many orthogonality methods , including the prior work on person Re-ID , enforce hard constraints on orthogonality of weights, whose computations rely on SVD. However, computing SVD on high-dimensional matrices is expensive, urging for the development of soft orthogonality regularizers. Many existing soft regularizers restrict the Gram matrix of F to be close to an identity matrix under Frobenius norm that can avoid the SVD step while being differentiable. However, the gram matrix for an overcomplete F cannot reach identity because of rank deficiency, making those regularizers biased. hence introduced the spectral norm-based regularizer that effectively alleviates the bias.
We propose a new option to enforce the orthogonality via directly regularizing the conditional number of :
where is the coefficient and denotes the condition number of F, defined as the ratio of maximum singular value to minimum singular value of . Naively solving will take one full SVD. To make it computationally more tractable, we convert (3) into a spectral value difference orthogonality (SVDO)The reason why we choose to penalize the difference between and rather than the ratio of them is to avoid numerical instability caused by dividing a very small , which we find happen frequently in our experiments. regularization:
where and denote the largest and smallest eigenvalues of , respectively.
We use auto-differentiation to obtain the gradient of SVDO, however, this computation still contains the expensive eigenvalue decomposition (EVD). To bypass EVD, we refer to the power iteration method to approximate the eigenvalues. We start with a random initialized , and then iteratively perform equation (5) (two times by default):
where X in equation (5) is for computing , and for . In that way, the computation of SVDO becomes practically efficient.
3 Network Architecture Overview
The overall architecture of the proposed ABD-Net is shown in Fig.~4. ABD-Net is compatible with most common feature extraction backbones, such as ResNet , InceptionNet , and Densenet . Unless otherwise specified, we use ResNet-50 as the default backbone network due to its popularity in Re-ID .
We add a CAM and O.F. on the outputs of res_conv_2 block. The regularized feature map is used as the input of res_conv_3. Next, after the res_conv_4 block, the network splits into a global branch and an attentive branch in parallel. We apply O.W. on all conv layers in our ResNet-50 backbone, , from res_conv_1 to res_conv_4 and the two res_conv_5 in both branches. The outputs of two branches are concatenated as the final feature embedding.
The attentive branch uses the same res_conv_5 layer as that in ResNet-50. The output feature map is then fed into a reduction layerA reduction layer consists of a linear layer, batch normalization, ReLU, and dropout. See: https://github.com/KaiyangZhou/deep-person-reid with O.F. applied yielding a smaller feature map . We feed into a CAM and a PAM simultaneously, both with O.F. constraints. The outputs from both attentive modules are concatenated with the input , and altogether go through a global average pooling layer, ending up with a -dimension feature vector.
In the global branch, after res_conv_5For both two res_conv_5 layers in two branches, we removed the down-sampling layer, in order for larger feature maps., the feature map is fed into a global average-pooling layer followed by a reduction layer, leading to a -dimension feature vector. The global branch intends to preserve global context information in addition to the attentive branch features.
Eventually, ABD-Net is trained under the loss function consisting of a cross entropy loss, a hard mining triplet loss, and orthogonal constraints on feature (O.F.) and on weights (O.W.) penalty terms:
where and stand for the SVDO penalty term applied to the hidden features and weights, respectively. and are hyper-parameters.
Experiments
To evaluate ABD-Net, we conducted experiments on three large-scale person re-identification datasets: Market-1501 , DukeMTMC-Re-ID and MSMT17 . First, we report a set of ablation study (mainly on Market-1501 and DukeMTMC-Re-ID) to validate the effectiveness of each component. Second, we compare the performance of ABD-Net against existing state-of-the-art methods on all three datasets. Finally, we provide more visualizations and analysis to illustrate how ABD-Net has achieved its effectiveness.
Market-1501 comprises 32,668 labeled images of 1,501 identities captured by six cameras. Following , 12,936 images of 751 identities are used for training, while the rest are used for testing. Among the testing data, the test probe set has 3,368 images of 750 identities. The test gallery set also includes 2,793 additional distractors.
DukeMTMC-Re-ID contains 36,411 images of 1,812 identities. These images are captured by eight cameras, among which 1,404 identities appear in more than two cameras and 408 identities (distractors) appear in only one camera. The 1,404 identities are randomly divided, with 702 identities for training and the others for testing. In the testing set, one query image for each ID per camera is chosen for the probe set, while all remaining images including distractors are in the gallery.
MSMT17 is the current largest publicly-available person Re-ID dataset. It has 126,441 images of 4,101 identities captured by a 15-camera network (12 outdoor, 3 indoor). We follow the training-testing split of . The video is collected with different weather conditions at three-time slots (morning, noon, afternoon). All annotations, including camera IDs, weathers and time slots, are available. MSMT17 is significantly more challenging than the other two, due to its massive scale, more complex and dynamic scenes. Additionally, the amount of methods that report on this dataset is limited since it is recently released.
2 Implementation Details and Evaluation
During training, the input images are re-sized to and then augmented by random horizontal flip, normalization, and random erasing . The testing images are re-sized to and augmented only by normalization. In our experiments, the sizes of feature maps and are , and , respectively. We set the dimension of features (, ) after global average-pooling both equal to 1024, leading to a 2048-dimensional final feature embedding for matching.
With the ImageNet-pretrained ResNet-50 backbone, we used the two-step transfer learning algorithm to fine-tune the model. First, we freeze the backbone weights and only train the reduction layers, classifiers and all attention modules for 10 epochs with only the cross entropy loss and triplet loss applied. Second, all layers are freed for training for another 60 epochs, with the full loss (6) applied. We set , and , and the margin parameter for triplet loss .
Our network is trained using 2 Tesla P100 GPUs with a batch size of 64. Each batch contains 16 identities, with 4 instances per identity. We use the Adam optimizer with the base learning rate initialized to , then decayed to , after 30, 40 epochs, respectively. The training takes about 4 hours on the Market-1501 dataset.
We adopt standard Re-ID metrics: top-1 accuracy, and the mean Average Precision (mAP). We consider mAP to be a more reliable indicator for Re-ID performance.
3 Ablation Study of ABD-Net
To verify the effects of attention modules and orthogonality regularization in ABD-Net, we incrementally evaluate each module on Market-1501 and DukeMTMC-Re-ID. We choose ResNet-50 For the fairness of ablation study, we use two duplicated branches with the same res_conv_5 like the structure in ABD-Net as shown in Fig.~4. Data augmentation and dropout are applied. with the cross entropy loss (XE) as the baseline. Nine variants are then constructed on top of the baselineNote that (1) CAM is used in two places of ABD-Net; (2) ABD-Net adopts O.F. + O.W. + PAM + CAM.: a) baseline (XE) + PAM; b) baseline (XE) + CAM; c) baseline (XE) + PAM + CAM; d) baseline (XE) + O.F.; e) baseline (XE) + O.W.; f) baseline (XE) + O.F. + O.W.; g) baseline + SVD layer (similar to SVD-Net ); h) ABD-Net (XE), that sets in (6); and i) ABD-Net, that uses the full loss (6).
Table 1 presents the ablation study results, from which several observations could be drawn:
Using either PAM or CAM improves the baseline on both datasets. The combination of the two different attention mechanisms gains further improvements, demonstrating their complementary power over utilizing either alone.
Using either O.F. or O.W. consistently outperforms the baseline on both datasets, and their combination leads to further gains which validates the effectiveness of our orthogonality regularizations. We also observe that the proposed SVDO-based O.W. empirically performs better than the SVD layer, potentially because SVD layer acts as a “hard constraint” and hence restricts the learning capability of the ResNet-50 backbone.
By combining “attention” and “diversity”, ABD-Net (XE) sees further boosts. For example, on Market-1501, ABD-Net (XE) outperforms the “no attention” counterpart (baseline (XE) + O.F. + O.W.) by a margin of 1.50% (top-1)/3.60% (mAP), and it outperforms “no diversity” counterpart (baseline (XE) + O.F. + O.W.) by 2.20% (top-1)/7.40% (mAP). Moreover, there are further performance improvements when we enforce diversity in the attention mechanism. Finally, the full ABD-Net further benefits from adding triplet loss.
4 Comparison to State-of-the-art Methods
We compare ABD-Net against the state-of-the-art methods on Market-1501, DukeMTMC-Re-ID and MSMT17, as shown in Tables 2, 3, and 4, respectively. For fair comparison, no post-processing such as re-ranking or multi-query fusion was used for our methods.
ABD-Net has clearly yielded overall state-of-the-art performance on all datasets. Specifically, on DukeMTMC-Re-ID, ABD-Net obtains top-1 accuracy and mAP, which significantly outperforms all existing methods. On MSMT17, ABD-Net presents a clear winner case too. On Market-1501, its top-1 accuracy (95.60%) slightly lags behind Local CNN (95.90%) and MGN (95.70%); yet ABD-Net clearly surpasses all existing methods in terms of mAP (88.28%, outperforming the closest competitor by a large margin of 0.88%).
Specifically, we emphasize the comparison between ABD-Net and existing attention-based methods (marked by in the Tables 2 3). As shown in Table 2 and 3, ABD-Net achieves at least top-1 and % mAP improvement on Market-1501, compared to the closest attention-based prior work CNet . On DukeMTMC, the margin becomes for top-1 and % for mAP. We also considered SVDNet and HA-CNN which also proposed to generate diverse and uncorrelated feature embeddings. ABD-Net surpasses both with significant top-1 and mAP improvement. Overall, our observations endorse the superiority of ABD-Net by combing “attentive” and “diverse”.
5 Visualizations
We conduct a set of attention visualizationsGrad-CAM visualization method : https://github.com/utkuozbulak/pytorch-cnn-visualizations; RAM visualization method for testing images. More results can be found in the supplementary. on the final output feature maps of the baseline (XE), baseline (XE) + PAM + CAM, and ABD-Net (XE), as shown in Fig.~5. We notice that the feature maps from the baseline show little attentiveness. PAM + CAM enforces the network to focus more on the person region, but the attention regions can sometimes overly emphasize some local regions (e.g., clothes), implying the risk of overfitting person-irrelevant nuisances. Most channels focus on the similar region may also cause a high correlation in the feature embeddings. In contrast, the attention of ABD-Net (XE) can strike a better balance: it focuses on more of the local parts of the person’s body while still being able to eliminate the person from backgrounds. The attention patterns now differ more from person to person, and the feature embeddings become more decorrelated and diverse.
Fig.~8 shows the t-SNE visualization on feature distributions from Baseline, Baseline + PAM + CAM and ABD-Net (XE) using t-SNE. Compared with Baseline, although attentive features from Baseline + PAM + CAM make ID 94 and ID 156 in cycle B slightly distinguishable, ABD-Net enlarges the intra-class distance of ID 521 in cycle A. It makes the features from ID 94 and ID 156 more discriminative, meanwhile the features from ID 521 also lie in a compact region.
Fig.~9 shows Re-ID visual examples of ABD-Net (XE), Baseline + PAM + CAM and Baseline on Market-1501. They indicate that ABD-Net succeeds in finding more true positives than Baseline + PAM + CAM model, even when the persons in the images are under significant view changes and appearance variations.
Conclusion
This paper proposes a novel Attentive but Diverse Network (ABD-Net) to learn more representative, robust, discriminative feature embeddings for person Re-ID. ABD-Net demonstrates its state-of-the-art performance through extensive experiments where the ablations and visualizations show that each added component substantially contributes to its final performance. In the future, we will generalize the design concept of ABD-Net to other computer vision tasks.