Learning Discriminative Features with Multiple Granularities for Person Re-Identification

Guanshuo Wang, Yufeng Yuan, Xiong Chen, Jiwei Li, Xi Zhou

Introduction

Person re-identification (Re-ID) is a challenging task to retrieve a given person among all the gallery pedestrian images captured across different security cameras. Due to the scene complexity of images from surveillance videos, the main challenges for person Re-ID come from large variations on persons such as pose, occlusion, clothes, background clutter, detection failure, etc. The prosperity of deep convolutional network has introduced more powerful representations with better discrimination and robustness for pedestrian images, which pushed the performance of Re-ID to a new level. Some recent deep Re-ID methods (Saquib Sarfraz et al., 2018; Sun et al., 2018; Chang et al., 2018; Shen et al., 2018a; Li et al., 2018; Shen et al., 2018b; Chen et al., 2018) have achieved breakthrough with high-level identification rates and mean average precision.

The intuitive approach of pedestrian representations is to extract discriminative features from the whole body on images. The aim of global feature learning is to capture the most salient clues of appearance to represent identities of different pedestrians. However, high complexities for images captured in surveillance scenes usually restrict the accuracy for feature learning in large scale Re-ID scenarios. Due to the limited scale and weak diversity of person Re-ID training datasets, some non-salient or infrequent detailed information can be easily ignored and make no contribution for better discrimination during global feature learning procedure, which makes global features hard to adapt similar inter-class common properties or large intra-class differences.

To relieve this dilemma, locating significant body parts from images to represent local information of identities has been confirmed to be an effective approach for better Re-ID accuracy in many previous works. Each located body part region only contains a small percentage of local information from the whole body, and at the same time distraction by other related or unrelated information outside the regions is actually filtered by locating operations, with which local features can be learned to concentrate more on identities and used as an important complement for global features. Part-based methods for person Re-ID can be divided into three main pathways according to their part locating methods: 1) Locating part regions with strong structural information such as empirical knowledge about human bodies (Cheng et al., 2016; Li et al., 2017b; Sun et al., 2018; Zhang et al., 2017) or strong learning-based pose information (Su et al., 2017; Zhao et al., 2017b); 2) Locating part regions by region proposal methods (Yao et al., 2017; Li et al., 2017a); 3) Enhancing features by middle-level attention on salient partitions (Zhao et al., 2017a; Liu et al., 2017b; Liu et al., 2017a; Li et al., 2018). However, obvious limitations impede the effectiveness of these methods. First, pose or occlusion variations can affect the reliability of local representation. Second, these methods almost only focus on specific parts with fixed semantics, but cannot cover all the discriminative information. Last but not least, most of these methods are not end-to-end learning process, which increases the complexity and difficulty of feature learning.

In this paper, we propose a feature learning strategy combining global and local information in different granularities. As shown in Figure 1, various numbers of partition stripes introduce a diversity of content granularity. We define the original image containing only one whole partition with global information as the coarsest case, and as the number of partitions increase, features of local parts can concentrate more on finer discriminative information in each part stripe, filtering information on the other stripes. Since deep learning mechanism can capture approximate response preferences on the main body from the whole image, it is also possible to capture more fine-grained saliency for local features extracted from smaller part regions. Notice that these part regions are not necessary to be located partitions with specific semantics, but only a piece of equally-split stripe on the original images. From the observation, we find that the granularity of discriminative responses indeed becomes finer as the number of horizontal stripes increases. Based on this motivation, we design the Multiple Granularity Network (MGN), a multi-branch network architecture divided into one global and two local branches with delicated parameters from the 4th residual stage of the ResNet-50 (He et al., 2016) backbone. In each local branch of MGN, we divide globally-pooled feature maps into different numbers of stripes as part regions to learn local feature representations independently, referring the methods in (Sun et al., 2018).

Comparing to the previous part-based methods, our method only utilize equally-divided parts for local representation, but can achieve outstanding performance exceeding all previous methods. Besides, our method is completely a end-to-end learning process, which is easy for learning and implementation. Extensive experiment results show that our method can achieve state-of-the-art performances on several mainstream Re-ID datasets, even with settings without any additional external data or re-ranking (Zhong et al., 2017) operation.

Related Works

With the prosperity of deep learning, feature learning by deep networks has become a common practice in person Re-ID tasks. (Li et al., 2014; Yi et al., 2014) first introduce deep siamese network architecture into Re-ID and combine the body part feature learning, achieving higher performances comparing to the contemporary hand-crafted methods. (Zheng et al., 2016) proposes ID-discriminative Embedding (IDE) with simple ResNet-50 backbones as a baseline of the performance level for modern deep Re-ID systems. A number of methods are proposed to improve the performance for deep person Re-ID. In (Ahmed et al., 2015; Varior et al., 2016), mid-level features of image pairs are computed to depict interrelation of local parts with carefully designed mechanism. (Xiao et al., 2016) introduces Domain Guided Dropout to enhance the generalization ability across different domains of pedestrian scenarios. (Zhong et al., 2017) brings re-ranking strategy into Re-ID tasks to modify the ranking results for accuracy improvement.

Recently some deep Re-ID methods pushed the performances to a new level comparing to the former systems. (Zhang et al., 2017) introduces a part-based alignment matching in training phase with shortest path programming and mutual learning to improve metric learning performance. (Bai et al., 2017; Sun et al., 2018) both equally slice the feature maps of input images into several stripes in vertical orientation. (Bai et al., 2017) merges slices of local features with LSTM network and combine with global features learned from classification metric learning. Instead (Sun et al., 2018) directly concatenates the features from local parts as the final representation, and applies refined part pooling to modify the mapping validation of part features. However, according to the report in (Zhang et al., 2017), these systems just achieve similar performances as human, which we still need a highway to surpass.

Among all the strategies for performance improvement, we argue that combining the local representations from parts of images is the most effective. As mentioned in Section 1, we summarize three main pathways for part-based learning: determining regions according to structural information about human body, locating body parts by region proposal methods and enhancing features by spatial attention. In (Cheng et al., 2016; Li et al., 2017b; Sun et al., 2018), images are all split into several stripes in horizontal orientation according to intrinsic human body structure knowledge, on which local feature representations are learned. (Zhao et al., 2017b; Su et al., 2017) utilize structural information of body landmarks predicted by pose estimation methods to crop more accurate region areas with semantics. To locate semantic partitions without strongly learning-based predictors, region proposal methods such as (Girshick, 2015; Jaderberg et al., 2015) are employed in some part-based methods (Liu et al., 2017b; Yao et al., 2017; Zhao et al., 2017a; Li et al., 2017a; Li et al., 2018). Attention information can be a powerful complement for discrimination, which are enhanced in (Liu et al., 2017a; Liu et al., 2017b; Li et al., 2018). In our proposed method, we only use simple horizontal stripes as part regions for local feature learning but achieve outstanding performances.

Loss functions are used as supervisory signals in feature learning. In the training phase for deep Re-ID systems, the most common loss functions are classification losses and metric losses. Softmax loss is almost the only choice of classification loss function for its strong robustness to various kinds of multi-class classification tasks, which can be used individually (Ahmed et al., 2015; Zheng et al., 2016; Xiao et al., 2016; Sun et al., 2018; Liu et al., 2017b; Yao et al., 2017; Li et al., 2017a; Li et al., 2018) or combined with other losses (Li et al., 2014; Cheng et al., 2016; Zhang et al., 2017; Bai et al., 2017) in embedding learning procedures for Re-ID. For metric losses used in embedding learning for Re-ID, there are more variants with different ranking metrics. Contrastive loss (Hadsell et al., 2006) is commonly used in siamese-liked networks (Varior et al., 2016), which focuses on maximizing the distances between inter-class pairs and minimizing that between intra-class pairs. Triplet loss (Hoffer and Ailon, 2015; Schroff et al., 2015) enforces a margin between the intra and inter distances to the same anchor sample with in a triplet. Based on triplet loss, many variants (Cheng et al., 2016; Hermans et al., 2017; Chen et al., 2017a; Song et al., 2016) are proposed to solve learning or performance issues in metric learning. We employ a setting of joint learning with both softmax and triplet losses in our proposed method.

Multiple Granularity Network

Figure 2 shows feature response maps of a certain image extracted from the IDE baseline model (Zheng et al., 2016) and a part-based model based on IDE. We can observe that even if no explicit attention mechanisms are imposed to enhance the preferences to some salient components, the deep network can still learn the preliminary distinction of response preferences on different body parts according to their inherent semantic meanings. However, to eliminate the distraction of unrelated patterns in pedestrian images with high complexity, higher responses are just concentrated on the main body of pedestrians instead of any concrete body parts with semantic patterns. As we narrow the area of represented regions and train as a classification task to learn local features, we can observe that responses on the local feature maps starts to cluster on some salient semantic patterns, which also varies with the sizes of represented regions. This observation reflects a relationship between the volume of image contents, i.e. the granularity of regions, and capability of deep networks to focus on specific patterns for representations. We believe this phenomenon comes from the limitation of information in restricted regions. In general, comparing from a global image, it is intuitively hard to discriminate the identities for pedestrians from a local part region. Supervising signals of the classification task enforce the features to be correctly classified as the target identity, which also push the learning procedure trying to explore useful fine-grained details among limited information.

Actually, local feature learning in previous part-based methods only introduces a basic granularity diversity of partitions to the total feature learning procedure with or without empirical prior knowledge. Assume that there are appropriate levels of granularities, details with most discriminative information might be almost concentrated by deep networks. Motivated by the observation and analysis above, we propose the Multiple Granularity Network (MGN) architecture to combine global and multi-granularity local feature learning for more powerful pedestrian representation.

The architecture of Multiple Granularity Network is shown in Figure 3. The backbone of our network is ResNet-50 which helps to achieve competitive performances in some Re-ID systems (Zhang et al., 2017; Bai et al., 2017; Sun et al., 2018). The most obvious modification different from the original version is that we divide the subsequent part after res_conv4_1 block into three independent branches, sharing the similar architecture with the original ResNet-50.

Table 1 lists the settings of these branches. In the upper branch, we employ down-sampling with a stride-2 convolution layer in res_conv5_1 block, following a global max-pooling (GMP) (Almazan et al., 2018) operation on the corresponding output feature map and a 1×11\times 1 convolution layer with batch normalization (Ioffe and Szegedy, 2015) and ReLU to reduce 2048-dim features zgG\mathbf{z}_{g}^{G} to 256-dim fgG\mathbf{f}_{g}^{G}. This branch learns the global feature representations without any partition information, so we name this branch as the Global Branch.

The middle and lower branches both share the similar network architecture with Global Branch. The difference is that we employ no down-sampling operations in res_conv5_1 block to preserve proper areas of reception fields for local features, and output feature maps in each branch are uniformly split into several stripes in horizontal orientation, on which we independently perform the same following operations as Global Branch to learn local feature representations. We call these branches Part-N Branch, where N refers to the number of partitions on the unreduced feature maps, e.g. the middle and lower branches in Figure 3 can be named as Part-2 and Part-3 Branch.

During testing phases, to obtain the most powerful discrimination, all the features reduced to 256-dim are concatenated as the final feature, combining both the global and local information to perfect the comprehensiveness for learned features.

2. Loss Functions

To unleash the discrimination ability of the learned representations of this network architecture, we employ softmax loss for classfication, and triplet loss for metric learning as the loss functions in training phases, which are both widely used in various deep Re-ID methods.

For basic discrimination learning, we regard the identification task as a multi-class classfication problem. For ii-th learned features fi\mathbf{f}_{i}, softmax loss is formulated as:

where Wk\mathbf{W}_{k} corresponds to a weight vector for class kk, with the size of mini-batch in training process N and the number of classes in the training dataset C. Different from the traditional softmax loss, the form we employ here abandons bias terms in linear multi-class classifiers according to (Wang et al., 2017), which contributes to better discrimination performances. Among all the learned embeddings, we employ the softmax loss to the global features before 1×11\times 1 convolution reduction {zgG,zgP2,zgP3}\{\mathbf{z}_{g}^{G},\mathbf{z}_{g}^{P2},\mathbf{z}_{g}^{P3}\} and part features after reduction {fpiP2∣i=12,fpiP3∣i=13}\{\mathbf{f}_{p_{i}}^{P2}|_{i=1}^{2},\mathbf{f}_{p_{i}}^{P3}|_{i=1}^{3}\}.

All the global features after reduction {fgG,fgP2,fgP3}\{\mathbf{f}_{g}^{G},\mathbf{f}_{g}^{P2},\mathbf{f}_{g}^{P3}\} are trained with triplet loss to enhance ranking performances. We use the batch-hard triplet loss (Hermans et al., 2017), an improved version based on the original semi-hard triplet loss. This loss function is formulated as follows:

where fa(i),fp(i),fn(i)\textbf{f}_{a}^{(i)},\textbf{f}_{p}^{(i)},\textbf{f}_{n}^{(i)} are the features extracted from anchor, positive and negative samples receptively, and α\alpha is the margin hyperparameter to control the differences of intra and inter distances. Here positive and negative samples refer to the pedestrians with same or different identity with the anchor. The candidate triplets are built by the furthest positive and closest negative sampled pairs, i.e. the hardest positive and negative pairs in a mini-batch with PP selected identities and KK images from each identity. This improved version of triplet loss enhances the robustness in metric learning, and further improve the performances at the same time.

In MGN achitecture, to avoid loss weight tuning troubles and difficulties in convergence, we novelly propose classfication-before-metric achitecture, which applies the softmax losses to reduced 256-dim local features in Part-2 and Part-3 Branches, and all the non-reduced global-pooled 2048-dim global features, but applies triplet losses to all the reduced features, different from existing methods using triplet losses. This setting is inspired from coarse-to-fine mechanism, regarding non-reduced features as coarse information to learn classification and reduced features as fine information with learned metric. The proposed setting achieves robust convergence comparing to that of imposing joint effects at the same level of reduced features. Besides, we employ no triplet loss on local features. Due to misalignment or other issues, the contents of local regions might vary dramatically, which makes the triplet loss tend to corrupt the model during training.

3. Discussions

In our proposed Multiple Granularity Network architecture, there are some issues worth our separate discussion. In this paragraph, we specifically discuss the issues as follows:

Multi-branch architecture According to our initial motivation for MGN architecture, it seems to be reasonable that the global and local representations are both learned in one single branch. We can directly split the same final feature maps extracted by res_conv5_3 in different numbers of stripes, and apply corresponding supervisory signals as our proposed methods. However, we find this setting is not efficient for further performance improving. Borrowing the ideas in (Sun et al., 2015), the reason might be that the branches sharing the similar network architecture(mainly the fourth residual stage of ResNet-50) just response to different levels of detailed information on images. Learning features in multiple granularities with one mixed single branch might dilute the importance of detailed information. Besides, we try to split the backbone network after shallower or deeper layers, which also achieve no better performances.

Diversity of granularity Three branches in our network architecture actually learn representing information with different perferences. Global Branch with larger reception field and global max-pooling captures integral but coarse features from the pedestrian images, and features learned by Part-2 and Part-3 Branches without strided convolution and split parts of stripes tend to be local but fine. The branch with more partitions will learn finer representation for pedestrian images. Branches learning different preferences can cooperatively supplement low-level discriminating information to the common backbone parts, which is the reason for performance boosting in any single branch.

Experiment

To capture more detailed information from pedestrian images, we refer to (Sun et al., 2018) and resize input images to 384×128384\times 128. We use the weights of ResNet-50 pretrained on ImageNet (Deng et al., 2009) to initialize the backbone and branches of MGN. Notice that different branches in the network are all initialized with the same pretrained weights of the corresponding layers after the res_conv4_1 block. During training phases, we only deploy random horizontal flipping to images in the training dataset for data augmentation. Each mini-batch is sampled with randomly selected P identities and randomly sampled K images for each identity from the training set to cooperate the requirement of triplet loss. Here we recommend to set P=16P=16 and K=4K=4 to train our proposed model. For the margin parameter for triplet loss, we set to 1.2 in all our experiments. We choose SGD as the optimizer with momentum 0.9. The weight decay factor for L2 regularization is set to 0.0005. As for the learning rate strategy, we set the initial learning rate to 0.01, and decay the learning rate to 1e-3 and 1e-4 after training for 40 and 60 epochs. The total training process lasts for 80 epochs. During evaluation, we both extract the features corresponding to original images and the horizontally flipped versions, then use the average of these as the final features. Our model is implemented on PyTorch framework. To conduct a complete training procedure on Market-1501 dataset, it takes about 2 hours with data-parallel acceleration by two NVIDIA TITAN Xp GPUs. All our experiments on different datasets follow the settings above.

2. Datasets and Protocols

The experiments to evaluate our proposed method are conducted on three mainstream Re-ID datasets: Market-1501 (Zheng et al., 2015), DukeMTMC-reID (Zheng et al., 2017b) and CUHK03 (Li et al., 2014). It is necessary to introduce these datasets and their evaluation protocols before we show our results.

Market-1501 This dataset includes images of 1,501 persons captured from 6 different cameras. The pedestrians are cropped with bounding-boxes predicted by DPM detector (Felzenszwalb et al., 2008). The whole dataset is divided into training set with 12,936 images of 751 persons and testing set with 3,368 query images and 19,732 gallery images of 750 persons. There are single-query and multiple-query modes in evaluation, the difference of which is the number of images from the same identity. In multiple-query mode, all features extracted from the images of a person captured by the same camera are merged by avg- or max-pooling, which contains more complete information than single query mode with only 1 query image.

DukeMTMC-reID This dataset is a subset of the DukeMTMC (Ristani et al., 2016) used for person re-identification by images. It consists of 36,411 images of 1,812 persons from 8 high-resolution cameras. 16,522 images of 702 persons are randomly selected from the dataset as the training set, and the remaining 702 persons are divided into the testing set where contains 2,228 query images and 17,661 gallery images. It might be the most challenging datasets for person Re-ID at present, with common situations in high similarity across persons and large variations within the same identity.

CUHK03 This dataset consists of 14,097 images of 1,467 persons from 6 cameras. Two types of annotations are provided in this dataset: manually labeled pedestrian bounding boxes and DPM-detected bounding boxes. Originally the whole dataset is divided into 20 random splits for cross-validation, which is designed for hand-crafted methods and very time-consuming to conduct experiments for deep-learning-based methods.

Protocols In our experiments, to evaluate the performances of Re-ID methods, we report the cumulative matching characteristics (CMC) at rank-1, rank-5 and rank-10, and mean average precision (mAP) on all the candidate datasets. On Market-1501 dataset, we conduct experiments both in single-query and multiple-query mode. On CUHK03 dataset, to simplify the evaluation procedure and meanwhile enhance the accuracy of the performance reflected by the results, we adopt the protocol used in (Zhong et al., 2017).

3. Comparison with State-of-the-Art Methods

We compare our proposed method with current state-of-the-art methods on all the candidate datasets to show our considerable performance advantage over all the existing competitors. Results in detail are given as follow:

Market-1501 The results on Market-1501 dataset is shown in Table 2. For the special effects of re-ranking method for improvement on mAP and rank-1 accuracy, we divide the results into two groups according to whether re-ranking is implemented or not. In single query mode, PCB+RPP (Sun et al., 2018) achieved the best published result without re-ranking , but our MGN achieves Rank-1/mAP=95.7%/86.9%, exceeding the former method by 1.9% in Rank-1 accuracy and 5.3% in mAP. After implementing re-ranking, the result can be improved to Rank-1/mAP=96.6%/94.2%, which surpasses all existing methods by a large margin.

Figure 4 shows top-10 ranking results for some given query pedestrian images. The first two results shows the great robustness: regardless of the pose or gait of these captured pedestrian, MGN features can robustly represent discriminative information of their identities. The third query image is captured in a low-resolution condition, losing an amount of important information. However, from some detailed clues such as the strap of the bag and his black suits, most of the ranking results are accurate and with high quality. The last pedestrian shows his back carrying a black backpack, but we can obtain his captured images in front view in rank-3, 6 and 9. We attribute this surprising result to the effects of local features, which establish relationships when some salient parts are lost.

DukeMTMC-reID According to Table 3, our MGN architecture also performs excellently on the challenging DukeMTMC-reID dataset. GP-reid (Almazan et al., 2018) is a good practice of many useful strategies combined in person Re-ID tasks and achieved the best published result. MGN achieves state-of-the-art result of Rank-1/mAP=88.7%/78.4%, outperforming GP-reid by +3.5% in Rank-1 and +5.6% in mAP. Standing on this level of performance on the most challenging datasets currently, we believe there are still some issues to be conquered for further perfect deep Re-ID systems.

CUHK03 As shown in Table 4, our MGN achieves Rank-1/mAP= 68.0%/67.4% on CUHK03 labeled setting and 66.8%/66.0% on CUHK03 detected setting, which outperform all the published results by a large margin. Here we can observe an obvious gap between results of labeled and detected conditions. We argue that it reflects an important affect of detection failure on person Re-ID performance, which emphasizes the importance of high-performance pedestrian detectors.

4. Effectiveness of Components

To verify effectiveness of each component in MGN, we conduct several ablation experiments with different component settings on Market-1501 dataset in single query mode. Notice that other unrelated settings in each comparative experiment are the same as MGN implementation in Section 4.1, and we have carefully tuned all the candidate models and report the best performance with our settings. Table 5 shows the comparison results in different settings related to components of MGN. We separately analyze each component as follows:

MGN vs ResNet-50 Comparing the results by the baseline ResNet-50 model with our MGN model without triplet loss, we can observe MGN makes a significant performance improvement from Rank-1/mAP=87.5%/71.4% to 95.3%/86.2% (+7.8%/14.8%). We also implement the same experiment with ResNet-101 model, which has a similar scale of weights with MGN. The deeper ResNet-101 network indeed brings a considerable performance boost (+2.9%/7.4%), but there is still a large gap with our MGN model, which shows that extra weights from additional branches are not the main contributors of the improvement, but the carefully-designed network architecture. Results above prove that our proposed MGN has incredible capability of feature representations for person Re-ID.

Multiple branch vs Multiple networks The multi-branch setting in one single network is very similar to ensemble of multiple independent networks, but we believe the cooperation of multiple branches can achieve a better performance than ensemble learning. We train three independent ResNet-50 networks separately, each of which respectively replicates the corresponding configuration of three branches, i.e. Global, Part-2 and Part-3 Branch in MGN . In our experiments, we explore the effects of multi-branch architecture in two aspects. On the one hand, from a global view, we compare the performance of MGN with the ensemble of three single networks. The ensemble strategy indeed achieves better performance than any single participating network, but MGN still outperforms about 1%∼2%1\%\sim 2\% on both Rank-1 and mAP. It shows that the cooperation of branches learns more discriminative feature representations than independent networks. On the other hand, from a local view, we respectively compare the performances of features learned by sub-branches of MGN with single networks in corresponding setting of branch. As our expect, the features from sub-branches also perform better than that from single networks. We argue that the mutual effects between sub-branches complement the blind spots in their individual learning procedure.

Multi-branch architecture settings The multi-branch deep network architectures are very common for person Re-ID tasks (Cheng et al., 2016). Our proposed MGN exceeds all the previous architectures by the power of learning representations with multiple granularities. Based on our proposed architecture, we can intuitively infer many variants architectures. On the one hand, based on the MGN model, we can add or reduce the number of local branches. Comparing the model removed the Part-3 branch with that added a Part-4 branch, we find that the removal of Part-3 Branch (MGN w/o Part-3) brings obvious performance degradation by Rank-1/mAP=−-0.9%/0.7%, and the addition of Part-4 Branch (MGN w/ Part-4) introduce no obvious performance boost. This confirms the necessity and efficiency of our proposed 3-branch setting. On the other hand, for the Part-N local branches, the number of partitions is a hyperparameter effecting the granularity of learning representations. We divide the feature maps into 2, 3, 4 in each local branches alternatively, and still find that the proposed Part-2/Part-3 setting is optimal. The Part-2/Part-4 setting skips the granularity level of Part-3, and brings performance degradation by Rank-1/mAP=−-0.9%/1.3%, which shows the importance of Part-3 granularity. Comparing the results of Part-2/Part-4 setting with Part-3/Part-4, the former setting causes greater performance loss than the later. The reason is that the uniform division by 2 and 4 introduces no overlap areas between stripes, but the Part-3/Part-4 setting does. We argue this overlap can introduce the correlation between different partitions, from which more discriminative information can be learned.

Triplet Loss A number of previous works (Li et al., 2014; Cheng et al., 2016; Zhang et al., 2017; Bai et al., 2017) have shown the effectiveness of joint training with softmax loss for classification and triplet loss for metric learning in person Re-ID tasks. In our experiments, we reproduce the boosting effect with both ResNet-50 and MGN models. With the help of triplet loss on all the candidate datasets, we can observe +1.2%/3.6% Rank-1/mAP improvement to the baseline model, and +0.4%/0.7% Rank-1/mAP improvement to MGN model. We can observe two interesting effects from the improvement figures: 1) The improvement on mAP is more obvious than that on rank-1 accuracy, which proves the ranking effects of metric learning losses. 2) Triplet loss brings larger improvement to the baseline model than that to MGN. Comparing to softmax loss, triplet loss helps to capture more detailed information to meet the margin condition (Schroff et al., 2015). Our proposed network architecture is initially designed to enhance the local representation, which dilutes the original effects of triplet loss. Notice that in the MGN without triplet loss setting (MGN w/o TP), we employ no softmax loss on 2048-dim features, and alternatively on the reduced 256-dim features.

Feature response with multiple granularity We insist on our proposed MGN model learns the global and local feature representations with multiple levels of granularities. Figure 5 shows some feature response maps for some input pedestrian images, extracted from all the top of branches in MGN. All the response maps filter most of the complex background, which contain no useful information about identities for pedestrians. The responses from Global Branch is mainly focused on main body parts, and the information on limbs, waists or feet is commonly ignored. On local branches, the global responses on main body are lost, but concentrate more on some particular parts. For example, we can observe preferences on body parts such as shoulders and joints in Part-2 cases. As for the Part-3 cases, responses are more scattered on body parts, but some pivotal semantic information is preferred. Notice that on the Part-3 response map of the last pedestrian images, we can observe a bright circle region in front of the chest of this pedestrian. It is a mark of his T-shirt, which can be regarded as a very discriminative characteristic for his identity.

Conclusion

In this paper, we propose the Multiple Granularity Network (MGN), a novel multi-branch deep network for learning discriminative representations in person re-identification tasks. Each branch in MGN learns global or local representation with certain granularity of body partition. Our method directly learns local features on horizontally-split feature stripes, which is completely end-to-end and introduces no part locating operations such as region proposal or pose estimation. Extensive experiments have indicated that our method not only achieves state-of-the-art results on several mainstream person Re-ID datasets, but also pushes the performance to an exceptional level comparing to existing methods.

References