Deep Learning for Person Re-identification: A Survey and Outlook
Mang Ye, Jianbing Shen, Gaojie Lin, Tao Xiang, Ling Shao, Steven C. H. Hoi
Introduction
Person re-identification (Re-ID) has been widely studied as a specific person retrieval problem across non-overlapping cameras . Given a query person-of-interest, the goal of Re-ID is to determine whether this person has appeared in another place at a distinct time captured by a different camera, or even the same camera at a different time instant . The query person can be represented by an image , a video sequence , and even a text description . Due to the urgent demand of public safety and increasing number of surveillance cameras, person Re-ID is imperative in intelligent surveillance systems with significant research impact and practical importance.
Re-ID is a challenging task due to the presence of different viewpoints , varying low-image resolutions , illumination changes , unconstrained poses , occlusions , heterogeneous modalities , complex camera environments, background clutter , unreliable bounding box generations, etc. These result in varying variations and uncertainty. In addition, for practical model deployment, the dynamic updated camera network , large scale gallery with efficient retrieval , group uncertainty , significant domain shift , unseen testing scenarios , incremental model updating and changing cloths also greatly increase the difficulties. These challenges lead that Re-ID is still unsolved problem. Early research efforts mainly focus on the hand-crafted feature construction with body structures or distance metric learning . With the advancement of deep learning, person Re-ID has achieved inspiring performance on the widely used benchmarks . However, there is still a large gap between the research-oriented scenarios and practical applications . This motivates us to conduct a comprehensive survey, develop a powerful baseline for different Re-ID tasks and discuss several future directions.
Though some surveys have also summarized the deep learning techniques , our survey makes three major differences: 1) We provide an in-depth and comprehensive analysis of existing deep learning methods by discussing their advantages and limitations, analyzing the state-of-the-arts. This provides insights for future algorithm design and new topic exploration. 2) We design a new powerful baseline (AGW: Attention Generalized mean pooling with Weighted triplet loss) and a new evaluation metric (mINP: mean Inverse Negative Penalty) for future developments. AGW achieves state-of-the-art performance on twelve datasets for four different Re-ID tasks. mINP provides a supplement metric to existing CMC/mAP, indicating the cost to find all the correct matches. 3) We make an attempt to discuss several important research directions with under-investigated open issues to narrow the gap between the closed-world and open-world applications, taking a step towards real-world Re-ID system design.
Unless otherwise specified, person Re-ID in this survey refers to the pedestrian retrieval problem across multiple surveillance cameras, from a computer vision perspective. Generally, building a person Re-ID system for a specific scenario requires five main steps (as shown in Fig. 1):
Step 1: Raw Data Collection: Obtaining raw video data from surveillance cameras is the primary requirement of practical video investigation. These cameras are usually located in different places under varying environments . Most likely, this raw data contains a large amount of complex and noisy background clutter.
Step 2: Bounding Box Generation: Extracting the bounding boxes which contain the person images from the raw video data. Generally, it is impossible to manually crop all the person images in large-scale applications. The bounding boxes are usually obtained by the person detection or tracking algorithms .
Step 3: Training Data Annotation: Annotating the cross-camera labels. Training data annotation is usually indispensable for discriminative Re-ID model learning due to the large cross-camera variations. In the existence of large domain shift , we often need to annotate the training data in every new scenario.
Step 4: Model Training: Training a discriminative and robust Re-ID model with the previous annotated person images/videos. This step is the core for developing a Re-ID system and it is also the most widely studied paradigm in the literature. Extensive models have been developed to handle the various challenges, concentrating on feature representation learning , distance metric learning or their combinations.
Step 5: Pedestrian Retrieval: The testing phase conducts the pedestrian retrieval. Given a person-of-interest (query) and a gallery set, we extract the feature representations using the Re-ID model learned in previous stage. A retrieved ranking list is obtained by sorting the calculated query-to-gallery similarity. Some methods have also investigated the ranking optimization to improve the retrieval performance .
According to the five steps mentioned above, we categorize existing Re-ID methods into two main trends: closed-world and open-world settings, as summarized in Table I. A step-by-step comparison is in the following five aspects:
Single-modality vs. Heterogeneous Data: For the raw data collection in Step 1, all the persons are represented by images/videos captured by single-modality visible cameras in the closed-world setting . However, in practical open-world applications, we might also need to process heterogeneous data, which are infrared images , sketches , depth images , or even text descriptions . This motivates the heterogeneous Re-ID in § 3.1.
Bounding Box Generation vs. Raw Images/Videos : For the bounding box generation in Step 2, the closed-world person Re-ID usually performs the training and testing based on the generated bounding boxes, where the bounding boxes mainly contain the person appearance information. In contrast, some practical open-world applications require end-to-end person search from the raw images or videos . This leads to another open-world topic, i.e., end-to-end person search in § 3.2.
Sufficient Annotated Data vs. Unavailable/Limited Labels: For the training data annotation in Step 3, the closed-world person Re-ID usually assumes that we have enough annotated training data for supervised Re-ID model training. However, label annotation for each camera pair in every new environment is time consuming and labor intensive, incurring high costs. In open-world scenarios, we might not have enough annotated data (i.e., limited labels) or even without any label information . This inspires the discussion of the unsupervised and semi-supervised Re-ID in § 3.3.
Correct Annotation vs. Noisy Annotation: For Step 4, existing closed-world person Re-ID systems usually assume that all the annotations are correct, with clean labels. However, annotation noise is usually unavoidable due to annotation error (i.e., label noise) or imperfect detection/tracking results (i.e., sample noise, partial Re-ID ). This leads to the analysis of noise-robust person Re-ID under different noise types in § 3.4.
Query Exists in Gallery vs. Open-set: In the pedestrian retrieval stage (Step 5), most existing closed-world person Re-ID works assume that the query must occur in the gallery set by calculating the CMC and mAP . However, in many scenarios, the query person may not appear in the gallery set , or we need to perform the verification rather than retrieval . This brings us to the open-set person Re-ID in § 3.5.
This survey first introduces the widely studied person Re-ID under closed-world settings in § 2. A detailed review on the datasets and the state-of-the-arts are conducted in § 2.4. We then introduce the open-world person Re-ID in § 3. An outlook for future Re-ID is presented in § 4, including a new evaluation metric (§ 4.1), a new powerful AGW baseline (§ 4.2). We discuss several under-investigated open issues for future study (§ 4.3). Conclusions will be drawn in § 5. A structure overview is shown in the supplementary.
Closed-world Person Re-Identification
This section provides an overview for closed-world person Re-ID. As discussed in § 1, this setting usually has the following assumptions: 1) person appearances are captured by single-modality visible cameras, either by image or video; 2) The persons are represented by bounding boxes, where most of the bounding box area belongs the same identity; 3) The training has enough annotated training data for supervised discriminative Re-ID model learning; 4) The annotations are generally correct; 5) The query person must appear in the gallery set. Typically, a standard closed-world Re-ID system contains three main components: Feature Representation Learning (§ 2.1), which focuses on developing the feature construction strategies; Deep Metric Learning (§ 2.2), which aims at designing the training objectives with different loss functions or sampling strategies; and Ranking Optimization (§ 2.3), which concentrates on optimizing the retrieved ranking list. An overview of the datasets and state-of-the-arts with in-depth analysis is provided in § 2.4.2.
We firstly discuss the feature learning strategies in closed-world person Re-ID. There are four main categories (as shown in Fig. 2): a) Global Feature (§ 2.1.1), it extracts a global feature representation vector for each person image without additional annotation cues ; b) Local Feature (§ 2.1.2), it aggregates part-level local features to formulate a combined representation for each person image ; c) Auxiliary Feature (§ 2.1.3), it improves the feature representation learning using auxiliary information, e.g., attributes , GAN generated images , etc. d) Video Feature (§ 2.1.4), it learns video representation for video-based Re-ID using multiple image frames and temporal information . We also review several specific architecture designs for person Re-ID in § 2.1.5.
Global feature representation learning extracts a global feature vector for each person image, as shown in Fig. 2(a). Since deep neural networks are originally applied in image classification , global feature learning is the primary choice when integrating advanced deep learning techniques into the person Re-ID field in early years.
To capture the fine-grained cues in global feature learning, A joint learning framework consisting of a single-image representation (SIR) and cross-image representation (CIR) is developed in , trained with triplet loss using specific sub-networks. The widely-used ID-discriminative Embedding (IDE) model constructs the training process as a multi-class classification problem by treating each identity as a distinct class. It is now widely used in Re-ID community . Qian et al. develop a multi-scale deep representation learning model to capture discriminative cues at different scales.
Attention Information. Attention schemes have been widely studied in literature to enhance representation learning . 1) Group 1: Attention within the person image. Typical strategies include the pixel level attention and the channel-wise feature response re-weighting , or background suppressing . The spatial information is integrated in . 2) Group 2: attention across multiple person images. A context-aware attentive feature learning method is proposed in , incorporating both an intra-sequence and inter-sequence attention for pair-wise feature alignment and refinement. The attention consistency property is added in . Group similarity is another popular approach to leverage the cross-image attention, which involves multiple images for local and global similarity modeling. The first group mainly enhances the robustness against misalignment/imperfect detection, and the second improves the feature learning by mining the relations across multiple images.
1.2 Local Feature Representation Learning
It learns part/region aggregated features, making it robust against misalignment . The body parts are either automatically generated by human parsing/pose estimation (Group 1) or roughly horizontal division (Group 2).
With automatic body part detection, the popular solution is to combine the full body representation and local part features . Specifically, the multi-channel aggregation , multi-scale context-aware convolutions , multi-stage feature decomposition and bilinear-pooling are designed to improve the local feature learning. Rather than feature level fusion, the part-level similarity combination is also studied in . Another popular solution is to enhance the robustness against background clutter, using the pose-driven matching , pose-guided part attention module , semantically part alignment .
For horizontal-divided region features, multiple part-level classifiers are learned in Part-based Convolutional Baseline (PCB) , which now serves as a strong part feature learning baseline in the current state-of-the-art . To capture the relations across multiple body parts, the Siamese Long Short-Term Memory (LSTM) architecture , second-order non-local attention , Interaction-and-Aggregation (IA) are designed to reinforce the feature learning.
The first group uses human parsing techniques to obtain semantically meaningful body parts, which provides well-align part features. However, they require an additional pose detector and are prone to noisy pose detections . The second group uses a uniform partition to obtain the horizontal stripe parts, which is more flexible, but it is sensitive to heavy occlusions and large background clutter.
1.3 Auxiliary Feature Representation Learning
Auxiliary feature representation learning usually requires additional annotated information (e.g., semantic attributes ) or generated/augmented training samples to reinforce the feature representation .
Semantic Attributes. A joint identity and attribute learning baseline is introduced in . Su et al. propose a deep attribute learning framework by incorporating the predicted semantic attribute information, enhancing the generalizability and robustness of the feature representation in a semi-supervised learning manner. Both the semantic attributes and the attention scheme are incorporated to improve part feature learning . Semantic attributes are also adopted in for video Re-ID feature representation learning. They are also leveraged as the auxiliary supervision information in unsupervised learning .
Viewpoint Information. The viewpoint information is also leveraged to enhance the feature representation learning . Multi-Level Factorisation Net (MLFN) also tries to learn the identity-discriminative and view-invariant feature representations at multiple semantic levels. Liu et al. extract a combination of view-generic and view-specific learning. An angular regularization is incorporated in in the viewpoint-aware feature learning.
Domain Information. A Domain Guided Dropout (DGD) algorithm is designed to adaptively mine the domain-sharable and domain-specific neurons for multi-domain deep feature representation learning. Treating each camera as a distinct domain, Lin et al. propose a multi-camera consistent matching constraint to obtain a globally optimal representation in a deep learning framework. Similarly, the camera view information or the detected camera location is also applied in to improve the feature representation with camera-specific information modeling.
GAN Generation. This section discusses the use of GAN generated images as the auxiliary information. Zheng et al. start the first attempt to apply the GAN technique for person Re-ID. It improves the supervised feature representation learning with the generated person images. Pose constraints are incorporated in to improve the quality of the generated person images, generating the person images with new pose variants. A pose-normalized image generation approach is designed in , which enhances the robustness against pose variations. Camera style information is also integrated in the image generation process to address the cross camera variations. A joint discriminative and generative learning model separately learns the appearance and structure codes to improve the image generation quality. Using the GAN generated images is also a widely used approach in unsupervised domain adaptation Re-ID , approximating the target distribution.
Data Augmentation. For Re-ID, custom operations are random resize, cropping and horizontal flip . Besides, adversarially occluded samples are generated to augment the variation of training data. A similar random erasing strategy is proposed in , adding random noise to the input images. A batch DropBlock randomly drops a region block in the feature map to reinforce the attentive feature learning. Bak et al. generate the virtual humans rendered under different illumination conditions. These methods enrich the supervision with the augmented samples, improving the generalizability on the testing set.
1.4 Video Feature Representation Learning
Video-based Re-ID is another popular topic , where each person is represented by a video sequence with multiple frames. Due to the rich appearance and temporal information, it has gained increasing interest in the Re-ID community. This also brings in additional challenges in video feature representation learning with multiple images.
The primary challenge is to accurately capture the temporal information. A recurrent neural network architecture is designed for video-based person Re-ID , which jointly optimizes the final recurrent layer for temporal information propagation and the temporal pooling layer. A weighted scheme for spatial and temporal streams is developed in . Yan et al. present a progressive/sequential fusion framework to aggregate the frame-level human region representations. Semantic attributes are also adopted in for video Re-ID with feature disentangling and frame re-weighting. Jointly aggregating the frame-level feature and spatio-temporal appearance information is crucial for video representation learning .
Another major challenge is the unavoidable outlier tracking frames within the videos. Informative frames are selected in a joint Spatial and Temporal Attention Pooling Network (ASTPN) , and the contextual information is integrated in . A co-segmentation inspired attention model detects salient features across multiple video frames with mutual consensus estimation. A diversity regularization is employed to mine multiple discriminative body parts in each video sequence. An affine hull is adopted to handle the outlier frames within the video sequence . An interesting work utilizes the multiple video frames to auto-complete occluded regions. These works demonstrate that handling the noisy frames can greatly improve the video representation learning.
It is also challenging to handle the varying lengths of video sequences, Chen et al. divide the long video sequences into multiple short snippets, aggregating the top-ranked snippets to learn a compact embedding. A clip-level learning strategy exploits both spatial and temporal dimensional attention cues to produce a robust clip-level representation. Both the short- and long-term relations are integrated in a self-attention scheme.
1.5 Architecture Design
Framing person Re-ID as a specific pedestrian retrieval problem, most existing works adopt the network architectures designed for image classification as the backbone. Some works have tried to modify the backbone architecture to achieve better Re-ID features. For the widely used ResNet50 backbone , the important modifications include changing the last convolutional stripe/size to 1 , employing adaptive average pooling in the last pooling layer , and adding bottleneck layer with batch normalization after the pooling layer .
Accuracy is the major concern for specific Re-ID network architecture design to improve the accuracy, Li et al. start the first attempt by designing a filter pairing neural network (FPNN), which jointly handles misalignment and occlusions with part discriminative information mining. Wang et al. propose a BraidNet with a specially designed WConv layer and Channel Scaling layer. The WConv layer extracts the difference information of two images to enhance the robustness against misalignments and Channel Scaling layer optimizes the scaling factor of each input channel. A Multi-Level Factorisation Net (MLFN) contains multiple stacked blocks to model various latent factors at a specific level, and the factors are dynamically selected to formulate the final representation. An efficient fully convolutional Siamese network with convolution similarity module is developed to optimize multi-level similarity measurement. The similarity is efficiently captured and optimized by using the depth-wise convolution.
Efficiency is another important factor for Re-ID architecture design. An efficient small scale network, namely Omni-Scale Network (OSNet) , is designed by incorporating the point-wise and depth-wise convolutions. To achieve multi-scale feature learning, a residual block composed of multiple convolutional streams is introduced.
With the increasing interest in auto-machine learning, an Auto-ReID model is proposed. Auto-ReID provides an efficient and effective automated neural architecture design based on a set of basic architecture components, using a part-aware module to capture the discriminative local Re-ID features. This provides a potential research direction in exploring powerful domain-specific architectures.
2 Deep Metric Learning
Metric learning has been extensively studied before the deep learning era by learning a Mahalanobis distance function or projection matrix . The role of metric learning has been replaced by the loss function designs to guide the feature representation learning. We will first review the widely used loss functions in § 2.2.1 and then summarize the training strategies with specific sampling designs § 2.2.2.
This survey only focuses on the loss functions designed for deep learning . An overview of the distance metric learning designed for hand-crafted systems can be found in . There are three widely studied loss functions with their variants in the literature for person Re-ID, including the identity loss, verification loss and triplet loss. An illustration of three loss functions is shown in Fig. 3.
Identity Loss. It treats the training process of person Re-ID as an image classification problem , i.e., each identity is a distinct class. In the testing phase, the output of the pooling layer or embedding layer is adopted as the feature extractor. Given an input image with label , the predicted probability of being recognized as class is encoded with a softmax function, represented by . The identity loss is then computed by the cross-entropy
where represents the number of training samples within each batch. The identity loss has been widely used in existing methods . Generally, it is easy to train and automatically mine the hard samples during the training process, as demonstrated in . Several works have also investigated the softmax variants , such as the sphere loss in and AM softmax in . Another simple yet effective strategy, i.e., label smoothing , is generally integrated into the standard softmax cross-entropy loss. Its basic idea is to avoid the model fitting to over-confident annotated labels, improving the generalizability .
Verification Loss. It optimizes the pairwise relationship, either with a contrastive loss or binary verification loss . The contrastive loss improves the relative pairwise distance comparison, formulated by
where represents the Euclidean distance between the embedding features of two input samples and . is a binary label indicator ( when and belong to the same identity, and , otherwise). is a margin parameter. There are several variants, e.g., the pairwise comparison with ranking SVM in .
Binary verification discriminates the positive and negative of a input image pair. Generally, a differential feature is obtained by , where and are the embedding features of two samples and . The verification network classifies the differential feature into positive or negative. We use to represent the probability of an input pair ( and ) being recognized as (0 or 1). The verification loss with cross-entropy is
The verification is often combined with the identity loss to improve the performance .
Triplet loss. It treats the Re-ID model training process as a retrieval ranking problem. The basic idea is that the distance between the positive pair should be smaller than the negative pair by a pre-defined margin . Typically, a triplet contains one anchor sample , one positive sample with the same identity, and one negative sample from a different identity. The triplet loss with a margin parameter is represented by
where measures the Euclidean distance between two samples. The large proportion of easy triplets will dominate the training process if we directly optimize above loss function, resulting in limited discriminability. To alleviate this issue, various informative triplet mining methods have been designed . The basic idea is to select the informative triplets . Specifically, a moderate positive mining with a weight constraint is introduced in , which directly optimizes the feature difference. Hermans et al. demonstrate that the online hardest positive and negative mining within each training batch is beneficial for discriminative Re-ID model learning. Some methods also studied the point to set similarity strategy for informative triplet mining . This enhances robustness against the outlier samples with a soft hard-mining scheme.
To further enrich the triplet supervision, a quadruplet deep network is developed in , where each quadruplet contains one anchor sample, one positive sample and two mined negative samples. The quadruplets are formulated with a margin-based online hard negative mining. Optimizing the quadruplet relationship results in smaller intra-class variation and larger inter-class variation.
The combination of triplet loss and identity loss is one of the most popular solutions for deep Re-ID model learning . These two components are mutually beneficial for discriminative feature representation learning.
OIM loss. In addition to the above three kinds of loss functions, an Online Instance Matching (OIM) loss is designed with a memory bank scheme. A memory bank contains the stored instance features, where denotes the class number. The OIM loss is then formulated by
where represents the corresponding stored memory feature for class , and is a temperature parameter that controls the similarity space . measures the online instance matching score. The comparison with a memorized feature set of unlabelled identities is further included to calculate the denominator , handling the large instance number of non-targeted identities. This memory scheme is also adopted in unsupervised domain adaptive Re-ID .
2.2 Training strategy
The batch sampling strategy plays an important role in discriminative Re-ID model learning. It is challenging since the number of annotated training images for each identity varies significantly . Meanwhile, the severely imbalanced positive and negative sample pairs increases additional difficulty for the training strategy design .
The most commonly used training strategy for handling the imbalanced issue is identity sampling . For each training batch, a certain number of identities are randomly selected, and then several images are sampled from each selected identity. This batch sampling strategy guarantees the informative positive and negative mining.
To handle the imbalance issue between the positive and negative, adaptive sampling is the popular approach to adjust the contribution of positive and negative samples, such as Sample Rate Learning (SRL) , curriculum sampling . Another approach is sample re-weighting, using the sample distribution or similarity difference to adjust the sample weight. An efficient reference constraint is designed in to transform the pairwise/triplet similarity to a sample-to-reference similarity, addressing the imbalance issue and enhancing the discriminability, which is also robust to outliers.
To adaptively combine multiple loss functions, a multi-loss dynamic training strategy adaptively reweights the identity loss and triplet loss, extracting appropriate component shared between them. This multi-loss training strategy leads to consistent performance gain.
3 Ranking Optimization
Ranking optimization plays a crucial role in improving the retrieval performance in the testing stage. Given an initial ranking list, it optimizes the ranking order, either by automatic gallery-to-gallery similarity mining or human interaction . Rank/Metric fusion is another popular approach for improving the ranking performance with multiple ranking list inputs.
The basic idea of re-ranking is to utilize the gallery-to-gallery similarity to optimize the initial ranking list, as shown in Fig. 4. The top-ranked similarity pulling and bottom-ranked dissimilarity pushing is proposed in . The widely-used -reciprocal reranking mines the contextual information. Similar idea for contextual information modeling is applied in . Bai et al. utilize the geometric structure of the underlying manifold. An expanded cross neighborhood re-ranking method is introduced by integrating the cross neighborhood distance. A local blurring re-ranking employs the clustering structure to improve neighborhood similarity measurement.
Query Adaptive. Considering the query difference, some methods have designed the query adaptive retrieval strategy to replace the uniform searching engine to improve the performance . Andy et al. propose a query adaptive re-ranking method using locality preserving projections. An efficient online local metric adaptation method is presented in , which learns a strictly local metric with mined negative samples for each probe.
Human Interaction. It involves using human feedback to optimize the ranking list . This provides reliable supervision during the re-ranking process. A hybrid human-computer incremental learning model is presented in , which cumulatively learns from human feedback, improving the Re-ID ranking performance on-the-fly.
3.2 Rank Fusion
Rank fusion exploits multiple ranking lists obtained with different methods to improve the retrieval performance . Zheng et al. propose a query adaptive late fusion method on top of a “L” shaped observation to fuse methods. A rank aggregation method by employing the similarity and dissimilarity is developed in . The rank fusion process in person Re-ID is formulated as a consensus-based decision problem with graph theory , mapping the similarity scores obtained by multiple algorithms into a graph with path searching. An Unified Ensemble Diffusion (UED) is recently designed for metric fusion. UED maintains the advantages of three existing fusion algorithms, optimized by a new objective function and derivation. The metric ensemble learning is also studied in .
4 Datasets and Evaluation
Datasets. We first review the widely used datasets for the closed-world setting, including 11 image datasets (VIPeR , iLIDS , GRID , PRID2011 , CUHK01-03 , Market-1501 , DukeMTMC , Airport and MSMT17 ) and 7 video datasets (PRID-2011 , iLIDS-VID , MARS , Duke-Video , Duke-Tracklet , LPW and LS-VID ). The statistics of these datasets are shown in Table II. This survey only focuses on the general large-scale datsets for deep learning methods. A comprehensive summarization of the Re-ID datasets can be found in and their websitehttps://github.com/NEU-Gou/awesome-reid-dataset. Several observations can be made in terms of the dataset collection over recent years:
1) The dataset scale (both #image and #ID) has increased rapidly. Generally, the deep learning approach can benefit from more training samples. This also increases the annotation difficulty needed in closed-world person Re-ID. 2) The camera number is also greatly increased to approximate the large-scale camera network in practical scenarios. This also introduces additional challenges for model generalizability in a dynamically updated network. 3) The bounding boxes generation is usually performed automatically detected/tracked, rather than mannually cropped. This simulates the real-world scenario with tracking/detection errors.
Evaluation Metrics. To evaluate a Re-ID system, Cumulative Matching Characteristics (CMC) and mean Average Precision (mAP) are two widely used measurements.
CMC- (a.k.a, Rank- matching accuracy) represents the probability that a correct match appears in the top- ranked retrieved results. CMC is accurate when only one ground truth exists for each query, since it only considers the first match in evaluation process. However, the gallery set usually contains multiple groundtruths in a large camera network, and CMC cannot completely reflect the discriminability of a model across multiple cameras.
Another metric, i.e., mean Average Precision (mAP) , measures the average retrieval performance with multiple grountruths. It is originally widely used in image retrieval. For Re-ID evaluation, it can address the issue of two systems performing equally well in searching the first ground truth (might be easy match as in Fig. 4), but having different retrieval abilities for other hard matches.
Considering the efficiency and complexity of training a Re-ID model, some recent works also report the FLoating-point Operations Per second (FLOPs) and the network parameter size as the evaluation metrics. These two metrics are crucial when the training/testing device has limited computational resources.
4.2 In-depth Analysis on State-of-The-Arts
We review the state-of-the-arts from both image-based and video-based perspectives. We include methods published in top CV venues over the past three years.
Image-based Re-ID. There are a large number of published papers for image-based Re-IDhttps://paperswithcode.com/task/person-re-identification. We mainly review the works published in 2019 as well as some representative works in 2018. Specifically, we include PCB , MGN , PyrNet , Auto-ReID , ABD-Net , BagTricks , OSNet , DGNet , SCAL , MHN , P2Net , BDB , SONA , SFT , ConsAtt , DenseS , Pyramid , IANet , VAL . We summarize the results on four datasets (Fig. 5). This overview motivates five major insights, as discussed below.
First, with the advancement of deep learning, most of the image-based Re-ID methods have achieved higher rank-1 accuracy than humans (93.5% ) on the widely used Market-1501 dataset. In particular, VAL obtains the best mAP of 91.6% and Rank-1 accuracy of 96.2% on Market-1501 dataset. The major advantage of VAL is the usage of viewpoint information. The performance can be further improved when using re-ranking or metric fusion. The success of deep learning on these closed-world datasets also motivates the shift focus to more challenging scenarios, i.e., large data size or unsupervised learning .
Second, part-level feature learning is beneficial for discriminative Re-ID model learning. Global feature learning directly learns the representation on the whole image without the part constraints . It is discriminative when the person detection/ tracking can accurately locate the human body. When the person images suffer from large background clutter or heavy occlusions, part-level feature learning usually achieves better performance by mining discriminative body regions . Due to its advantage in handling misalignment/occlusions, we observe that most of the state-of-the-art methods developed recently adopt the features aggregation paradigm, combining the part-level and full human body features .
Third, attention is beneficial for discriminative Re-ID model learning. We observe that all the methods (ConsAtt , SCAL , SONA , ABD-Net ) achieving the best performance on each dataset adopt an attention scheme. The attention captures the relationship between different convolutional channels, multiple feature maps, hierarchical layers, different body parts/regions, and even multiple images. Meanwhile, discriminative , diverse , consistent and high-order properties are incorporated to enhance the attentive feature learning. Considering the powerful attention schemes and the specificity of the Re-ID problem, it is highly possible that attentive deeply learned systems will continue dominating the Re-ID community, with more domain specific properties.
Fourth, multi-loss training can improve the Re-ID model learning. Different loss functions optimize the network from a multi-view perspective. Combining multiple loss functions can improve the performance, evidenced by the multi-loss training strategy in the state-of-the-art methods, including ConsAtt , ABD-Net and SONA . In addition, a dynamic multi-loss training strategy is designed in to adaptively integrated two loss functions. The combination of identity loss and triplet loss with hard mining is the primary choice. Moreover, due to the imbalanced issue, sample weighting strategy generally improves the performance by mining informative triplets .
Finally, there is still much room for further improvement due to the increasing size of datasets, complex environment, limited training samples. For example, the Rank-1 accuracy (82.3%) and mAP (60.8%) on the newly released MSMT17 dataset are much lower than that on Market-1501 (Rank-1: 96.2% and mAP 91.7%) and DukeMTMC (Rank-1: 91.6% and mAP 84.5%). On some other challenging datasets with limited training samples (e.g., GRID and VIPeR ), the performance is still very low. In addition, Re-ID models usually suffers significantly on cross-dataset evaluation , and the performance drops dramatically under adversarial attack . We are optimistic that there would be important breakthroughs in person Re-ID, with increasing discriminability, robustness, and generalizability.
Video-based Re-ID. Video-based Re-ID has received less interest, compared to image-based Re-ID. We review the deeply learned Re-ID models, including CoSeg , GLTR , STA , ADFD , STC , DRSA , Snippet , ETAP , DuATM , SDM , TwoS , ASTPN , RQEN , Forest , RNN and IDEX . We also summarize the results on four video Re-ID datasets, as shown in Fig. 6. From these results, the following observations can be drawn.
First, a clear trend of increasing performance can be seen over the years with the development of deep learning techniques. Specifically, the Rank-1 accuracy increases from 70% (RNN in 2016) to 95.5% (GLTR in 2019) on PRID-2011 dataset, and from 58% (RNN ) to 86.3% (ADFD ) on iLIDS-VID dataset. On the large-scale MARS dataset, the Rank-1 accuracy/mAP increase from 68.3%/49.3% (IDEX ) to 88.5%/82.3% (STC ). On the Duke-Video dataset , STA also achieves a Rank-1 accuracy of 96.2%, and the mAP is 94.9%.
Second, spatial and temporal modeling is crucial for discriminative video representation learning. We observe that all the methods (STA , STC , GLTR ) design spatial-temporal aggregation strategies to improve the video Re-ID performance. Similar to image-based Re-ID, the attention scheme across multiple frames also greatly enhances the discriminability. Another interesting observation in demonstrates that utilizing multiple frames within the video sequence can fill in the occluded regions, which provides a possible solution for handling the challenging occlusion problem in the future.
Finally, the performance on these datases has reached a saturation state, usually about less than 1% accuracy gain on these four video datasets. However, there is still large room for improvements on the challenging cases. For example, on the newly collected video dataset, LS-VID , the Rank-1 accuracy/mAP of GLTR are only 63.1%/44.43%, while GLTR can achieve state-of-the-art or at least comparable performance on the other four daatsets. LS-VID contains significantly more identities and video sequences. This provides a challenging benchmark for future breakthroughs in video based Re-ID.
Open-world Person Re-Identification
This section reviews open-world person Re-ID as discussed in § 1, including heterogeneous Re-ID by matching person images across heterogeneous modalities (§ 3.1), end-to-end Re-ID from the raw images/videos (§ 3.2), semi-/unsupervised learning with limited/unavailable annotated labels (§ 3.3), robust Re-ID model learning with noisy annotations (§ 3.4) and open-set person Re-ID when the correct match does not occur in the gallery (§ 3.5).
This subsection summarizes four main kinds of heterogeneous Re-ID, including Re-ID between depth and RGB images (§ 3.1.1), text-to-image Re-ID (§ 3.1.2), visible-to-infrared Re-ID (§ 3.1.3) and cross resolution Re-ID (§ 3.1.4).
Depth images capture the body shape and skeleton information. This provides the possibility for Re-ID under illumination/clothes changing environments, which is also important for personalized human interaction applications.
A recurrent attention-based model is proposed in to address the depth-based person identification. In a reinforcement learning framework, they combine the convolutional and recurrent neural networks to identify small, discriminative local regions of the human body.
Karianakis et al. leverage the large RGB datasets to design a split-rate RGB-to-Depth transfer method, which bridges the gap between the depth images and the RGB images. Their model further incorporates a temporal attention to enhance video representation for depth Re-ID.
Some methods have also studied the combination of RGB and depth information to improve the Re-ID performance, addressing the clothes-changing challenge.
1.2 Text-to-Image Re-ID
Text-to-image Re-ID addresses the matching between a text description and RGB images . It is imperative when the visual image of query person cannot be obtained, and only a text description can be alternatively provided.
A gated neural attention model with recurrent neural network learns the shared features between the text description and the person images. This enables the end-to-end training for text to image pedestrian retrieval. Cheng et al. propose a global discriminative image-language association learning method, capturing the identity discriminative information and local reconstructive image-language association under a reconstruction process. A cross projection learning method also learns a shared space with image-to-text matching. A deep adversarial graph attention convolution network is designed in with graph relation mining. However, the large semantic gap between the text descriptions and the visual images is still challenging. Meanwhile, how to combine the texts and hand-painting sketch image is also worth studying in the future.
1.3 Visible-Infrared Re-ID
Visible-Infrared Re-ID handles the cross-modality matching between the daytime visible and night-time infrared images. It is important in low-lighting conditions, where the images can only be captured by infrared cameras .
Wu et al. start the first attempt to address this issue, by proposing a deep zero-padding framework to adaptively learn the modality sharable features. A two stream network is introduced in to model the modality-sharable and -specific information, addressing the intra- and cross-modality variations simultaneously. Besides the cross-modality shared embedding learning , the classifier-level discrepancy is also investigated in . Recent methods adopt the GAN technique to generate cross-modality person images to reduce the cross-modality discrepancy at both image and feature level. Hierarchical cross-Modality disentanglement factors are modeled in . A dual-attentive aggregation learning method is presented in to capture multi-level relations.
1.4 Cross-Resolution Re-ID
Cross-Resolution Re-ID conducts the matching between low-resolution and high-resolution images, addressing the large resolution variations . A cascaded SR-GAN generates the high-resolution person images in a cascaded manner, incorporating the identity information. Li et al. adopt the adversarial learning technique to obtain resolution-invariant image representations.
2 End-to-End Re-ID
End-to-end Re-ID alleviates the reliance on additional step for bounding boxes generation. It involves the person Re-ID from raw images or videos, and multi-camera tracking.
Re-ID in Raw Images/Videos This task requires that the model jointly performs the person detection and re-identification in a single framework . It is challenging due to the different focuses of two major components.
Zheng et al. present a two-stage framework, and systematically evaluate the benefits and limitations of person detection for the later stage person Re-ID. Xiao et al. design an end-to-end person search system using a single convolutional neural network for joint person detection and re-identification. A Neural Person Search Machine (NPSM) is developed to recursively refine the searching area and locate the target person by fully exploiting the contextual information between the query and the detected candidate region. Similarly, a contextual instance expansion module is learned in a graph learning framework to improve the end-to-end person search. A query-guided end-to-end person search system is developed using the Siamese squeeze-and-excitation network to capture the global context information with query-guided region proposal generation. A localization refinement scheme with discriminative Re-ID feature learning is introduced in to generate more reliable bounding boxes. An Identity DiscriminativE Attention reinforcement Learning (IDEAL) method selects informative regions for auto-generated bounding boxes, improving the Re-ID performance.
Yamaguchi et al. investigate a more challenging problem, i.e., searching for the person from raw videos with text description. A multi-stage method with spatio-temporal person detection and multi-modal retrieval is proposed. Further exploration along this direction is expected.
Multi-camera Tracking End-to-end person Re-ID is also closely related to multi-person, multi-camera tracking . A graph-based formulation to link person hypotheses is proposed for multi-person tracking , where the holistic features of the full human body and body pose layout are combined as the representation for each person. Ristani et al. learn the correlation between the multi-target multi-camera tracking and person Re-ID by hard-identity mining and adaptive weighted triplet learning. Recently, a locality aware appearance metric (LAAM) with both intra- and inter-camera relation modeling is proposed.
3 Semi-supervised and Unsupervised Re-ID
Early unsupervised Re-ID mainly learns invariant components, i.e., dictionary , metric or saliency , which leads to limited discriminability or scalability.
For deeply unsupervised methods, cross-camera label estimation is one the popular approaches . Dynamic graph matching (DGM) formulates the label estimation as a bipartite graph matching problem. To further improve the performance, global camera network constraints are exploited for consistent matching. Liu et al. progressively mine the labels with step-wise metric promotion . A robust anchor embedding method iteratively assigns labels to the unlabelled tracklets to enlarge the anchor video sequences set. With the estimated labels, deep learning can be applied to learn Re-ID models.
For end-to-end unsupervised Re-ID, an iterative clustering and Re-ID model learning is presented in . Similarly, the relations among samples are utilized in a hierarchical clustering framework . Soft multi-label learning mines the soft label information from a reference set for unsupervised learning. A Tracklet Association Unsupervised Deep Learning (TAUDL) framework jointly conducts the within-camera tracklet association and model the cross-camera tracklet correlation. Similarly, an unsupervised camera-aware similarity consistency mining method is also presented in a coarse-to-fine consistency learning scheme. The intra-camera mining and inter-camera association is applied in a graph association framework . The semantic attributes are also adopted in Transferable Joint Attribute-Identity Deep Learning (TJ-AIDL) framework . However, it is still challenging for model updating with newly arriving unlabelled data.
Besides, several methods have also tried to learn a part-level representation based on the observation that it is easier to mine the label information in local parts than that of a whole image. A PatchNet is designed to learn discriminative patch features by mining patch level similarity. A Self-similarity Grouping (SSG) approach iteratively conducts grouping (exploits both the global body and local parts similarity for pseudo labeling) and Re-ID model training in a self-paced manner.
Semi-/Weakly supervised Re-ID. With limited label information, a one-shot metric learning method is proposed in , which incorporates a deep texture representation and a color metric. A stepwise one-shot learning method (EUG) is proposed in for video-based Re-ID, gradually selecting a few candidates from unlabeled tracklets to enrich the labeled tracklet set. A multiple instance attention learning framework uses the video-level labels for representation learning, alleviating the reliance on full annotation.
3.2 Unsupervised Domain Adaptation
Unsupervised domain adaptation (UDA) transfers the knowledge on a labeled source dataset to the unlabeled target dataset . Due to the large domain shift and powerful supervision in source dataset, it is another popular approach for unsupervised Re-ID without target dataset labels.
Target Image Generation. Using GAN generation to transfer the source domain images to target-domain style is a popular approach for UDA Re-ID. With the generated images, this enables supervised Re-ID model learning in the unlabeled target domain. Wei et al. propose a Person Transfer Generative Adversarial Network (PTGAN), transferring the knowledge from one labeled source dataset to the unlabeled target dataset. Preserved self-similarity and domain-dissimilarity is trained with a similarity preserving generative adversarial network (SPGAN). A Hetero-Homogeneous Learning (HHL) method simultaneously considers the camera invariance with homogeneous learning and domain connectedness with heterogeneous learning. An adaptive transfer network decomposes the adaptation process into certain imaging factors, including illumination, resolution, camera view, etc. This strategy improves the cross-dataset performance. Huang et al. try to suppress the background shift to minimize the domain shift problem. Chen et al. design an instance-guided context rendering scheme to transfer the person identities from source domain into diverse contexts in the target domain. Besides, a pose disentanglement scheme is added to improve the image generation . A mutual mean-teacher learning scheme is also developed in . However, the scalability and stability of the image generation for practical large-scale changing environment are still challenging.
Bak et al. generate a synthetic dataset with different illumination conditions to model realistic indoor and outdoor lighting. The synthesized dataset increases generalizability of the learned model and can be easily adapted to a new dataset without additional supervision .
Target Domain Supervision Mining. Some methods directly mine the supervision on the unlabeled target dataset with a well trained model from source dataset. An exemplar memory learning scheme considers three invariant cues as the supervision, including exemplar-invariance, camera invariance and neighborhood-invariance. The Domain-Invariant Mapping Network (DIMN) formulates a meta-learning pipeline for the domain transfer task, and a subset of source domain is sampled at each training episode to update the memory bank, enhancing the scalability and discriminability. The camera view information is also applied in as the supervision signal to reduce the domain gap. A self-training method with progressive augmentation jointly captures the local structure and global data distribution on the target dataset. Recently, a self-paced contrastive learning framework with hybrid memory is developed with great success, which dynamically generates multi-level supervision signals.
The spatio-temporal information is also utilized as the supervision in TFusion . TFusion transfers the spatio-temporal patterns learned in the source domain to the target domain with a Bayesian fusion model. Similarly, Query-Adaptive Convolution (QAConv) is developed to improve cross-dataset accuracy.
3.3 State-of-The-Arts for Unsupervised Re-ID
Unsupervised Re-ID has achieved increasing attention in recent years, evidenced by the increasing number of publications in top venues. We review the SOTA for unsupervised deeply learned methods on two widely-used image-based Re-ID datasets. The results are summarized in Table III. From these results, the following insights can be drawn.
First, the unsupervised Re-ID performance has increased significantly over the years. The Rank-1 accuracy/mAP increases from 54.5%/26.3% (CAMEL ) to 90.3%/76.7% (SpCL ) on the Market-1501 dataset within three years. The performance for DukeMTMC dataset increases from 30.0%/16.4% to 82.9%/68.8%. The gap between the supervised upper bound and the unsupervised learning is narrowed significantly. This demonstrates the success of unsupervised Re-ID with deep learning.
Second, current unsupervised Re-ID is still under-developed and it can be further improved in the following aspects: 1) The powerful attention scheme in supervised Re-ID methods has rarely been applied in unsupervised Re-ID. 2) Target domain image generation has been proved effective in some methods, but they are not applied in two best methods (PAST , SSG ). 3) Using the annotated source data in the training process of the target domain is beneficial for cross-dataset learning, but it is also not included in above two methods. These observations provide the potential basis for further improvements.
Third, there is still a large gap between the unsupervised and supervised Re-ID. For example, the rank-1 accuracy of supervised ConsAtt has achieved 96.1% on the Market-1501 dataset, while the highest accuracy of unsupervised SpCL is about 90.3%. Recently, He et al. have demonstrated that unsupervised learning with large-scale unlabeled training data has the ability to outperform the supervised learning on various tasks . We expect that several breakthroughs in future unsupervised Re-ID.
4 Noise-Robust Re-ID
Re-ID usually suffers from unavoidable noise due to data collection and annotation difficulty. We review noise-robust Re-ID from three aspects: Partial Re-ID with heavy occlusion, Re-ID with sample noise caused by detection or tracking errors, and Re-ID with label noise caused by annotation error.
Partial Re-ID. This addresses the Re-ID problem with heavy occlusions, i.e., only part of the human body is visible . A fully convolutional network is adopted to generate fix-sized spatial feature maps for the incomplete person images. Deep Spatial feature Reconstruction (DSR) is further incorporated to avoid explicit alignment by exploiting the reconstructing error. Sun et al. design a Visibility-aware Part Model (VPM) to extract sharable region-level features, thus suppressing the spatial misalignment in the incomplete images. A foreground-aware pyramid reconstruction scheme also tries to learn from the unoccluded regions. The Pose-Guided Feature Alignment (PGFA) exploits the pose landmarks to mine discriminative part information from occlusion noise. However, it is still challenging due to the severe partial misalignment, unpredictable visible regions and distracting unshared body regions. Meanwhile, how to adaptively adjust the matching model for different queries still needs further investigation.
Re-ID with Sample Noise. This refers to the problem of the person images or the video sequence containing outlying regions/frames, either caused by poor detection/inaccurate tracking results. To handle the outlying regions or background clutter within the person image, pose estimation cues or attention cues are exploited. The basic idea is to suppress the contribution of the noisy regions in the final holistic representation. For video sequences, set-level feature learning or frame level re-weighting are the commonly used approaches to reduce the impact of noisy frames. Hou et al. also utilize multiple video frames to auto-complete occluded regions. It is expected that more domain-specific sample noise handling designs in the future.
Re-ID with Label Noise. Label noise is usually unavoidable due to annotation error. Zheng et al. adopt a label smoothing technique to avoid label overfiting issues . A Distribution Net (DNet) that models the feature uncertainty is proposed in for robust Re-ID model learning against label noise, reducing the impact of samples with high feature uncertainty. Different from the general classification problem, robust Re-ID model learning suffers from limited training samples for each identity . In addition, the unknown new identities increase additional difficulty for the robust Re-ID model learning.
5 Open-set Re-ID and Beyond
Open-set Re-ID is usually formulated as a person verification problem, i.e., discriminating whether or not two person images belong to the same identity . The verification usually requires a learned condition , i.e., . Early researches design hand-crafted systems . For deep learning methods, an Adversarial PersonNet (APN) is proposed in , which jointly learns a GAN module and the Re-ID feature extractor. The basic idea of this GAN is to generate realistic target-like images (imposters) and enforce the feature extractor is robust to the generated image attack. Modeling feature uncertainty is also investigated in . However, it remains quite challenging to achieve a high true target recognition and maintain low false target recognition rate .
Group Re-ID. It aims at associating the persons in groups rather than individuals . Early researches mainly focus on group representation extraction with sparse dictionary learning or covariance descriptor aggregation . The multi-grain information is integrated in to fully capture the characteristics of a group. Recently, the graph convoltuional network is applied in , representing the group as a graph. The group similarity is also applied in the end-to-end person search and the individual re-identification to improve the accuracy. However, group Re-ID is still challenging since the group variation is more complicated than the individuals.
Dynamic Multi-Camera Network. Dynamic updated multi-camera network is another challenging issue , which needs model adaptation for new cameras or probes. A human in-the-loop incremental learning method is introduced in to update the Re-ID model, adapting the representation for different probe galleries. Early research also applies the active learning for continuous Re-ID in multi-camera network. A continuous adaptation method based on sparse non-redundant representative selection is introduced in . A transitive inference algorithm is designed to exploit the best source camera model based on a geodesic flow kernel. Multiple environmental constraints (e.g., Camera Topology) in dense crowds and social relationships are integrated for an open-world person Re-ID system . The model adaptation and environmental factors of cameras are crucial in practical dynamic multi-camera network. Moreover, how to apply the deep learning technique for the dynamic multi-camera network is still less investigated.
An Outlook: Re-ID in Next Era
This section firstly presents a new evaluation metric in § 4.1, a strong baseline (in § 4.2) for person Re-ID. It provides an important guidance for future Re-ID research. Finally, we discuss some under-investigated open issues in § 4.3.
For a good Re-ID system, the target person should be retrieved as accurately as possible, i.e., all the correct matches should have low rank values. Considering that the target person should not be neglected in the top-ranked retrieved list, especially for multi-camera network, so as to accurately track the target. When the target person appears in the gallery set at multiple time stamps, the rank position of the hardest correct match determines the workload of the inspectors for further investigation. However, the currently widely used CMC and mAP metrics cannot evaluate this property, as shown in Fig. 7. With the same CMC, rank list 1 achieves a better AP than rank list 2, but it requires more efforts to find all the correct matches. To address this issue, we design a computationally efficient metric, namely a negative penalty (NP), which measures the penalty to find the hardest correct match
where indicates the rank position of the hardest match, and represents the total number of correct matches for query . Naturally, a smaller NP represents better performance. For consistency with CMC and mAP, we prefer to use the inverse negative penalty (INP), an inverse operation of NP. Overall, the mean INP of all the queries is represented by
The calculation of mINP is quite efficient and can be seamlessly integrated in the CMC/mAP calculating process. mINP avoids the domination of easy matches in the mAP/CMC evaluation. One limitation is that mINP value difference for large gallery size would be much smaller compared to small galleries. But it still can reflect the relative performance of a Re-ID model, providing a supplement to the widely-used CMC and mAP metrics.
2 A New Baseline for Single-/Cross-Modality Re-ID
According to the discussion in § 2.4.2, we design a new AGWDetails are in https://github.com/mangye16/ReID-Survey and comprehensive comparison is shown in the supplementary material. baseline for person Re-ID, which achieves competitive performance on both single-modality (image and video) and cross-modality Re-ID tasks. Specifically, our new baseline is designed on top of BagTricks , and AGW contains the following three major improved components:
(1) Non-local Attention (Att) Block. As discussed in § 2.4.2, the attention scheme plays a crucial role in discriminative Re-ID model learning. We adopt the powerful non-local attention block to obtain the weighted sum of the features at all positions, represented by
where is a weight matrix to be learned, represents a non-local operation, and formulates a residual learning strategy. Details can be found in . We adopt the default setting from to insert the non-local attention block.
(2) Generalized-mean (GeM) Pooling. As a fine-grained instance retrieval, the widely-used max-pooling or average pooling cannot capture the domain-specific discriminative features. We adopt a learnable pooling layer, named generalized-mean (GeM) pooling , formulated by
where represents the feature map, and is number of feature maps in the last layer. is the set of activations for feature map . is a pooling hyper-parameter, which is learned in the back-propagation process . The above operation approximates max pooling when and average pooling when .
(3) Weighted Regularization Triplet (WRT) loss. In addition to the baseline identity loss with softmax cross-entropy, we integrate with another weighted regularized triplet loss,
where represents a hard triplet within each training batch. For anchor , is the corresponding positive set, and is the negative set. / represents the pairwise distance of a positive/negative sample pair. The above weighted regularization inherits the advantage of relative distance optimization between positive and negative pairs, but it avoids introducing any additional margin parameters. Our weighting strategy is similar to , but our solution does not introduce additional hyper-parameters.
The overall framework of AGW is shown in Fig 8. Other components are exactly the same as . In the testing phase, the output of BN layer is adopted as the feature representation for Re-ID. The implementation details and more experimental results are in the supplementary material.
Results on Single-modality Image Re-ID. We first evaluate each component on two image-based datasets (Market-1501 and DukeMTMC) in Table IV. We also list two state-of-the-art methods, BagTricks and ABD-Net . We report the results on CUHK03 and MSMT17 datasets in Table V. We obtain the following two observations:
1) All the components consistently contribute the accuracy gain, and AGW performs much better than the original BagTricks under various metrics. AGW provides a strong baseline for future improvements. We have also tried to incorporate part-level feature learning , but extensive experiments show that it does not improve the performance. How to aggregate part-level feature learning with AGW needs further study in the future. 2) Compared to the current state-of-the-art, ABD-Net , AGW performs favorably in most cases. In particular, we achieve much higher mINP on DukeMTMC dataset, 45.7% vs. 42.1%. This demonstrates that AGW requires less effort to find all the correct matches, verifying the ability of mINP.
Results on Single-modality Video Re-ID. We also evaluate the proposed AGW on four widely used single modality video-based datasets ( MARS , DukeVideo , PRID2011 and iLIDS-VID , as shown in Table VI. We also compare two state-of-the-art methods, BagTricks and Co-Seg . For video data, we develop a variant (AGW+) to capture the temporal information with frame-level average pooling for sequence representation. Meanwhile, constraint random sampling strategy is applied for training. Compared to Co-Seg , our AGW+ obtains better Rank-1, mAP and mINP in most cases.
Results on Partial Re-ID. We also test the performance of AGW on two partial Re-ID datasets, as shown in Table VII. The experimental setting are from DSR . We also achieve comparable performance with the state-of-the-art VPM method . This experiment further demonstrates the superiority of AGW for the open-world partial Re-ID task. Meanwhile, the mINP also shows the applicability for this open-world Re-ID problem.
Results on Cross-modality Re-ID. We also test the performance of AGW using a two-stream architecture on the cross-modality visible-infrared Re-ID task. The comparison with the current state-of-the-arts on two datasets is shown in Table VIII. We follow the settings in AlignG to perform the experiments. Results show that AGW achieves much higher accuracy than existing cross-modality Re-ID models, verifying the effectiveness for the open-world Re-ID task.
3 Under-Investigated Open Issues
We discuss the open-issues from five different aspects according to the five steps in §1, including uncontrollable data collection, human annotation minimization, domain-specific/generalizable architecture design, dynamic model updating and efficient model deployment.
Most existing Re-ID works evaluate their method on a well-defined data collection environment. However, the data collection in real complex environment is uncontrollable. The data might be captured from unpredictable modality, modality combinations, or even cloth changing data .
Multi-Heterogeneous Data. In real applications, the Re-ID data might be captured from multiple heterogeneous modalities, i.e., the resolutions of person images vary a lot , both the query and gallery sets may contain different modalities (visible, thermal , depth or text description ). This results in a challenging multiple heterogeneous person Re-ID. A good person Re-ID system would be able to automatically handle the changing resolutions, different modalities, various environments and multiple domains. Future work with broad generalizability is expected, evaluating their method for different Re-ID tasks.
Cloth-Changing Data. In practical surveillance system, it is very likely to contain a large number of target persons with changing clothes. A cloth-Clothing Change Aware Network (CCAN) addresses this issue by separately extracting the face and body context representation, and similar idea is applied in . Yang et al. present a spatial polar transformation (SPT) to learn cross-cloth invariant representation. However, they still rely heavily on the face and body appearance, which might be unavailable and unstable in real scenarios. It would be interesting to further explore the possibility of other discriminative cues (e.g., gait, shape) to address the cloth-changing issue.
3.2 Human Annotation Minimization
Besides the unsupervised learning, active learning or human interaction provides another possible solution to alleviate the reliance on human annotation.
Active Learning. Incorporating human interaction, labels are easily provided for newly arriving data and the model can be subsequently updated . A pairwise subset selection framework minimizes human labeling effort by firstly constructing an edge-weighted complete -partite graph and then solving it as a triangle free subgraph maximization problem. Along this line, a deep reinforcement active learning method iteratively refines the learning policy and trains a Re-ID network with human-in-the-loop supervision. For video data, an interpretable reinforcement learning method with sequential decision making is designed. The active learning is crucial in practical Re-ID system design, but it has received less attention in the research community. Additionally, the newly arriving identities is extremely challenging, even for human. Efficient human in-the-loop active learning is expected in the future.
Learning for Virtual Data. This provides an alternative for minimizing the human annotation. A synthetic dataset is collected in for training, and they achieve competitive performance on real-world datasets when trained on this synthesized dataset. Bak et al. generate a new synthetic dataset with different illumination conditions to model realistic indoor and outdoor lighting. A large-scale synthetic PersonX dataset is collected in to systematically study the effect of viewpoint for a person Re-ID system. Recently, the 3D person images are also studied in , generating the 3D body structure from 2D images. However, how to bridge the gap between synthesized images and real-world datasets remains challenging.
3.3 Domain-Specific/Generalizable Architecture Design
Re-ID Specific Architecture. Existing Re-ID methods usually adopt architectures designed for image classification as the backbone. Some methods modify the architecture to achieve better Re-ID features . Very recently, researchers have started to design domain specific architectures, e.g., OSNet with omni-scale feature learning . It detects the small-scale discriminative features at a certain scale. OSNet is extremely lightweight and achieves competitive performance. With the advancement of automatic neural architecture search (e.g., Auto-ReID ), more domain-specific powerful architectures are expected to address the task-specific Re-ID challenges. Limited training samples in Re-ID also increase the difficulty in architecture design.
Domain Generalizable Re-ID. It is well recognized that there is a large domain gap between different datsets . Most existing methods adopt domain adaptation for cross-dataset training. A more practical solution would be learning a domain generalized model with a number of source datasets, such that the learned model can be generalized to new unseen datasets for discriminative Re-ID without additional training . Hu et al. studied the cross-dataset person Re-ID by introducing a part-level CNN framework. The Domain-Invariant Mapping Network (DIMN) designs a meta-learning pipeline for domain generalizable Re-ID, learning a mapping between a person image and its identity classifier. The domain generalizability is crucial to deploy the learned Re-ID model under an unknown scenario.
3.4 Dynamic Model Updating
Fixed model is inappropriate for practical dynamically updated surveillance system. To alleviate this issue, dynamic model updating is imperative, either to a new domain/camera or adaptation with newly collected data.
Model Adaptation to New Domain/Camera. Model adaptation to a new domain has been widely studied in the literature as a domain adaptation problem . In practical dynamic camera network, a new camera may be temporarily inserted into an existing surveillance system. Model adaptation is crucial for continuous identification in a multi-camera network . To adapt a learned model to a new camera, a transitive inference algorithm is designed to exploit the best source camera model based on a geodesic flow kernel. However, it is still challenging when the newly collected data by the new camera has totally different distributions. In addition, the privacy and efficiency issue also need further consideration.
Model Updating with Newly Arriving Data. With the newly collected data, it is impractical to training the previously learned model from the scratch . An incremental learning approach together with human interaction is designed in . For deeply learned model, an addition using covariance loss is integrated in the overall learning function. However, this problem is not well studied since the deep model training require large amount of training data. Besides, the unknown new identities in the newly arriving data is hard to be identified for the model updating.
3.5 Efficient Model Deployment
It is important to design efficient and adaptive models to address scalability issue for practical model deployment.
Fast Re-ID. For fast retrieval, hashing has been extensively studied to boost the searching speed, approximating the nearest neighbor search . Cross-camera Semantic Binary Transformation (CSBT) transforms the original high-dimensional feature representations into compact low-dimensional identity-preserving binary codes. A Coarse-to-Fine (CtF) hashing code search strategy is developed in , complementarily using short and long codes. However, the domain-specific hashing still needs further study.
Lightweight Model. Another direction for addressing the scalability issue is to design a lightweight Re-ID model. Modifying the network architecture to achieve a lightweight model is investigated in . Model distillation is another approach, e.g., a multi-teacher adaptive similarity distillation framework is proposed in , which learns a user-specified lightweight student model from multiple teacher models, without access to source domain data.
Resource Aware Re-ID. Adaptively adjusting the model according to the hardware configurations also provides a solution to handle the scalability issue. Deep Anytime Re-ID (DaRe) employs a simple distance based routing strategy to adaptively adjust the model, fitting to hardware devices with different computational resources.
Concluding Remarks
This paper presents a comprehensive survey with in-depth analysis from a both closed-world and open-world perspectives. We first introduce the widely studied person Re-ID under the closed-world setting from three aspects: feature representation learning, deep metric learning and ranking optimization. With powerful deep learning, the closed-world person Re-ID has achieved performance saturation on several datasets. Correspondingly, the open-world setting has recently gained increasing attention, with efforts to address various practical challenges. We also design a new AGW baseline, which achieves competitive performance on four Re-ID tasks under various metrics. It provides a strong baseline for future improvements. This survey also introduces a new evaluation metric to measure the cost of finding all the correct matches. We believe this survey will provide important guidance for future Re-ID research.
References
A. Experiments on Single-modality Image-based Re-ID
Architecture Design. The overall structurehttps://github.com/mangye16/ReID-Survey of our proposed AGW baseline for single-modality Re-ID is illustrated in § 4 (Fig. R1). We adopt ResNet50 pre-trained on ImageNet as our backbone network and change the dimension of the fully connected layer to be consistent with the number of identities in the training dataset. The stride of the last spatial down-sampling operation in the backbone network is changed from 2 to 1. Consequently, the spatial size of the output feature map is changed from to , when feeding an image of resolution as input. In our method, we replace the Global Average Pooling in the original ResNet50 with the Generalized-mean (GeM) pooling. The pooling hyper parameter for generalized-mean pooling is initialized as 3.0. A BatchNorm layer, named BNNeck is plugged between the GeM pooling layer and the fully connected layer. The output of the GeM pooling layer is adopted for computing center loss and triplet loss in the training stage, while the feature after BNNeck is used for computing distance between pedestrian images during testing inference stage.
Non-local Attention. The ResNet contains 4 residual stages, i.e. , , and , each containing stacks of bottleneck residual blocks. We inserted five non-local blocks after , , , and respectively. We adopt the Dot Product version of non-local block with a bottleneck of 512 channels in our experiment. For each non-local block, a BatchNorm layer is added right after the last linear layer that represents . The affine parameter of this BatchNorm layer is initialized as zeros to ensure that the non-local block can be inserted into any pre-trained networks while maintaining its initial behavior.
Training Strategy. In the training stage, we randomly sample 16 identities and 4 images for each identity to form a mini-batch of size 64. Each image is resized into pixels, padding 10 pixels with zero values, and then randomly cropped into pixels. Random horizontally flipping and random erasing with 0.5 probability respectively are also adopted for data augmentation. Specifically, random erasing augmentation randomly selects a rectangle region with area ratio to the whole image, and erase its pixels with the mean value of the image. Besides, the aspect ratio of this region is randomly initialized between and . In our method, we set the above hyper-parameter as , and . At last, we normalize the RGB channels of each image with mean 0.485, 0.456, 0.406 and stand deviation 0.229, 0.224, 0.225, respectively, which are the same with settings in .
Training Loss. In the training stage, three types of loss are combined for optimization, including identity classification loss (), center loss () and our proposed weighted regularization triplet loss ().
The balanced weight of the center loss () is set to 0.0005 and the one () of the weighted regularized triplet loss is set to 1.0. Label smoothing is adopted to improve the original identity classification loss, which encourages the model to be less confident during training and prevent overfitting for classification task. Concretely, it changes the one-hot label as follow:
where is the total number of identities, is a small constant to reduce the confidence for the true identity label and is treated as a new classification target for training. In our method, we set to be 0.1.
Optimizer Setting. Adam optimizer with a weight decay is adopted to train our model. The initial learning rate is set as 0.00035 and is decreased by 0.1 at the 40th epoch and 70th epoch, respectively. The model is trained for 120 epochs in total. Besides, a warm-up learning rate scheme is also employed to improve the stability of training process and bootstrap the network for better performance. Specifically, in the first 10 epochs, the learning rate is linearly increased from to . The learning rate at epoch can be computed as:
B. Experiments on Video-based Re-ID
Implementation Details. We extend our proposed AGW baseline to a video-based Re-ID model by several minor changes to the backbone structure and training strategy of single-modality image-based Re-ID model. The video-based AGW baseline takes a video sequence as input and extracts the frame-level feature vectors, which are then averaged to be a video-level feature vector before the BNNeck layer. Besides, the video-based AGW baseline is trained for 400 epochs totally to better fit the video person Re-ID datasets. The learning rate is decayed by 10 times every 100 epochs. To form an input video sequence, we adopt the constraint random sampling strategy to sample 4 frames as a summary for the original pedestrian tracklet. The BagTricks baseline is extended to a video-based Re-ID model in the same way as AGW baseline for fair comparison. In addition, we also develop a variant of AGW baseline, termed as AGW+, to model more abundant temporal information in a pedestrian tracklet. AGW+ baseline adopts the dense sampling strategy to form an input video sequence in the testing stage. Dense sampling strategy takes all the frames in a pedestrian tracklet to form input video sequence, resulting better performance but higher computational cost. To further improve the performance of AGW+ baseline on video re-ID datasets, we also remove the warm-up learning rate strategy and add dropout operation before the linear classification layer.
Detailed Comparison. In this section, we conduct the performance comparison between AGW baseline and other state-of-the-art video-based person Re-ID methods, including ETAP , DRSA , STA Snippet , VRSTC , ADFD , GLTR and CoSeg . The comparison results on four video person Re-ID datasets (MARS, DukeVideo, PRID2011 and iLIDS-VID) are listed in Table R1. As we can see, by simply taking video sequence as input and adopting average pooling to aggregate frame-level feature, our AGW baseline achieves competitive results on two large-scale video Re-ID dataset, MARS and DukeVideo. Besides, AGW baseline also performs significantly better than BagTricks baseline under multiple evaluation metrics. By further modeling more temporal information and adjusting training strategy, AGW+ baseline gains huge improvement and also achieves competitive results on both PRID2011 and iLIDS-VID datasets. AGW+ baseline outperforms most state-of-the-art methods on MARS, DukeVideo and PRID2011 datasets. Most of these video-based person Re-ID methods achieve state-of-the-art performance by designing complicate temporal attention mechanism to exploit temporal dependency in pedestrian video. We believe that our AGW baseline can help video Re-ID model achieve higher performance with properly designed mechanism to further exploit spatial and temporal dependency.
C. Experiments on Cross-modality Re-ID
Architecture Design. We adopt a two-stream network structure as the backbone for cross-modality visible-infrared Re-IDhttps://github.com/mangye16/Cross-Modal-Re-ID-baseline. Compared to the one-stream architecture in single-modality person Re-ID (Fig. 8), the major difference is that, i.e., the first block is specific for two modalities in order to capture modality-specific information, while the remaining blocks are shared to learn modality sharable features. Compared to the two-stream structure widely used in , which only has one shared embedding layer, our design captures more sharable components. An illustration for cross-modality visible-infrared Re-ID is shown in Fig. R2.
Training Strategy. At each training step, we random sample 8 identities from the whole dataset. Then 4 visible and 4 infrared images are randomly selected for each identity. Totally, each training batch contains 32 visible and 32 infrared images. This guarantees the informative hard triplet mining from both modalities, i.e., we directly select the hard positive and negative from both intra- and inter-modalities. This approximates the idea of bi-directional center-constrained top-ranking loss, handling the inter- and intra-modality variations simultaneously.
For fair comparison, we follow the settings in exactly to conduct the image processing and data augmentation. For infrared images, we keep the original three channels, just like the visible RGB images. All the input images from both modalities are first resized to , and random crop with zero-padding together with random horizontal flipping are adopted for data argumentation. The cropped image sizes are for both modality. The image normalization are exactly following the single-modality setting.
Training Loss. In the training stage, we combine with the identity classification loss () and our proposed weighted regularization triplet loss (). The weight of combining the identity loss and weighted regularized triplet loss is set to 1, the same as the single-modality setting. The pooling parameter is set to 3. For stable training, we adopt the same identity classifier for two heterogeneous modalities, mining sharable information.
Optimizer Setting. We set the initial learning rate as 0.1 on both datasets, and decay it by 0.1 and 0.01 at 20 and 50 epochs, respectively. The total number of training epoch is 60. We also adopt a warm-up learning rate scheme. We adopt the stochastic gradient descent (SGD) optimizer for optimization, and the momentum parameter is set to 0.9. We have tried the same Adam optimizer (used in single-modality Re-ID) on cross-modality Re-ID task, but the performance is much lower than that of SGD optimizer by using a large learning rate. This is crucial since ImageNet initialization is adopted for the infrared images.
Detailed Comparison This section conducts the comparison with the state-of-the-art cross-modality VI-ReID methods, including eBDTR , HSME , D2RL , MAC , MSR and AlignGAN . These methods are published in the past two years. AlignGAN , published in ICCV 2019, achieves the state-of-the-art performance by aligning the cross-modality representation at both the feature level and pixel level with GAN generated images. The results on two datasets are shown in Tables R2 and R3. We observe that the proposed AGW consistently outperforms the current state-of-the-art, without the time-consuming image generation process. For different query settings on RegDB dataset, our proposed baseline generally keeps the same performance. Our proposed baseline has been widely used in many recently developed methods. We believe our new baseline will provide a good guidance to boost the cross-modality Re-ID.
D. Experiments on Partial Re-ID
Implementation Details. We also evaluate the performance of our proposed AGW baseline on two commonly-used partial Re-ID datasets, Partial-REID and Partial-iLIDS. The overall backbone structure and training strategy for partial Re-ID AGW baseline model are the same as the one for single-modality image-based Re-ID model. Both Partial-REID and Partial-iLIDS datasets offer only query image set and gallery image set. So we train AGW baseline model on the training set of Market-1501 dataset, then evaluate its performance on the testing set of two partial Re-ID datasets. We adopt the same way to evaluate the performance of BagTricks baseline on these two partial Re-ID datasets for better comparison and analysis.
Detailed Comparison. We compare the performance of AGW baseline with other state-of-the-art partial Re-ID methods, including DSR , SFR and VPM . All these methods are published in recent years. The comparison results on both Partial-REID and Partial-iLIDS datasets are shown in Table R4. The VPM achieves a very high performance by perceiving the visibility of regions through self-supervision and extracting region-level features. Considering only global features, our proposed AGW baseline still achieves competitive results compared to the current state-of-the-arts on both datasets. Besides, AGW baseline brings significant improvement comparing to BagTricks under multiple evaluation metrics, demonstrating its effectiveness for partial Re-ID problem.
E. Overview of This Survey
The overview figure of this survey is shown in Fig. R3. According to the five steps in developing a person Re-ID system, we conduct the survey from both closed-world and open-world settings. The closed-world setting is detailed in three different aspects: feature representation learning, deep metric learning and ranking optimization. We then summarize the datasets and SOTAs from both image- and video-based perspectives. For open-world person Re-ID, we summarize it into five aspects: including heterogeneous data, Re-ID from raw images/videos, unavailable/limited labels, noisy annotation and open-set Re-ID.
Following the summary, we present an outlook for future person Re-ID. We design a new evaluation metric (mINP) to evaluate the difficulty to find all the correct matches. By analyzing the advantages of existing Re-ID methods, we develop a strong AGW baseline for future developments, which achieves competitive performance on four Re-ID tasks. Finally, some under-investigated open issues are discussed. Our survey provides a comprehensive summarization of existing state-of-the-art in different sub-tasks. Meanwhile, the analysis of future directions is also presented for further development guidance.
Acknowledgement. The authors would like to thank the anonymous reviewers for providing valuable feedbacks to improve the quality of this survey. The authors also would like to thank the pioneer researchers in person re-identification and other related fields. This work is sponsored by CAAI-Huawei MindSpore Open Fund.