Person Search via A Mask-Guided Two-Stream CNN Model

Di Chen, Shanshan Zhang, Wanli Ouyang, Jian Yang, Ying Tai

Introduction

The task of person search is first introduced by , which unifies the pedestrian detection and person re-identification in a coherent system. A typical person re-identification method aims to find matchings between the query probe and the cropped person image patches from the gallery, thus requiring perfect person detection results, which are hard to obtain in practice. In contrast, person search, which searches the queried person over the whole image instead of comparing with manually cropped person image locally, is closer to the real world applications. However, considering the tasks of detection and re-ID together brings domain-specific difficulties: large appearance variance across cameras, low resolution, occlusion, etc. In addition, sharing features between detection and re-ID also accumulates errors from each of them, e.g. false alarms, misalignments and inexpressive person descriptors, which further jeopardizes the final person search performance.

Following , a few other works have also been proposed for person search. Most of them focus on an end-to-end solution based on Faster R-CNN . Specifically, an axillary fully-connected (FC) layer is added upon the top convolutional layer of Faster R-CNN to extract discriminative features for re-identification. During training, they optimize a joint loss which is composed of the Faster R-CNN losses and a person categorization loss. However, we argue that it is not appropriate to share representations between the detection and re-ID tasks, as their goals contradict with each other. For the detection task, all people are treated as one class, and the goal is to distinguish them from the background, thus the representations focus on the commonness of different people, e.g. the body shape; while for the re-ID task, different people are deemed as different categories, and the goal is to maximize the differences in between, thus the representations focus on the characteristics of each identity, e.g. clothing, hairstyle, etc. In other words, the detection and re-ID tasks aim to model the inter-class and intra-class difference for people respectively. Therefore, it makes more sense to separate the two tasks rather than solving them jointly.

In the community of person re-ID, it is widely believed that discriminative information lies in foreground, while background is one of the detrimental factors and ought to be neglected or removed during feature extraction . An intuitive idea would be to extract features on the foreground person patch only while ignoring the background area. However, simply abandoning all background information may harm the re-ID performance from two aspects. Firstly, the feature extraction procedure may gather errors from imperfect or noisy segmentation masks, i.e. identification information loss caused by fractional body shape. Secondly, background information sometimes acts as useful context, e.g. attendant suitcases, handbags or companions. Casting out all background area would neglect some informative cues for the re-ID problem. Therefore, we argue that it is more suitable to consider a compromised strategy of paying extra attention on the foreground person while also using the background as a complementary cue.

Inspired by the above discussions, we propose a new approach for person search. It consists of two stages: pedestrian detection and person re-identification. We solve them separately, without sharing any representations. Furthermore, we propose a Two-stream CNN to model foreground person and original image independently, which aims to extract more informative features for each identity and still consider the complementarity of the background. The whole framework is demonstrated in Fig. 1, and we will talk about more details in Sec. 3.

In summary, our contributions are three-folds:

To the best of our knowledge, this paper is the first work showing that for the person search problem, better performance can be achieved by solving the pedestrian detection and person re-identification tasks separately rather than jointly.

We propose a Mask-guided Two-Stream CNN Model (MGTS) for person re-id, which explicitly makes use of one stream from the foreground as the emphasized information and enriches the representation by incorporating another separate stream from the original image.

Our proposed method achieves mAP of 83.0%83.0\% and 32.6%32.6\% on CUHK-SYSU and PRW benchmarks respectively, which improves over the previous state-of-the-arts by a large margin (more than 5pp).

Related Work

We first review existing works on person search, which is a recently proposed topic. Since our person search method is composed of two stages: pedestrian detection and person re-identification, we also review some recent works in both fields.

Person search. Person search has drawn much research interest since the publication of two large scale datasets: CUHK-SYSU and PRW . Zheng et al. conduct a detailed survey on various separated models and propose to solve the person search problem in two separate models, one for detection and another for re-ID. Other works propose to solve the problem in an end-to-end fashion by employing the Faster R-CNN detector for pedestrian detection and share the base network between detection and re-identification . In , feature discriminative power is increased by introducing center loss during training. Liu et al. improve the localization policy of Faster R-CNN by recursively shrink the search area from the whole image till achieving precise location of the target person. In this paper, we first made a systematic comparison between separated models and joint models, and show that a separated solution improves both the detection and re-identification results.

Pedestrian detection. Pedestrian detection is canonical object detection, especially when hand-crafted features are widely used. The classic HOG descriptor is based on local image differences, and successfully represents the special head-shoulder shape of pedestrians. A deformable part model (DPM) is proposed to handle deformations and still uses HOG as basic features. More recently, the integral channel feature (ICF) detectors become popular, as they achieve remarkable improvements while running fast. In recent years, convnets are also employed in pedestrian detection and further push forward the progress . Some works use the R-CNN architecture, which relies on ICF for proposal generation . Aiming for an end-to-end procedure, Faster R-CNN is adopted and it achieves top results by applying proper adaptations . Therefore, we use the adapted Faster R-CNN detector in this paper.

Person re-ID. Early person re-identification methods focus on manually designing discriminative features , using salient regions , and learning distance metrics . For instance, Zhao et al. propose to densely combine color histogram and SIFT features as the final multi-dimensional descriptor vector. Kostinger et al. present KISS method to learn a distance metric from equivalence constraints. CNN-based models have attracted extensive attentions since the successful applications by two pioneer works . Most of those CNN models can be categorized into two groups. The first group uses the siamese model with image pairs or triplets as inputs. The main idea of these works is to minimize the feature distance between the same person and maximize the distance between different people. The second group of works formulate the re-identification task as a classification problem . The main drawback of classification models is that they require more training data. Xiao et al. propose to combine multiple datasets for training and improve feature learning by domain guided dropout. Zheng et al. point out that classification models are able to reach higher accuracy than siamese model, even without careful sample choosing.

Recently, attention mechanism has been adopted to learn better features for person re-ID. For instance, HydraPlus-Net aggregates multiple feature layers within the spatial attentive regions extracted from multiple layers in the network. PDC model enriches the person representation with pose-normalized images and re-weights the features by channel-wise attention. In this paper, we also formulate person re-identification as a classification problem, and we propose to emphasize foreground information in the aggregated representation by adding an axillary stream with spatial attention (instance mask) and channel-wise re-weighting (SEBlock), which is similar to HydraPlus-Net and PDC model. However, our work differs from them in that the attention mechanism in our work is introduced with a different motivation, which is to consider the foreground-background relationship instead of local-global or part-whole relationship. In addition, the architecture of our model is more clear and concise, along with more practical training strategy without multi-staged fine-tuning.

Method

As shown in Fig. 1, our proposed person search method consists of two stages: pedestrian detection and re-identification. In this section, we first give an overview of our framework, and then describe more details for both stages individually.

A panoramic image is first fed into a pedestrian detector, which outputs several bounding boxes along with their confidence scores. We remove the bounding boxes whose confidence scores are lower than a given threshold. Only the remaining ones are used by the re-ID network.

A post-processing is implemented on the detected persons before they are sent to the re-ID stage. Specifically, we expand each RoI (Region of Interest) with a ratio of γ(γ>1)\gamma(\gamma>1) to include more context and crop out the person from the whole image. In order to separate the foreground person from background, we apply an off-the-shelf instance segmentation method FCIS on the whole image, and then designate the person to the right mask via majority vote. After that, for each person, we obtain a pair of images, one containing only the foreground person, and the other containing both the foreground and the background. See an illustration in Fig. 2.

Next up, in the re-ID model, the pair images go through two different paths, namely F-Net and O-Net, for individual modeling. Feature maps from the two paths are then concatenated and re-weighted by an SEBlock . After channel re-weighting, we pool the two dimensional feature maps into feature vectors using Global Average Pooling (GAP). Finally, the feature vectors are projected to an L2-normalized dd-dimensional subspace as the final identity descriptor.

The pedestrian detector and re-ID model are trained independently. In order to avoid the mistakes resulting from the detector, we use the ground truth annotations instead of detections to train the re-ID model.

2 Pedestrian Detection

We use a Faster R-CNN detector for pedestrian detection. The Faster R-CNN architecture is composed of a base network for feature extraction, a region proposal network (RPN) for proposal generation and a classification network for final predictions.

In this paper, we use VGG16 as our base network. The top convolutional layer ‘conv5_3’ produces 512 channels of feature maps, where the image resolution is reduced by a factor of 16. According to , up-sampling the input image is a reasonable way for compensation.

RPN is built upon ‘conv5_3’ to predict pedestrian candidate boxes. We follow the anchor settings in and set uniform scales ranging from the smallest and biggest persons we want to detect. The RPN produces a large number of proposals, so we apply a humble Non-Maximum Suppression (NMS) with an intersection over union (IoU) threshold of 0.7 to remove repeating ones and also cut off low-scoring ones by a given threshold.

The remaining proposals are then sent to the classification network, where an RoI pooling layer (512×7×7512\times 7\times 7) is used to generate the same length of features for each proposal. The final detection confidence and corresponding bounding box regression parameters are regressed by fully connected layers. After bounding box regression, another NMS with IoU threshold of 0.45 is applied and low-scoring detections are cut off.

The base net, RPN and classification network are jointly trained using Stochastic Gradient Descent (SGD).

3 Person Re-ID via A Mask-guided Two-Stream CNN Model

After RoIs for each person are obtained (either from a detector or ground truth), we aim to extract discriminative features. First of all, we expand each RoI by a ratio of γ\gamma to include more context. Then, we propose a two-stream structure to extract features for foreground person and whole image individually. The features from both streams are concatenated as enriched representations for the RoIs and a re-weighting operation is applied to highlight more informative features while suppressing less useful ones.

Foreground separation. The key step is to separate foreground and background for each RoI. We first apply an instance segmentation method FCIS on the whole image to obtain segmentation masks for persons. After that, we associate each RoI with its corresponding mask by majority vote. Those pixels inside and outside the mask boundary are considered as foreground and background respectively. We describe the detailed separation procedure in Algorithm 1 and show an example in Fig. 2.

Feature re-weighting. We further re-weight all the feature maps with an SEBlock , which models the interdependencies between channels of convolutional features. The architecture of an SEBlock is illustrated in Fig. 3. It is defined as a transformation from F to F′\textbf{F}^{\prime}:

σ\sigma and δ\delta refer to the Sigmoid activation and the ReLU function respectively. W1\textbf{W}_{1} and W2\textbf{W}_{2} are the weight matrix of two FC layers. fGAPf_{GAP} is the operation of GAP and ⋅\cdot denotes channel-wise multiplication. SEBlock learns to selectively emphasis informative features and suppress less useful ones by re-weighting channel features using the weighting vector w. In this way, foreground and background information are fully explored and re-calibrated, and hence help to optimize the final feature descriptor for person re-identification.

The whole MGTS model is trained with ground truth RoIs and supervised by an Online Instance Matching loss (OIM) . The OIM objective is to maximize the expected log-likelihood:

ptp_{t} denotes the probability of x belonging to class tt. τ\tau is a temperature factor similar to the one in Softmax function. vt\textbf{v}_{t} is the class central feature vector of the tt-th class. It is stored in a lookup table with size LL and incrementally updated during training with a momentum of η\eta:

where uk\textbf{u}_{k} is a feature vector for an unlabeled person. A circular queue of size QQ is used to store uk\textbf{u}_{k} vectors. It pops out old features and pushes in new features during training.

Experiments

In this section, we first introduce the datasets and evaluation protocols we use in our experiments, followed by some implementation details. After that, we show experimental results with comparison to state-of-the-art methods, followed by an ablation study to verify the design of our approach.

CUHK-SYSU. CUHK-SYSU is a large-scale person search database consisted of street/urban scene images captured by a hand-held camera or selected from movie snapshots. It contains 18,18418,184 images and 96,14396,143 pedestrian bounding boxes. There are a total of 8,4328,432 labeled identities, and the rest of the pedestrians are served as negative samples for identification. We adopt the standard train/test split provided by the dataset, where the training set includes 11,20611,206 images and 5,5325,532 identities, while the testing set contains 2,9002,900 probe persons and 6,9786,978 gallery images. Moreover, each probe person corresponds to several gallery subsets with different sizes, which are defined in the dataset.

PRW. The PRW dataset is extracted from video frames captured with six cameras in a university campus. There are 11,81611,816 frames annotated with 34,30434,304 bounding boxes. Among all the pedestrians, 932932 identities are tagged and the rest of them are marked as unknown persons similar to CUHK-SYSU. The training set includes 5,1345,134 images with 482482 different persons. The testing set contains 2,0572,057 probe persons and 6,1126,112 gallery images. Different from CUHK-SYSU, the whole gallery set serves as the search space for each probe person.

2 Evaluation Protocol

Pedestrian detection. Average Precision (AP) and recall are used to measure the performance of pedestrian detection. A detection bounding box is considered as a true positive if and only if its overlap ratio with any ground truth annotation is above 0.50.5.

Person search. We adopt the mean Average Precision (mAP) and cumulative matching characteristics (CMC top-K) as performance metrics for re-identification and person search. The mAP metric reflects the accuracy and matching rate of searching a probe person from gallery images. CMC top-K is widely used for person re-identification task, where a matching is counted if there is at least one of the top-K predicted bounding boxes overlapping with the ground truth with an intersection-over-union (IoU) larger than or equal to a threshold. The threshold is set to 0.5 throughout the paper.

3 Implementation Details

We implement our system with Pytorch. The VGG-based pedestrian detector is initialized with an ImageNet-pretrained model. It is trained using SGD with a batch size of 1. Input images are resized to have at least 900900 pixels on the short side and at most 1,5001,500 pixels on the long side. The initial learning rate is 0.001, decayed by a factor of 0.1 at 60K and 80K iterations respectively and kept unchanged until the model converges at 100K iterations. The first two convolutional blocks (‘conv1’ and ‘conv2’) are frozen during training, while other layers are updated.

The RoI expand ratio γ\gamma is set to 1.3. Both F-Net and O-Net of our MGTS model are based on ResNet50 and truncated at the last convolutional layer (‘conv5_3’). The input image patches are re-scaled to an arbitrary size of 256×128256\times 128 and the batch size is set to 128. The model is trained with an initial learning rate of 0.001, decayed to 0.0001 after 11 epochs and kept identical until early cutting after epoch 15. The temperature scalar τ\tau, circular queue size QQ and momentum η\eta in OIM loss are set to 1/301/30, 50005000 and 0.50.5 respectively. The feature dimension dd is set to 128 through out the paper if not specified.

As of foreground extraction, we use the off-the-shelf instance segmentation method FCIS trained on COCO trainval35k without any fine-tuninghttps://github.com/msracver/FCIS. Sample results of instance masks from CUHK-SYSU and PRW are shown in Fig. 4, where we can see FCIS generalizes well to both datasets.

4 Comparison with State-of-the-Art Methods

In this subsection, we report our person search results on the CUHK-SYSU and PRW datasets, with a comparison to several state-of-the-art methods, including OIM , IAN , NPSM and IDE . Other than the above joint methods, we also compare with some methods which also solve the person search problem in two steps of pedestrian detection and person re-identification, similar to our method. These methods use different pedestrian detectors (DPM , ACF , CCF , LDCF ), person descriptors (BoW , LOMO , DSIFT ) and distance metrics (KISSME , XQDA ).

Results on CUHK-SYSU. Table 1 shows the person search results on CUHK-SYSU with a gallery size of 100. We follow the notations defined in , where “CNN” denotes the Faster R-CNN detector based on ResNet50 and “IDNet” represents a re-identification net defined in . Our VGG-based detector is marked as “CNNv”. IDNetOIM is a re-identification net trained with ground truth RoIs and supervised by an OIM loss. Compared with OIM, CNNv + IDNetOIM slightly improves the performance by solving detection and re-identification tasks in two independent models. By further employing our proposed MGTS model, we achieve 83.0%83.0\% mAP. Our final model outperforms the state-of-the-art method by more than 55 pp w.r.t. mAP, and 2.52.5 pp w.r.t. CMC top-1.

Moreover, we evaluate the proposed method (CNNv + MGTS) under different gallery sizes along with other competitive methods. Figure 5 shows how the mAP changes with a varying gallery size of . We can see that all methods suffer from a performance degeneration as the gallery size increases. However, our method outperforms others under different settings, which indicates the robustness of our method. Besides, we notice that the gap between our method and others become even larger as gallery size increases.

We also show some qualitative results of our method and the competitive baseline OIM in Fig. 7. As can be seen in the figure, our method performs better on hard cases where gallery persons wear similar clothes with the probe person, possibly with the help of context information in the expanded RoI, e.g. accompanied person (Fig. 7, 7), handrail (Fig. 7), baby carriage (Fig. 7) etc. It is also more robust on hard cases where gallery entries share both similar foreground and background to the probe person (Fig.7, 7), where more subtle differences like hairstyle and gender shall be excavated from the emphasized foreground person. Fig. 7 shows a failure case where both OIM and MGTS suffer from bad illumination condition, which is rather challenging and needs more efforts in future works.

Results on PRW. In Table 2 we report the evaluation results on the PRW dataset. A number of combinations of detection methods and re-identification models are explored in . Among them, DPM + AlexNet -based R-CNN with ID-discriminative Embedding (IDEdet) and Confidence Weighted Similarity(CWS) achieves the best performance. In contrast, joint methods including OIM, IAN and NPSM, all achieve better results. But it is unclear whether the improvement comes from a joint solution or the usage of deeper networks (ResNet50/ResNet101) and a better performed detector (Faster R-CNN).

For fair comparison, we also employ ResNet50 and Faster R-CNN in our framework, and achieve significant improvements compared to joint methods. Specifically, we outperform the state of the art by 8.48.4 pp and 10.210.2 pp w.r.t. mAP and top-1 accuracy. These results again demonstrate the effectiveness of our proposed method.

5 Ablation Study

From the experimental results in Sec. 4.4, we obtain significant improvement to our baseline method OIM on both standard benchmarks. The major differences between our method and OIM are as follows: (1) We solve the pedestrian detection and person re-identification tasks separately rather than jointly, i.e. we do not share features between them. (2) In the re-identification net, we model the foreground and original image in two parallel streams so as to obtain enriched representations.

In order to understand the impact of the above two changes, we conduct analytical experiments on CUHK-SYSU at a gallery size of 100, and provide discussions in the following.

Integration vs. Separation. We investigate the impact of sharing features between detection and re-identification task on both performance.

In Table 3, we compare the detection performance of a jointly trained model and a vanilla detector. We can see that the jointly trained detector under-performs the vanilla one by 8.58.5 pp w.r.t. AP while reaching the same recall.

Similarly, we make a comparison on re-ID performance between a jointly trained model and a vanilla re-ID net in Table 3. The person search performance of a jointly trained OIM method is 0.60.6 pp and 1.21.2 pp lower in mAP and top-1 accuracy than a vanilla re-ID net (IDNetOIM).

From the above comparisons, we conclude that joint training harms both the detection and re-ID performance, thus it is a better solution to solve them separately.

Visual Component Study. In this part, we study the contributions of foreground and original image information to a re-ID system. To exclude the influence of detectors, all the following models are trained and tested using ground truth RoIs on CUHK-SYSU with a gallery size of 100. They are based on ResNet50 and supervised by an OIM loss.

Four variants to the input RoI patch and their combinations are considered: (1) Original RoI (O); (2) Masked foreground (F); (3) Masked background (B); (4) Expanded RoI (E) with a ratio of γ\gamma.

Comparison results are shown in Table 4, from which we make the following observations:

Background is an important cue for re-ID. The mAP drops by 2.82.8 pp when only the foreground is used, while discarding all background. More interestingly, using only background information yields an mAP of 34.2%34.2\%, which can be further pushed to 38.7%38.7\% if RoI is expanded.

Modeling foreground and original image in two streams improves the results significantly. The two-stream model O+F+E reaches 89.1%89.1\% mAP, surpassing the one-stream model O+E by 11.411.4 pp.

6 Model Inspection

To further understand the respective impact of the two streams, we provide an analysis on the SEBlock weights of foreground vs. original image representations. The analysis is implemented based on the trained models in Sec.4.4, to which we feed all training samples in CUHK-SYSU (96,14396,143 proposals) and PRW (42,87142,871 proposals) respectively. For sample ii, we compute three metrics: (1) average weight of F-Net channels, Avgi(F)Avg_{i}(F); (2) average weight of O-Net channels, Avgi(O)Avg_{i}(O); (3) number of channels that among top 20 of the whole network while come from F-Net, N20(F)N^{20}(F), the histograms of N20(F)N^{20}(F) for all training samples from two datasets are shown in Fig. 6.

Based on analyzing the above statistics, we have the following findings:

The inequation Avgi(F)>Avgi(O)Avg_{i}(F)>Avg_{i}(O) holds for all samples. It demonstrates that in general the foreground patch contributes more than the original patch to the final feature vector, as it involves more informative cues for each identity.

From Fig. 6, the majority of the top 20 channels come from the foreground stream for most samples. This observation indicates that the most informative cues are from the foreground patch.

Although the majority of the top 20 channels are represented by the foreground patch, we still observe that quite a few top channels are from the original patch. This is a good evidence showing the context information contained in the original image patch is helpful for the re-identification task.

Moreover, the impact of the amount of context information is inspected by changing the RoI expansion value γ\gamma. We conduct a set of experiments on CUHK-SYSU with a gallery size of 100 and list the results in Table 5, from which we can draw the intuitive conclusion that 1) the γ\gamma is relatively stable when γ∈[1.2 1.5]\gamma\in[1.2\ 1.5]; and 2) a proper amount of context information is better than no context, while too much background could be harmful.

7 Runtime Analysis and Acceleration

We implement our runtime analysis on a Tesla K80 GPU. For testing a 1500×9001500\times 900 input image, our proposed approach takes around 1.3 seconds on average, including 626 ms for pedestrian detection, 579 ms for segmentation mask generation and another 64 ms for person re-identification.

We notice half of the computational time is used to generate the segmentation mask. In order to accelerate, we propose an alternative option to use the tight ground truth bounding boxes as ‘weak’ masks instead of the ‘accurate’ FCIS masks. The results are presented in Tab. 6, from which we can see that using ‘accurate’ FCIS masks yields better performance than using bounding box masks. However, using bounding boxes as weak masks still achieves promising results, which outperforms the single stream model without using any masks by a large margin (∼\sim 7pp mAP) at comparable time cost. Therefore, our proposed method can be accelerated by a factor of ∼\sim2x with an acceptable performance drop, while still surpassing the state-of-the-art results.

Conclusion

In this paper, we propose a novel deep learning based approach for person search. The task is accomplished in two steps: we first apply a Faster R-CNN detector for pedestrian detection on gallery images; and then make matchings between the probe image and output detections via a re-identification network. We obtain better results by training the detector and re-identification network separately, without sharing representations. We further enhance the re-identification accuracy by modeling the foreground and original image patches in two subnets to obtain enriched representations. Experimental results show that our proposed method significantly outperforms state-of-the-art methods on two standard benchmarks.

Inspired by the success of utilizing the segmented foreground patch for additional feature extraction, for future work, we will explore to optimize the segmentation masks and identification accuracy in a joint framework, so as to obtain finer masks.

Acknowledgments

This work was supported by the National Science Fund of China under Grant Nos. U1713208, 61472187 and 61702262, the 973 Program No.2014CB349303, Program for Changjiang Scholars, and “the Fundamental Research Funds for the Central Universities” No.30918011322.

References