Recover and Identify: A Generative Dual Model for Cross-Resolution Person Re-Identification
Yu-Jhe Li, Yun-Chun Chen, Yen-Yu Lin, Xiaofei Du, Yu-Chiang Frank Wang
Introduction
Person re-identification (re-ID) aims at recognizing the same person across images taken by different cameras, and is an active research topic in computer vision. A variety of applications ranging from person tracking , video surveillance system , to computational forensics are highly correlated this research topic. Nevertheless, due to the presence of background clutter, occlusion, illumination or viewpoint changes, person re-ID remains a challenging task for practical applications.
Driven by the recent success of convolutional neural networks (CNNs), several learning-based methods have been proposed. Despite promising performances, these methods are typically developed under the assumption that both query and gallery images are of similar or sufficiently high resolutions. This assumption, however, may not hold in practice since image resolutions would vary drastically. For instance, query images captured by surveillance cameras are often of low resolution (LR) whereas those in the gallery set are carefully selected beforehand and are of high resolution (HR). As a result, direct matching of LR query images and HR gallery ones would lead to non-trivial resolution mismatch problems.
To address cross-resolution person re-ID, most existing methods employ super-resolution (SR) models to convert LR inputs into their HR versions followed by person re-ID. However, these methods suffer from two limitations. First, each employed SR model is designed to upscale image resolutions by a particular factor. Thus, these methods need to pre-determine the resolutions of LR queries so that the corresponding SR models can be applied. However, designing SR models for each possible resolution input makes these methods hard to scale. Second, in the real-world scenario, queries can be with various resolutions even with the resolutions that are unseen during training. As illustrated in Figure 1, queries with varying or unseen resolutions would restrict the applicability of the person re-ID methods that employ SR models since one cannot assume the resolutions of the input images will be known in advance.
In this paper, we propose Cross-resolution Adversarial Dual Network (CAD-Net) for cross-resolution person re-ID. The key characteristics of CAD-Net are two-fold. First, to address the resolution variations, CAD-Net derives the resolution-invariant representations via adversarial learning. This allows our model to handle images of varying and even unseen resolutions. Second, CAD-Net learns to recover the missing details in LR input images. Together with the resolution-invariant features, our model generates HR images preferable for person re-ID, achieving the state-of-the-art performance on cross-resolution person re-ID. It is worth noting that the above image resolution recovery and cross-resolution person re-ID are realized by a single model learned in an end-to-end fashion.
The contributions of this paper are highlighted below:
We propose an end-to-end trainable network which advances adversarial learning strategies for cross-resolution person re-ID.
Our model learns resolution-invariant representations while recovering the missing details in LR input images, resulting in improved cross-resolution person re-ID performance.
Our model is able to handle query images with varying or even unseen resolutions without the need to pre-determine the input resolutions.
Extensive experimental results on five challenging datasets confirm that our method performs favorably against the state-of-the-art person re-ID approaches.
Related Work
A variety of existing methods are developed to address various challenges in person re-ID, such as background clutter, viewpoint changes, and pose variations. For instance, Yang et al. learn a camera-invariant subspace to deal with the style variations caused by different cameras. Liu et al. develop a pose-transferable framework based on the generative adversarial network (GAN) to yield pose-specific images for tackling the pose variations. Several methods addressing background clutter leverage attention mechanisms to emphasize the discriminative parts. Another research trend focuses on domain adaptation for person re-ID . By viewing image-to-image translation methods as a data augmentation technique, these methods employ image translation modules, e.g., CycleGAN , to generate viewpoint specific images with labels. However, the above approaches typically assume that both query and gallery images are of similar or sufficiently high resolutions, which might not be practical for real-world applications.
Cross-resolution person re-ID.
A number of methods have been proposed to address the problem of resolution mismatch in person re-ID. Li et al. jointly perform multi-scale distance metric learning and cross-scale image domain alignment. Jing et al. develop a semi-coupled low-rank dictionary learning framework to seek a mapping between HR and LR images. Wang et al. learn a discriminating scale-distance function space by varying the image scale of LR images when matching with the HR ones. Nevertheless, these methods adopt hand-crafted descriptors, which cannot easily adapt the developed models to the tasks of interest, and thus may lead to sub-optimal person re-ID performance.
Recently, three CNN-based methods are presented for cross-resolution person re-ID. The network of SING is composed of several SR sub-networks and a person re-ID module to carry out LR person re-ID. On the other hand, CSR-GAN cascades multiple SR-GANs and progressively recovers the details of LR images to address the resolution mismatch problem. In spite of their promising results, such methods require the training of pre-defined SR models. As mentioned earlier, the degree of resolution mismatch, i.e., the resolution difference between the query and gallery images, is typically unknown beforehand. Moreover, if the resolution of the input LR query is unseen during training, the above methods cannot be easily applied or might not lead to satisfactory performance. Apart from these methods, RAIN aligns the feature distributions of HR and LR images, showing some performance improvement over existing algorithms.
Similar to RAIN , our method also performs feature distribution alignment between HR and LR images. Our model differs from RAIN in two aspects. First, our model derives resolution-invariant representations and recovers the missing details in LR input images. By jointly considering features of both modalities, our algorithm further improves the performance. Second, the HR image recovery is learned in an end-to-end fashion, allowing our model to recover HR images preferable for person re-ID. Experimental results demonstrate that our approach can be applied to input images of varying and even unseen resolutions using only a single model.
Cross-resolution vision applications.
The issues regarding cross-resolution handling have been studied in the literature. For face recognition, existing approaches typically rely on face hallucination algorithms or SR mechanisms to super-resolve the facial details. Unlike the above existing methods that focus on synthesizing the facial details, our model learns to recover re-ID oriented discriminative details. Together with the derived resolution-invariant features, our model would considerably boost the person re-ID performance while allowing query images with varying and even unseen resolutions.
Proposed Method
In this section, we first provide an overview of our proposed approach. We then describe the details of each network component as well as the loss functions.
2 Cross-Resolution GAN (CRGAN)
In CRGAN, we have a cross-resolution encoder which converts input images across different resolutions into resolution-invariant representations, followed by a high-resolution decoder recovering the associated HR versions.
High-resolution decoder 𝒢𝒢\mathcal{G}.
In addition to learning the resolution-invariant representation , our CRGAN further synthesizes the associated HR images. This is to recover the missing details in LR input images, together with the person re-ID task to be performed later in the cross-modal re-ID network.
It is also worth repeating that the goal of this HR decoder is not simply to recover the missing details in LR input images, but also to have such recovered HR images aligned with the learning task of interest (i.e., person re-ID). Namely, we encourage the HR decoder to perform re-ID oriented HR recovery, which is further realized by the following cross-modal re-ID network.
3 Cross-Modal Re-ID
The total loss function for training our proposed CAD-Net is summarized as follows:
Experiments
We first provide the implementation details, followed by dataset descriptions and settings. Both quantitative and qualitative results are presented, including ablation studies.
2 Datasets
We evaluate the proposed method on five datasets, each of which is described as follows.
The CUHK03 dataset comprises images of identities with different camera views. Following CSR-GAN , we use the training/test identity split.
VIPeR [17].
The VIPeR dataset contains person-image pairs captured by cameras. Following SING , we randomly divide this dataset into two non-overlapping halves based on the identity labels. Namely, images of a subject belong to either the training set or the test set.
CAVIAR [11].
The CAVIAR dataset is composed of images of person identities captured by cameras. Following SING , we discard people who only appear in the closer camera, and split this dataset into two non-overlapping halves according to the identity labels.
Market-150115011501 [48].
The Market- dataset consists of images of identities with camera views. We use the widely adopted training/test identity split.
DukeMTMC-reID [50].
The DukeMTMC-reID dataset contains images of identities captured by cameras. We adopt the benchmarking training/test identity split.
3 Experimental Settings and Evaluation Metrics
We evaluate the proposed method using cross-resolution person re-ID setting where the test (query) set is composed of LR images while the gallery set contains HR images only. In all of the experiments, we adopt the standard single-shot person re-ID setting and use the average cumulative match characteristic as the evaluation metric.
4 Evaluation and Comparisons
Following SING , we consider multiple low-resolution (MLR) person re-ID and evaluate the proposed method on four synthetic and one real-world benchmarks. To construct the synthetic MLR datasets (i.e., MLR-CUHK03, MLR-VIPeR, MLR-Market-, and MLR-DukeMTMC-reID), we follow SING and down-sample images taken by one camera by a randomly selected down-sampling rate (i.e., the size of the down-sampled image becomes ), while the images taken by the other camera(s) remain unchanged. The CAVIAR dataset inherently contains realistic images of multiple resolutions, and is a genuine and more challenging dataset for evaluating MLR person re-ID.
We compare our approach with methods developed for cross-resolution person re-ID, including JUDEA , SLD2L , SDF , SING , and CSR-GAN , and methods developed for standard person re-ID, including CamStyle and FD-GAN . For methods developed for cross-resolution person re-ID, the training set contains HR images and LR ones with all three down-sampling rates for each person. For methods developed for standard person re-ID, the training set contains HR images for each identity only.
Table 1 reports the quantitative results recorded at ranks , , and on all five adopted datasets. For CSR-GAN on MLR-CUHK03, CAVIAR, MLR-Market-, and MLR-DukeMTMC-reID, and CamStyle and FD-GAN on all five adopted datasets, their results are obtained by running the released code with the default implementation setup. For SING , we reproduce their results on MLR-Market- and MLR-DukeMTMC-reID.
We note that the performance of our method can be further improved by applying pre-/post-processing methods, attention mechanisms, or re-ranking. For fair comparisons, no such techniques are used in all of our experiments.
In Table 1, our method performs favorably against all competing methods on all five datasets. We observe that our method consistently outperforms the best competitors by at rank . The performance gains can be ascribed to three main factors. First, unlike most existing person re-ID methods, our model performs cross-resolution person re-ID in an end-to-end learning fashion. Second, our method learns resolution-invariant representations, allowing our model to recognize persons in images of different resolutions. Third, our model learns to recover the missing details in LR input images, thus providing additional discriminative evidence for person re-ID.
The advantage of deriving joint representation can be assessed by comparing with two of our variant methods, i.e., Ours ( only) and Ours ( only). In method “Ours ( only)”, the classifier only takes the resolution-invariant representation as input. In method “Ours ( only)”, the classifier only takes the HR representation as input. We observe that deriving joint representation consistently improves the performance over these two baseline methods. We note that method “Ours ( only)” achieves a better performance than method “Ours ( only)” on the CAVIAR dataset. We attribute the results to the higher resolution variations exhibited in the CAVIAR dataset.
5 Evaluation of the Recovered HR Images
To demonstrate that our CRGAN is capable of recovering the missing details in LR images of varying and even unseen resolutions, we evaluate the quality of the recovered HR images on the MLR-CUHK03 test set using SSIM, PSNR, and LPIPS metrics. We employ the ImageNet-pretrained AlexNet when computing LPIPS. We compare our CRGAN with CycleGAN , SING , and CSR-GAN . For CycleGAN , we train the model to learn a mapping between LR and HR images. We report the quantitative results of the recovered image quality and person re-ID in Table 2 with two different settings: (1) LR images of resolutions seen during training, i.e., , and (2) LR images of unseen resolution, i.e., .
For seen resolutions (i.e., left block), we observe that our results using SSIM and PSNR metrics are slightly worse than CSR-GAN while compares favorably against SING and CycleGAN . However, our method performs favorably against these three methods using LPIPS metric and achieves the state-of-the-art performance when evaluating on cross-resolution person re-ID task. These results indicate that (1) SSIM and PSNR metrics are low-level pixel-wise metrics, which do not reflect high-level perceptual tasks and (2) the end-to-end learning of cross-resolution person re-ID would result in better person re-ID performance and recover more perceptually realistic HR images as reflected by LPIPS.
For unseen resolution (i.e., right block), our method performs favorably against all three competing methods on all the adopted evaluation metrics. These results suggest that our method is capable of handling unseen resolution (i.e., ) with favorable performance in terms of both image quality and person re-ID. Note that we only train our model with HR images and LR ones with .
Figure 3 presents six examples. For each person, there are four different resolutions (i.e., ). Note that images with down-sampling rate indicate that the images remain their original sizes and are the corresponding HR images of the LR ones. We observe that when LR images with down-sampling rate are given, our model recovers the HR details with the highest visual quality among all competing methods. Both quantitative and qualitative results above confirm that our model can handle a range of seen resolutions and generalize well to unseen resolutions using just one single model, i.e., CRGAN.
6 Ablation Study
To analyze the importance of each developed loss function, we conduct an ablation study on the MLR-CUHK03 dataset. Table 3 reports the quality of the recovered HR images and the performance of cross-resolution person re-ID recorded at rank .
7 Resolution-Invariant Representation f𝑓f
We note that the projected feature vectors in Figure 4(b) are well separated, suggesting that sufficient person re-ID ability can be exhibited by our model. On the other hand, for Figure 4(c), we colorize each image resolution with a unique color in each identity cluster (four different down-sampling rates ). We observe that the projected feature vectors of the same identity but different down-sampling rates are all well clustered. We note that images with down-sampling rate are not present in the training set (i.e., unseen resolution).
The above visualizations demonstrate that our model learns resolution-invariant representations and generalizes well to unseen image resolution (e.g., ) for cross-resolution person re-ID.
Conclusions
We have presented an end-to-end trainable generative adversarial network, CAD-Net, for addressing the resolution mismatch issue in person re-ID. The core technical novelty lies in the unique design of the proposed CRGAN which learns the resolution-invariant representations while being able to recover re-ID oriented HR details. Our cross-modal re-ID network jointly considers the information from two feature modalities, leading to better person re-ID capability. Extensive experimental results demonstrate that our approach performs favorably against existing cross-resolution and standard person re-ID methods on five challenging benchmarks, and produces perceptually higher quality HR images using only a single model. Visualization of the resolution-invariant representations further verifies our ability in handling query images with varying or even unseen resolutions. Thus, the use of our model for practical person re-ID applications can be strongly supported.
This work is supported in part by Ministry of Science and Technology (MOST) under grants 107-2628-E-001-005-MY3, 108-2634-F-007-009, and 108-2634-F-002-018, and Umbo Computer Vision.