Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association

Dapeng Chen, Hongsheng Li, Xihui Liu, Yantao Shen, Zejian Yuan, Xiaogang Wang

Introduction

Person re-identification (re-ID) is a critical task in intelligent video surveillance, aiming to associate the same people across different cameras. Encouraged by the remarkable success of deep Convolutional Neural Network (CNN) in image classification , the re-ID community has made great process by developing various networks, yielding quite effective visual representations. To further boost the identification accuracy, diverse auxiliary information has been incorporated in the deep neural networks, such as the camera ID information , human poses , person attributes , depth maps, and infrared person images . These data are utilized as either the augmented information for an enhanced inter-image similarity estimation or the training supervisions that can regularize the feature learning process . Our work belongs to the latter category and proposes to use language descriptions as training supervisions to improve the person visual features. Compared with other types of auxiliary information, natural language provides a flexible and compact way of describing the salient visual aspects for distinguishing different persons. Previous efforts on language-based person re-ID is about cross-modal image-text retrieval, aiming to search the target image from a gallery set by a text query. Instead, we are interested in how the language can help the image-to-image search when they are only utilized in the training stage. This task is non-trivial because it requires a detailed understanding of the content of images, language, and their cross-modal correspondences.

To exploit the semantic information conveyed in the language descriptions, we not only need to identify the final image representation but also propose to optimize the global and local association between the intermediate features and linguistic features. The global image-language association is learned from their ID labels. That is, the overall image feature and text feature should have high relevance for the same person, and have low relevance when they are from different persons (Fig. 1, left). The local image-language association is based on the implicit correspondences between image regions and noun phrases (Fig . 1, right). As in a coupled image-text pair, a noun phrase in the text usually describes a specific region in the image, thus the phrase feature is more related to some local visual features. We design a deep neural network to automatically associate related phrases and local visual features via the attention mechanism, then aggregate these visual features to reconstruct the phrase. Reasoning such latent and inter-modal correspondence makes the feature embedding interpretable, can be employed as a regularization scheme for feature learning.

In summary, our contributions are three-fold: (1) We propose to use language description as training supervisions for learning more discriminative visual representation for person re-ID. This is different from existing text-image embedding methods aiming at cross-modal retrieval. (2) We provide two effective and complementary image-language association schemes, which utilize semantic, linguistic information to guide the learning of visual features in different granularities. (3) Extensive ablation studies validate the effectiveness and complementarity of the two association schemes. Our method achieves state-of-the-art performance on person re-ID and outperforms conventional cross-modal embedding methods.

Related Work

Early works on person re-ID concentrated on either feature extraction or metric learning . Recent methods mainly benefit from the advances of CNN architectures , which combine the above two aspects to produce robust and ID-discriminative image representation . Our work aims to further improve the deep visual representation by making use of language descriptions as training supervisions.

Diverse auxiliary information has been introduced to improve the visual feature representations for person re-ID. Several works detected person pose landmarks to obtain the human body regions. They firstly decomposed the feature maps according to the regions, then fused them to create the well-aligned feature maps. Lin et al. utilized Camera ID information to assist inter-image similarity estimation by keeping consistencies in a camera network. Also, different types of sensors such as depth cameras , or infrared cameras have been employed in person re-ID to generate more reliable visual representations. For these methods, the auxiliary information is used in both training and testing stage, requiring an additional model or data acquisition device for algorithm deployment. Differently, human attributes usually serve as a kind of training supervisions. For example, Su et al. learned a semi-supervised discriminative model to predict the binary attribute feature for re-ID. Lin et al. improved the interpretability of the intermediate feature maps by jointly optimizing the identification loss and attribute classification loss. Although attributes proves helpful for feature learning, they are quite difficult to obtain as people need to remember tens of attribute labels for annotations. They are also less flexible to describe diverse variations in human appearance.

Associating image and language helps establish correspondences for their inter-relations. It has attracted great attention in recent years because of its wide applications in image captioning , visual QA , and text-image retrieval . These cross-modal associations can be modeled by either generative methods or discriminative methods. Generative models utilize probabilistic models to capture the temporal or spatial dependencies within the image or text , and have popular applications like caption generation and image generation . On the other hand, discriminative models have also been developed for image-text association. Karpathy and Fei-Fei formulated a bidirectional ranking loss to associate the text and image fragments. Reed et al. proposed deep symmetric structured joint embeddings, and enforced the embedding of matched image-text pair should be higher than those of unmatched pairs. Our method combines the merits of both discriminative and generative methods to build image-text association in different granularities, where the language descriptions act as training supervisions to improve visual representation.

Our Approach

To improve the visual representation for person re-ID with deep neural networks, we aim to exploit language descriptions of person images as the training supervisions in addition to the original ID labels. The visual representations are not only required to be discriminative for different persons but also need to keep consistencies with the linguistic representations. We, therefore, propose the global and local image-language association schemes. The global visual feature of one person should be more relevant to the language description features of the same person than those of a different person. Unlike existing cross-modal joint embedding methods, we do not require the visual and linguistic features to be mapped to a unified embedding space. Furthermore, based on the assumption that the image and language are spatially decomposable and temporally decomposable, we also try to find the mutual correspondences between the features of the image regions and the noun-phrases. The overall framework is illustrated in Fig. 2.

Given a dataset D={(In,Tn,ln)}n=1N\mathcal{D}=\{(I_{n},T_{n},l_{n})\}_{n=1}^{N} containing NN tuples, each tuple has an image II, a text description TT, and an ID label ll. To improve the learned visual feature ϕ(I)\phi(I), we build global and local correspondences between the intermediate visual feature maps Ψ(I)\Psi(I) and linguistic representation Θ(T)\Theta(T).

The visual representation. The visual feature ϕ(I)\phi(I) and the intermediate feature map Ψ(I)\Psi(I) are obtained from standard convolutional neural network (CNN), which takes ResNet-50 as the backbone network. Ψ(I)\Psi(I) is the feature map obtained with the 1 ⁣× ⁣11\!\times\!1 convolution over the last residual-block. Suppose the Ψ(I)\Psi(I) has KK bins, the feature vector at the kkth bin is denoted by ψk(I)\psi_{k}(I), then Ψ(I)\Psi(I) can be represented as Ψ(I)={ψk(I)}k=1K\Psi(I)=\{\psi_{k}(I)\}_{k=1}^{K}. The objective visual feature vector ϕ(I)\phi(I) is linear projection from the average feature map ψˉ(I)=1K∑k=1Kψk(I)\bar{\psi}(I)=\frac{1}{K}\sum_{k=1}^{K}\psi_{k}(I):

We employ the ID loss over ϕ(I)\phi(I), aiming to make it distinctive for different persons. Specifically, given NN images belonging to II persons, the ID loss is the average negative log-likelihood of the feature maps being correctly classified to its ID:

2 Global Discriminative Image-language Association

The ID losses in the previous section only enforce the visual and linguistic feature be discriminative within each modality but do not establish image-language correspondences to enhance the visual feature. As the global description is usually related to multiple and diverse regions in the image, θg(T)\theta^{g}(T) can be associated to ψˉ(I)\bar{\psi}(I) (Eqn. (1)) in a discriminative fashion. Specifically, ψˉ(I)\bar{\psi}(I) and θg(T)\theta^{g}(T) firstly form a joint representation φ(I,T)\varphi(I,T):

where ∘\circ denotes the Hadamard product. The joint representation is then projected into a scalar value within the range (0,1)(0,1) by:

To build the relevance between ψˉ(I)\bar{\psi}(I) and θg(T)\theta^{g}(T), we expect s(I,T)s(I,T) to be 1 when II and TT belong to the same person and to be when they belong to different persons. We thus impose the binary cross-entropy loss over the scores:

where N^\hat{N} is the number of sampled image-text pairs. li,j=1l_{i,j}=1 if IiI_{i} and TjT_{j} are describing a same person and li,j=0l_{i,j}=0 otherwise.

Discussion. Here, we draw a distinction between the proposed discriminative scheme and the bi-directional ranking , which is formulated by:

where ki,j=ψˉ(Ii)⊤θg(Tj)k_{i,j}=\bar{\psi}(I_{i})^{\top}\theta^{g}(T_{j}). The loss stipulates that the cosine similarity ki,ik_{i,i} for one image-text tuple should be higher than ki,jk_{i,j} or kj,ik_{j,i} for any i ⁣≠ ⁣ji\!\neq\!j by at least a margin of α\alpha. We highlight two main differences between the proposed Ldis\mathcal{L}_{dis}(Eqn. (7)) and Lrank\mathcal{L}_{rank}: (1) As Lrank\mathcal{L}_{rank} is originally applied in the image-text retrieval task, it associates the image and text description features by simply checking whether they are from the same tuple. Differently, Ldis\mathcal{L}_{dis} is based on person ID, which is more reasonable as one description can well correspond to different images of the same person. (2) Lrank\mathcal{L}_{rank} estimates the image-text relevance by cosine similarity, requiring ψˉ(Ii)\bar{\psi}(I_{i}) and θg(Tj)\theta^{g}(T_{j}) lie in the same feature space. Meanwhile, our scheme employs a projection over the joint representation, being able to capture more complicated correlations between image and text description.

3 Local Reconstructive Image-language Association

A phrase usually only describes one part of an image and could be contained in the descriptions of different persons. For this reason, a phrase is disjoint with the person ID, but can still build correspondences with a certain region in the image it describes. We therefore propose a reconstruction scheme. That is, the phrase feature θl(P)\theta^{l}(P) can select relevant feature vectors in visual feature map Ψ(In)\Psi(I_{n}) if P∈P(Tn)P\in\mathcal{P}(T_{n}), and the selected feature vectors are able to reconstruct the phrase PP in turn.

Image feature aggregation. Suppose PP is a phrase that describes a specific region in image InI_{n}, we aim to estimate a vector ψ^P(In)\hat{\psi}_{P}(I_{n}) that can reflect the features in the region. For this purpose, we compute ψ^P(In)\hat{\psi}_{P}(I_{n}) by weighted aggregation of the feature vectors {ψk(In)}k=1K\{\psi_{k}(I_{n})\}_{k=1}^{K} in the feature map Ψ(In)\Psi(I_{n}):

where rk(P,In)r_{k}(P,I_{n}) is the attention weight reflecting the relevance between the phrase PP and the feature vector ψk(In)\psi_{k}(I_{n}) . It is estimated by an attention function f_{att}\big{(}\psi_{k}(I_{n}),\theta^{l}(P)\big{)}, which first computes the the unnormalized weight rˉk(P,In)\bar{r}_{k}(P,I_{n}) with a linear projection over the joint representation of ψk(In)\psi_{k}(I_{n}) and rk(P,In)r_{k}(P,I_{n}):

then normalizes the values by using a softmax operation over all the KK bins:

In practice, the attention model is easy to overfit with limited training data. Besides, the spatially adjacent feature maps possibly represent one phrase, they are more reasonable to be merged. For these reasons, we reduce the training burden by average pooling the neighboring feature maps in Ψ(In)\Psi(I_{n}) before the weighted aggregation, which is also illustrated in Fig. 4.

4 Training and Testing

The final loss function is a combination of the image ID loss, the text ID loss as well as the discriminative and reconstructive image-language association losses:

where λT,λdis\lambda_{T},\lambda_{dis} and λrec\lambda_{rec} are balancing parameters. For network training, we adopt stochastic gradient descent (SGD) with an initial learning rate of 10−210^{-2}, which is further decayed to 10−310^{-3} after the 20th epoch. We organize the training batch as follows. The data tuple (In,Tn,dn)(I_{n},T_{n},d_{n}) is firstly transformed to (In,Tn,P(Tn),dn)(I_{n},T_{n},\mathcal{P}(T_{n}),d_{n}). Each batch contains the samples from 32 randomly selected persons, and each person has two randomly sampled tuples. For global discrimination, we form 32×432\times 4 positive image-description pairs by exploiting all the intra-tuple and inter-tuple image-description compositions, and sample 6 negative pairs for each image, yielding 64×664\times 6 negative pairs, keeping the pos/neg ratio to be 1:3. Meanwhile, the local reconstruction is performed within each tuple.

In testing, only image features are extracted, and no language descriptions are used. The distance between two image features are simply the Euclidean distance, i.e.,

Person Re-ID is performed by ranking the distances between the probe image and gallery images in ascending order.

Experiments

We evaluate the proposed approach on three standard person re-ID datasets, whose language annotations can be fully or partially obtained from the CUHK-PEDES dataset . Ablation studies are mainly conducted on Market-1501 and CUHK-SYSU , which are convenient for extensive evaluation as with fixed training/testing splits. We also report the overall results on Market-1501, CUHK03 and CUHK01 to compare with the state-of-the-art approaches.

Datasets and Metrics. To verify the utility of language descriptions in person re-ID, we augment four standard person re-ID datasets (Market-1501, CUHK03, CUHK01, and CUHK-SYSU) with language descriptions. The language descriptions are obtained from the CUHK-PEDES dataset, which is originally developed for cross-modal text-based person search and contains 40,206 images of 13,003 persons from five existing person re-ID datasets. Since persons in Market-1501 and CUHK03 have many similar samples, only four images of each person in this two datasets have language descriptions.

Among the four datasets, Market-1501 consists of 32,668 images of 1,501 persons and provides a standard protocol for training and testing. CUHK03 contains 13,164 images of 1,360 persons. Following , we use 1,260 persons for training and the rest 100 persons for testing. CUHK01 contains 971 persons captured from two views, and each person has two images in each view. 485 persons are randomly selected for training, and the remaining 486 persons are used for testing. CUHK-SYSU is a new dataset used for joint detection and identification. According to separation in CUHK-PEDES, 15,080 images from 5,532 identities are used for training, 8,341 images from 2,900 persons are used for testing with 2,900 query images and 5,441 gallery images. Mean average precision (mAP) and CMC top-1, top-5, top-10 accuracies are adopted as the evaluation metrics.

Implementation details. All the person images are resized to 256×\times128. For data augmentation, random horizontal flipping and random cropping are adopted. We empirically set the dimensions of feature embeddings ϕ(I)\phi(I), θl(P)\theta^{l}(P) and θg(T)\theta^{g}(T) to be 256256, and set the balancing parameters λT=0.1\lambda_{T}=0.1, λdis=1\lambda_{dis}=1, λrec=1\lambda_{rec}=1, respectively. As some images in Market-1501 and CUHK03 do not have language descriptions, we employ the description of the same person (in the same camera if possible) for them to compose the data tuple (In,Tn,dn)(I_{n},T_{n},d_{n}). The ResNet-50 backbone is initialized by the parameters pre-trained on ImageNet.

Baseline and variants. The baseline is just the visual CNN that produces the feature map ϕ(I)\phi(I), indicated by the red lines in Fig. 2. We additionally build 4 variants on the baseline for ablation study. The loss configuration of them are displayed in Table. 1. Among them, basel. only imposes the ID loss to make ϕ(I)\phi(I) be separable for different persons. Both basel.+rank and basel.+GDA additionally impose the ID loss over the global description feature θg(T)\theta^{g}(T) but have different global image-language association schemes. basel.+rank employs the Lrank\mathcal{L}_{rank} in Eqn.(8), while basel.+GDA utilizes the proposed Ldis\mathcal{L}_{dis} in Eqn.(7). The variant basel.+LRA employs the reconstruction loss Lrec\mathcal{L}_{rec} in Eqn.(13) to build the local association between the aggregated feature vector ψ^P(In)\hat{\psi}_{P}(I_{n}) and the phrase feature θl(P)\theta^{l}(P). Our proposed method takes advantages of both global and local image-language association schemes.

2 The Effect of Global Discriminative Association (GDA)

Comparison with non-discriminative variants. We evaluate the effects of global discriminative image-language association by comparing the variants with and without using the description feature θg(T)\theta^{g}(T). Among them, basel.+GDA improves basel. by 5.6% and 4.4% in term of mAP on Market-1501 and CUHK-SYSU respectively (Table. 2), which shows that GDA can benefit the learning of visual representation. Furthermore, our proposed method yields better performance than basel.+LRA, indicating the effect of global discriminative association is complementary to that of the local reconstructive association.

Comparison with bi-directional ranking loss . Ldis\mathcal{L}_{dis} in GDA aims to discriminate the matched image-text pairs from the unmatched ones. It has the similar functions with the bidirectional ranking loss Lrank\mathcal{L}_{rank} (Eqn. (8)) for image-language cross-modal retrieval. We implement two types of ranking losses for comparison. The first one is more similar to the loss in , where a positive image-text pair is composed of the image and text from the same tuple. The other one adopts the loss in , where the positive image-text pairs are obtained by arbitrary image-text combinations from the same person. We modify basel.+GDA by replacing Ldis\mathcal{L}_{dis} with the two loss functions, and denote them by basel.+rank1\emph{basel.+rank}^{1} and basel.+rank2\emph{basel.+rank}^{2}, respectively. The results in Table 2 show that both ranking losses can boost the baseline. Besides, basel. ⁣+ ⁣ rank2\emph{basel.\!+\! rank}^{2} is better than basel. ⁣+ ⁣ rank1\emph{basel.\!+\! rank}^{1} by incorporating more abundant positive samples for discrimination. The proposed basel.​+​ GDA further improves the mAP by 2.3% and 1.4% on Market-1501 and CUHK-SYSU, verifying the effectiveness of our relevance estimation strategy (Eqns. 5 and 6 ).

The importance of LT\mathcal{L}_{T}. To preserve separability of the visual feature, the associated linguistic feature θg(T)\theta^{g}(T) is supposed to be discriminative for different persons, thus LT\mathcal{L}_{T} is employed along with Ldis\mathcal{L}_{dis}. We investigate the importance of LT\mathcal{L}_{T} based upon basel.+GDA and observe how the performance changes with λT\lambda_{T} in Table 3. Slightly worse results are observed when λT=0\lambda_{T}=0, indicating LT\mathcal{L}_{T} is indispensable. On the other hand, the optimal results are achieved when λT\lambda_{T} is around 0.10.1. One possible reason is that language description is sometimes more ambiguous to describe a specific person, making LI\mathcal{L}_{I} and LT\mathcal{L}_{T} not equally important. For example, “The man wears a blue shirt” can simultaneously describe different persons wearing a dark blue shirt and a light blue shirt.

3 The Effect of Local Reconstructive Association (LRA)

Comparison with non-reconstructive variants. We evaluate the effects of local reconstructive association by comparing the variants with and without using the local phrase feature θl(P)\theta^{l}(P). The performance gap between basel. and basel.+LRA proves the effectiveness of LRA for visual feature learning. Employing LRA brings 5.2%5.2\% and 3.9%3.9\% mAP gain over the two datasets, which is close to the gain of employing GDA. Besides, the fact that the proposed method is better than basel.+GDA also indicates the effectiveness of LRA.

Visualization of phrase-guided attention weights. We compute the attention weights for a specific phrase (Eqn. (4)), align the weights to the corresponding image, and obtain the heat map for the phrase. The heat maps are displayed in Fig. 5, showing that the attention weights can roughly capture the local regions described by the phrases.

4 Results on Text-to-image Retrieval

As a by-product, our method can also be utilized for text-to-image retrieval, which is fulfilled by ranking the cross-modal relevance (Eqn. (6)). We report the retrieval results on CUHK-PEDES following the standard protocol, where there are 3,074 test images with 6,156 captions, 3,078 validation images with 6,158 captions, and 34,054 training images with 68,126 captions. The quantitative and qualitative results are reported in Table 4 and Fig. 6, respectively. Although our method is not specifically designed for this task, it achieves competitive results to the current state-of-the-art methods.

5 Comparison with the State-of-the-Art Approaches

We compare our method with the current state-of-the-arts on the Market1501, CUHK03, and CUHK01 datasets. The results on Market-1501 are reported in Table 6 left. Our method outperforms all the other approaches regarding mAP and top-1 accuracy under both single-query and multi-query protocols. Note that the baseline of our method is quite competitive to the most of the previous methods, which is partly because of well initialized ResNet-50 backbone and proper data augmentation strategies. The proposed image-language association scheme can largely boost the well-performed baseline, making our method better than the recent state-of-the-arts . CUHK03 has two types of person bounding boxes: one is manually labeled, and the other is obtained by a pedestrian detector. We compare our methods and others on both types, and report the top-1 and top-5 accuracies in Table 6 right. It can be seen that our method has significant advantages over the top-1 accuracy, but is 0.2% less than D-person on the top-5 accuracy for the labeled bounding boxes. As D-person only utilizes image data, it is promising to apply our language association scheme to D-person for better performance. Compared with Market-1501 and CUHK03, CUHK01 has fewer images for training as described in Sec. 5. As in Table 5, the proposed association schemes have 7.8% top-1 accuracy gain over the baseline on CUHK01. The results confirm the effectiveness of language description, and indicate the schemes may be more useful when the image data are not enough.

Among the compared approaches, Spindle and PDC utilize pose landmarks, CADL employs the camera ID labels, and ACN makes use of the attributes for training. We achieve better results than them on all the three datasets (Tables 5 and 6). The results indicate language description is also a kind of useful auxiliary information for person re-ID. With the proposed schemes, it can achieve the superior performance with the standard CNN architecture.

Conclusions

We utilized language descriptions as additional training supervisions to improve the visual features for person re-identification. The global and local image-language association schemes have been proposed. The former learns better global visual features with the discriminative supervision of the overall language descriptions, while the latter enforces the semantic consistencies between local visual features and noun phrases by phrase reconstruction. Our ablation studies show that the proposed image-language association schemes can remarkably improve the learning of the visual feature and are more effective than the existing image-text joint embedding methods. The proposed method achieves state-of-the-art performance on three public person re-ID datasets.

Acknowledgement. This work is supported by SenseTime Group Limited, the General Research Fund sponsored by the Research Grants Council of Hong Kong (Nos. CUHK14213616, CUHK14206114, CUHK14205615, CUHK14203015, CUHK14239816, CUHK419412, CUHK14207814, CUHK14208417, CUHK14202217), the Hong Kong Innovation and Technology Support Program (No.ITS/121/15FX).

References