Unifying Identification and Context Learning for Person Recognition

Qingqiu Huang, Yu Xiong, Dahua Lin

Introduction

Person recognition is a key task in computer vision and has been extensively studied over the past decades. Thanks to the advances in deep learning, recent years have witnessed remarkable progress in face recognition techniques . On LFW , a challenging public benchmark, the accuracy has been pushed to over 99.8%99.8\% . Nonetheless, the success on benchmarks does not mean that the problem has been well solved. Recent studies have shown that recognizing persons under an unconstrained setting remains very challenging. Substantial difficulties arise in unfavorable conditions, e.g. when the faces are in a non-frontal position, subject to extreme lighting, or too far away from the camera. Such conditions are very common in practice.

The difficulties above are essentially due to the fact that facial appearance is highly sensitive to environmental conditions. To tackle this problem, a natural idea is to leverage another important source of information, namely the context. It is our common experience that we can easily recognize a familiar person by looking at the wearing, the surrounding environment, or the people who are nearby. On the other hand, cognitive neuroscience studies have shown that context plays a crucial role when we, as human beings, recognize a person or an object. A familiar context often allows much greater accuracy in recognition.

Exploiting context to help recognition is not a new story in computer vision. Previous efforts mainly follow two lines. The first line of research attempts to incorporate additional visual cues, e.g. clothes and hairstyles, as additional channels of features. The other line, instead, focuses on social relationships, e.g. group priors or people co-occurrence . There are also studies that try to integrate both visual cues and social relations .

Whereas previous works have shown the utility of context in person recognition, especially in unconstrained environments, a key question remains open, that is, how to discover and leverage contexts robustly. Specifically, existing methods usually rely on simple heuristics, e.g. feature-based clustering, to establish contextual priors, and hand-crafted rules to combine contextual cues from different channels. Moreover, the construction of the context model is typically done separately and before person identification. The limitations of such approaches lie in two important aspects: (1) Heuristics designed manually are difficult to capture the diversity and complexity in unconstrained context modeling. (2) The identities of the people in a scene are also an important part of the context. Constructing the context separately would lose this significant connection.

In this work, we aim to explore more effective ways to leverage the context in person recognition (see Figure 1). Inspired by previous efforts, we consider two kinds of contexts, namely the visual context, e.g. additional visual cues, and the social context, e.g. the events that a person often attends. But, we move beyond the limitations of existing methods, by considering context learning and person identification as a unified process and solving both jointly. Driven by this idea, we propose novel methods for leveraging visual and social contexts respectively. Particularly, we develop a Region Attention Network, which is learned end-to-end to combine various visual cues adaptively with instance-dependent weights. We also develop a unified formulation, where the social context model is learned online jointly with the reasoning of people identities. As a by-product, the solution to this problem also comes with a set of “events” discovered from the given photo collection – an event not only share similar scenes but also a consistent set of attendants.

On PIPA , a large public benchmark for person recognition, our proposed method consistently outperform existing methods, under all evaluation policies. Particularly, in the most challenging day split, our method raised the state-of-the-art performance from 59.77%59.77\% to 67.16%67.16\%. To assess our method in more diverse settings and to promote future research on this topic, we construct another large dataset, Cast In Movies (CIM), by annotating the characters in 192192 movies. This dataset contains more than 150K150K person instances and 12181218 labeled identities. Our approach also demonstrated its effectiveness on CIM.

Our contributions mainly lie in three aspects: (1) For visual context, we propose a Region Attention Network, which combines visual cues with instance-dependent weights. (2) For social context, We propose a unified formulation that couples context learning with people identification. It also discovers events from photo collections automatically. These two techniques together result in remarkable performance gains over the state-of-the-art. (3) We construct Cast In Movies (CIM), a large and diverse dataset for person recognition research.

Related Work

The significance of context in person recognition has long been recognized by the vision community. Early methods mainly tried to use additional visual cues, such as clothing , or additional metadata . Yet, the improvement was limited. Later, more sophisticated frameworks that integrate multiple cues (clothing, timestamps, scenes, etc) have been developed. Some of these works formulated the task as a joint inference process over a Markov random field and obtained further performance gains. Note that these MRF-based methods assume the same set of people and social relations in both training and testing, and thus the learned models are difficult to generalize to new collections.

Recent efforts.

The rise of deep learning has led to new innovations on this topic. Zhang et al. proposed a Pose Invariant Person Recognition method (PIPER), which combines three types of visual recognizers based on ConvNets, respectively on face, full body, and poselet-level cues. The PIPA dataset published in has been widely adopted as a standard benchmark to evaluate person recognition methods. Oh et al. evaluated the effectiveness of different body regions, and used a weighted combination of the scores obtained from different regions for recognition.

Recently, Li et al. proposed a multi-level contextual model, which integrates person-level, photo-level and group-level contexts. Although this framework also considers the combination of visual cues and social context, it differs from ours essentially in two key aspects: (1) The visual cues are combined with a simple heuristic rule, instead of a learned network. (2) The groups are identified by spectral clustering of scene features before person recognition, as a separate step. Our framework, instead, formulates event discovery and people identification as a unified optimization problem and solves both jointly.

Another way of integrating both visual cues and social relations was proposed in . This work formulates the recognition of multiple people into a sequence prediction problem and tries to capture the relational cues with a recurrent network. As there is no inherent order among the people in a scene, it is unclear how a sequential model can capture their relations. Note that we compared the proposed method with all of the four methods above in our experiments on PIPA. As we shall see in Section 5, our method consistently outperforms them under all evaluation policies.

Person Re-identification.

Another relevant task is person re-identification , which is to match pedestrian images from different cameras, within a relatively short period. This task is essentially different, where visual cues are likely to remain consistent and social context is weak. General person recognition, instead, requires recognizing across events, where visual cues may vary significantly and thus the social context can be crucial.

Methodology

In general, the task of person recognition in a photo collection can be formalized as follows. Consider a collection of photos I1,…,IMI_{1},\ldots,I_{M}, where ImI_{m} contains NmN_{m} person instances. All person instances are divided into two disjoint subsets, the gallery set, in which the instances are all labeled (i.e. their identities are provided). and the query set, in which the instances are unlabeled. The task is to predict the identities for those instances in the query set.

As discussed, person recognition in an unconstrained setting is very challenging. In this work, we leverage two kinds of contexts, the visual context and the social context. Particularly, the visual context involves different regions of the person instances, including face, head, upper body, and whole body. These regions often convey complementary cues for visual matching. The social context, instead, captures the social behavior of people, e.g. the events they usually attend or the people whom they often stay with. It is worth noting that unlike visual cues, the social relations are reflected collectively by multiple photos and can not be reliably derived from a single photo in isolation.

We devise a framework that incorporates both the visual context and the social context for person recognition. As shown in Figure 2, the framework recognizes the identities for all instances in the query set jointly, in two stages.

Visual matching. This stage computes a matching score for each pair of instances. For this, a Region Attention Network is learned to adaptively combine the visual cues from different regions, with instance-dependent weights.

Joint optimization. This is the key stage of our framework. In this stage, the social context model, which captures both event-people and people-people relations, will be jointly learned along with the identification of query instances, by solving a unified optimization problem.

2 Visual Matching

We combine the visual observations from different regions to compute the matching score between two instances. Particularly, we consider four regions: face, head, upper body, and whole body. These regions are often complementary to each other. This strategy has also been shown to be effective in previous work .

However, existing methods mostly adopt uniform weighting schemes, where each region is assigned a fixed weight that is shared by all instances. Let s(i,j)s(i,j) be the overall matching score between instances ii and jj, and sr(i,j)s^{r}(i,j) be the matching score based on the rr-th region. Then, such a scheme can generally be expressed as

where RR is the number of distinct regions. The weights {wr}\{w^{r}\} are often decided by empirical rules or optimized over a validation set .

The uniform schemes as described above are limited in two aspects, as illustrated in Figure 3. (1) Some regions may be invisible for an instance. The missing of such regions may be due to various reasons, e.g. limited scope of the camera and occlusion. With a uniform scheme, one would be forced to locate the missing parts with rigid rules and compute matching scores for them, which often leads to inaccurate results. (2) The contributions of different parts vary significantly across instances. For example, the facial features play a key role when the frontal face is visible. However, when we can only see one’s back, we will have to resort to the clothing in the body region. A uniform scheme can not effectively handle such variations.

We propose to tackle this problem using instance-dependent weights, where the weight of a region is determined by whether it is visible and how much it contributes. Specifically, given an instance, we get the bounding boxes of the regions by either the annotation of the dataset or detectors, then resize each region to a standard size, and apply region-specific CNNs to extract their features.

To combine these features adaptively, we devise a Region Attention Network (RANet) as shown in Figure 2 to compute the fusion weights. Here, the RANet is a small neural network that takes the stacked features from all regions as input, feeds them through a convolution layer, a fully-connected layer, and a sigmoid layer, and finally yields four positive coefficients as the region weights. Then the combined matching score is given by

Here, wirw_{i}^{r} and wjrw_{j}^{r} are instance-dependent weights of the rr-th region respectively for instances ii and jj; sr(i,j)s^{r}(i,j) is the cosine similarity between the corresponding features. We use the product wirwjrw_{i}^{r}w_{j}^{r} to weight a region score, which reflects the rationale that a region type should be active only when it is clearly visible in both instances. All region-specific CNNs together with the RANet are jointly trained in an end-to-end manner with the cross-entropy loss.

3 Unified Formulation with Social Context

In an unconstrained environment, certain instances are very difficult to recognize purely based on their appearance. For such cases, one can leverage the social context to help. Specifically, the social context refers to a set of social relations. We consider two types of social relations:

Event-person relations. Generally, an event can be conceptually understood as an activity that occurs at a certain place with a certain set of attendants . Over a large photo collection, an event may involve just a small fraction of the people. Hence, an event can provide a strong prior for recognition if we can infer the event that a photo is capturing, as illustrated by the photo in Figure 4a.

Person-person relations. It is commonly observed that certain groups of people often stay together. For such groups, the presence of a person may indicate the presence of others in the group, as illustrated by the photo in Figure 4b. Note that person-person relations are complementary to the event-person relations, as such relations do not require the match of scene features.

Taking both the visual context and the social context into account, we can formulate a unified optimization problem where person identifications are coupled with event association and contextual relation learning. The objective function of this problem is given below:

The notations involved here are described below:

Among these quantities, the matching scores S\mathbf{S} and the scene features F\mathbf{F} are provided in the visual analysis stage, while others are jointly solved by optimizing this problem.

Potential Terms.

The joint objective in Eq.(3) comprises three potential terms, which are introduced below.

Visual consistency: ψv(X∣S)\psi_{v}(\mathbf{X}|\mathbf{S}) encourages the consistency between person identities and the visual matching scores, and is formulated as:

Event consistency: ϕep(Y,X;F~,P∣F)\phi_{ep}(\mathbf{Y},\mathbf{X};\widetilde{\mathbf{F}},\mathbf{P}|\mathbf{F}) concerns about the assignments of photos to events, and encourages them to be consistent in both scenes and attendants. This term is formulated as:

People cooccurrence: ϕpp(X;Q)\phi_{pp}(\mathbf{X};\mathbf{Q}) takes into account the person-person relations, i.e. which identities tend to coexist in a photo. This term is formulated as:

This formula considers all pairs of distinct instances in each image, and sums up their person-person relation value. In particular, if xj\mathbf{x}_{j} indicates label ll and xj′\mathbf{x}_{j^{\prime}} indicates label l′l^{\prime}, then xjTQxj′=Q(l,l′)\mathbf{x}_{j}^{T}\mathbf{Q}\mathbf{x}_{j^{\prime}}=Q(l,l^{\prime}).

To balance the contributions of these potential terms, we introduce two coefficients α\alpha and β\beta, which are decided via cross validation.

4 Joint Estimation and Inference

We solve this problem using coordinate ascent. Specifically, our algorithm alternates between the updates of (1) people identities (X\mathbf{X}), (2) assignments of events to photos (Y\mathbf{Y}), and (3) social relation parameters (F~\widetilde{\mathbf{F}}, P\mathbf{P}, and Q\mathbf{Q}). These steps are presented below.

Given both the event assignments X\mathbf{X} and the social context parameters, the inference of people identities can be done separately for each photo, by maximizing the sub-objective as:

where y^i\hat{y}_{i} indicates the assigned event. Note that xj\mathbf{x}_{j} here is constrained to be an indicator vector, i.e. only one of its entry is set to one, while others are zeros. When there is only one person instance, its identity can be readily derived as

When there are two or more instances, we treat it as an MRF over their identities and solve them jointly using the max-product algorithm.

Event Assignment.

We found that the granularity of the events has significant impact on the identification performance. If we group the photos into coarse-grained events such that each event may contain many people or scenes, then the event-person relations may not be able to provide a strong prior. However, for fine-grained events, it would be difficult to estimate their parameters reliably. Hence, it is advisable to seek a good balance.

In this work, we use two parameters νmin\nu_{min} and νmax\nu_{max} to control the granularity, and require that the number of photos assigned to an event fall in the range [νmin,νmax][\nu_{min},\nu_{max}]. Then the problem of event assignment can be written as

Here, aika_{i}^{k} Eq.(9) follows Eq.(2). Eq.(10) enforces the constraint that each photo is associated to at most one event; Eq.(11) enforces the granularity constraint above. This is a linear programming problem, and can be readily solved by an LP solver. Also, the optima is guaranteed to be integral.

Context Learning.

As mentioned, the social context model, which is associated with three parameters F~\widetilde{\mathbf{F}}, P\mathbf{P}, and Q\mathbf{Q}, are learned along the inference of people identities and event assignments. Given X\mathbf{X} and Y\mathbf{Y}, we can easily derive the optimal solution of the parameters listed above.

where Ek={i∣aik=1}\mathcal{E}_{k}=\{i\mid a_{i}^{k}=1\} refers to the set of photos that are assigned to the kk-th event. For the identity distributions P\mathbf{P}, we have the optimal pk\mathbf{p}_{k} given (the kk-th column of P\mathbf{P}) given by

It is worth emphasizing that all sub-tasks presented above are steps in the coordinate ascent procedure to optimize the unified objective in Eq.(3). We run these steps iteratively, and it usually takes about 55 iterations to converge.

New Dataset: Cast In Movies

In addition to photo albums, the proposed method can also be applied to other settings with strong contexts, e.g. recognizing actors in movies. To test our method in such settings, we constructed the Cast In Movies (CIM) dataset from 192192 movies. We divide each movie into shots using an existing technique , sample one frame from each shot, and retain all those that contain persons. This procedure results in a dataset with 72,87572,875 photos.

We manually annotated all person instances in these photos with bounding boxes for the body locations. We also annotated the identities of those instances that correspond to the 12181218 main actorsThe main actors are chosen according to two criteria: 1) ranked top 10 in the cast list of IMDb for the corresponding movie, and 2) occur for more than 5 times in our sampled frames.. In this way, we obtained 77,59877,598 instances with known identities, while other instances are labeled as “others”. Figure 5 shows some examples of our dataset. Table 1 shows the statistics of CIM in comparison with PIPA . To our best knowledge, CIM is the first large-scale dataset for person recognition in movies.

Experiments

We tested our method on both PIPA , a dataset widely used for person recognition, and CIM, our new dataset presented above.

The PIPA dataset is partitioned into three disjoint sets: training, validation and test sets. The test set is further split into two subsets, one as the gallery set and the other as the query set. To evaluate a method’s performance, we first use it to predict the identities of the instances in the query set and compute the prediction accuracy. Then, we switch the gallery and the query set and compute the accuracy in the same way. The average of both accuracies will be reported as the performance metric.

There are four different ways to split the test set, namely original, album, time, and day, for evaluating an algorithm under different application scenarios. In the original setup of PIPA , a query instance may have similar instances in the gallery. defines the other three splits, which are more challenging. For example, day split requires that the query and the gallery instances of the same subject need to have notable differences in visual appearance. For CIM, we follow the rule in , dividing it into three disjoint subsets respectively for training, validation, and testing. Also, the test set is randomly split into a gallery set and a query set.

Implementation Details

We use four regions of each instance: face, head, upperbody, and body. PIPA provides the head locations, while CIM provides the locations of whole body. Other regions are obtained by simple geometric rules based on the results from a face detector and OpenPose . Note that we only keep those bounding boxes that lie mostly within the photo. For those regions that are largely invisible, we simply use a black image to represent their appearance. We will see that our RANet can learn to assign such regions with negligible weights in our experiments. We adopt ResNet-50 as our base model and train the feature extractor with OIM loss . We chose design parameters empirically on the validation set. The coefficients α\alpha and β\beta in Eq.(3) are set to 0.050.05 and 0.010.01. The number of events KK is set to 300300 for both PIPA and CIM.

2 Results on PIPA

We set up a baseline for comparison, which relies on a uniformly weighted combination of visual cues from all regions, where the weights are optimized by grid search. We tested three configurations of the proposed methods: (1) +RANet: This config combines region-specific scores following Eq.(2), using the instance-dependent weights from the Region Attention Network (see Sec. 3.2). (2) +RANet+P: In addition to the visual matching score RANet, it also uses the person-person relations in joint inference. (3) +RANet+P+E: This is the full configuration of our framework, which takes visual matching, person-person relations, and person-event relations into account. Moreover, we also compared with four previous methods: PIPER , Naeil , RNN , and MLC .

Table 2 shows the results under all the four splits, from which we can see that: (1) RANet, with adaptive weights, significantly outperforms the baseline with uniform weights. On the most challenging day split, it remarkably raises the accuracy from 47.09%47.09\% to 65.49%65.49\%. (2) With our proposed joint inference method, the use of social contexts leads to consistent improvement across all splits. (3) Our method also outperforms all previous works, including the state-of-the-art MLC , by a considerable margin on all splits. Particularly, the performance gain is especially remarkable on the most challenging day split (67.16%67.16\% with ours vs. 59.77%59.77\% with MLC).

Table 2 clearly shows the effectiveness of the Region Attention Network (RANet). To learn more about the RANet, we study the distributions of region-specific weights on the test set of PIPA, and show them in Figure 6. This study reveals some interesting characteristics of RANet: (1) For each of the following region types, face, body, and upper body, there exist a fraction of instances with very low weights because the particular regions of them are out of scope. (2) The average weight of faces is the highest among all region types. This is not surprising, as faces are often the strongest indicators of identities when they are visible. (3) A small portion of instances have very high weights assigned to the head regions, because for such instances all other parts are largely invisible.

Analysis on Event

Events are automatically discovered during joint inference and they play an important role in person recognition. Figure 7 shows some example events with their associated photos. We can see that our method can discover events in a reasonable way, and they can provide strong prior in a considerable portion of cases. More examples will be provided in the supplemental materials.

Case Study

Figure 8 shows some photos and associated recognition results. We can see that 1) For an instance with frontal and clear face (1st row), all methods predict correctly. 2) When the face is not clearly visible (2nd row), our method with RANet can still correctly recognize the person with other visual cues, e.g. the head or upper body. 3) For the most challenging case where all visual cues fail (3rd row), our full model can still make a correct prediction by exploiting the social context.

3 Results on CIM

Table 3 shows the results on CIM, which again demonstrates the effectiveness of our approach. Only with RANet, it already outperforms the baseline (with uniform weighting) by over 4%4\%. The whole framework, with social context taken into account, further improves the accuracy (6.3%6.3\% higher than the baseline). Recognition results on example photos will be provided in the supplemental materials. It is also worth noting that the accuracies we obtained on CIM are generally lower than those on PIPA, implying that this is a more challenging dataset which can help to drive the progress on this task.

4 Computational Cost Analysis

Our method obtains the improvement on recognition accuracy with substantially lower computing cost compared to some previous works. Note that PIPER uses more than 100 deep CNNs and Naeil uses 17 deep CNNs for feature extraction. While our model uses only 4 CNNs and a fusion module whose computing cost is negligible. Although MLC uses just 3 deep CNNs for feature extraction, it additionally requires to train thousands of group-specific SVMs, which is also a costly procedure.

Compared with the CNN-based feature extraction components, the cost of the joint estimation and inference procedure is insignificant. Particularly, it takes about 30 minutes to perform inference over the whole test set of PIPA, with one single 2.22.2 GHz CPU, while the feature extractors take over 4040 hours to detect regions and compute CNN features for all test photos, with a Titan X GPU.

Conclusions

We presented a new framework for person recognition, which integrates a Region Attention Network to adaptively combine visual cues and a model that unifies person identification and context learning in joint inference. We conducted experiments on both PIPA and a new dataset CIM constructed from movies. On PIPA, our method consistently outperformed previous state-of-the-art methods by a notable margin, under all splits. On CIM, the new components developed in this work also demonstrated strong effectiveness in raising the recognition accuracy. Both quantitative and qualitative studies showed that adaptive combination of visual cues is important in a generic context and that the social context often conveys useful information especially when the visual appearance causes ambiguities.

Acknowledgement

This work is partially supported by the Big Data Collaboration Research grant from SenseTime Group (CUHK Agreement No. TS1610626), the General Research Fund (GRF) of Hong Kong (No. 14236516). We are grateful to Shuang Li and Hongsheng Li for helpful discussions.

References