Learning Social Relation Traits from Face Images

Zhanpeng Zhang, Ping Luo, Chen Change Loy, Xiaoou Tang

Introduction

Social relation manifests when we establish, reciprocate, or deepen relationships with one another in either physical or virtual world. Studies have shown that implicit social relations can be discovered from texts and microblogs . Images and videos are becoming the mainstream medium to share information, which capture individuals with different social connections. Effectively exploiting such socially-rich sources can provide social facts other than the conventional medium like text (Fig. 1).

The aim of this study is to characterise and quantify social relation traits from computer vision point of view. Inspired by extensive psychological studies , which show that face emotional expressions can serve as social predictive functions, we wish to automatically recognise fine-grained and high-level social relation traits (e.g., friendliness, warm, and dominance) from face images. Such a capability promises a wide spectrum of applications. For instance, automatic social relation inference allows for relation mining from image collection in social network, personal album, and films.

Profiling unscripted social relation from face images is non-trivial. Among the most significant challenges are: (1) as suggested by psychological studies , relations of face images are related to high-level facial factors. Thus we need a rich face representation that captures various attributes such as expression and head pose; (2) no single dataset is presently available, which encompasses all the required facial attribute annotations to learn such a rich representation. In particular, some datasets only contain face expression labels, whilst other datasets may only contain the gender label. Moreover, these datasets are collected from different environments and exhibit different statistical distributions. How to effectively train a model on such heterogeneous data remains an open problem.

To this end, we carefully formulate a deep model to learn a face representation for social relation prediction, driven by rich facial attributes such as expression, head pose, gender, and age. We devise a new deep architecture that is capable of (1) dealing with missing attribute labels from different datasets, and (2) bridging the gap of heterogeneous datasets by weak constraints derived from the association of face part appearances. This allows the model to learn more effectively from heterogeneous datasets with different annotations and statistical distributions. Unlike existing face analyses that mostly consider single subject, our network is formulated with a Siamese-like architecture , it is thus capable of jointly considering pairwise faces for relation reasoning, where each face serves as the mutual context to the other.

The contributions of this study are three-fold: (1) to our knowledge, this is the first work that investigates face-driven social relation inference, of which the relation traits are defined based on psychological study . We carefully investigate the detectability and quantification of such traits from a pair of face images. (2) we carefully construct a new social relation dataset labeled with pairwise relation traits supported by psychological studies , which can facilitate future research on high-level face interpretation. (3) we formulate a new deep architecture for learning face representation driven by multiple tasks, bridging the gap from heterogeneous sources with potentially missing target attribute labels. It is also demonstrated that the model can be extended to utilize additional cues such as the faces’ relative location, besides face images.

Related Work

Social signal processing. Understanding social relation is an important research topic in social signal processing , an important multidisciplinary problem that has attracted a surge of interest from computer vision community. Social signal processing mainly involves facial expression recognition and affective behaviour analysis . On the other hand, there exists a number of studies that aim to infer social relation from images and videos . Many of these studies focus on the coarser level of social connection other than the one defined by Kiesler in the interpersonal circle . For instance, Ding and Yilmaz only discover social group without inferring relation between individuals. Fathi et al. only detect three social interaction classes, i.e., ‘dialogue, monologue and discussion’. Wang et al. define social relation by several social roles, such as ‘father-child’ and ‘husband-wife’. Other related problems also include image communicative intents prediction and social role inference , usually applied on news and talks shows , or meetings to infer dominance .

Our work differs significantly from the aforementioned studies. Firstly, most affective analysis approaches are based on single person therefore cannot be directly employed for interpersonal relation inference. In addition, these studies mostly focus on recognizing prototypical expressions (happy, angry, sad, disgust, surprise, fear). Social relation is far more complex involving many factors such as age and gender. Thus, we need to consider more attributes jointly in our problem. Secondly, in comparison to the existing social relation studies , our work aims to recognize fine-grained and high-level social relation traits . Thirdly, many of the social relation studies did not use face images directly for relation inference, but visual concepts discovered by detectors or people spatial proximity in 2D or 3D space . All these information sources are valuable for learning human interactions but social relation is fundamentally limited by the input sources.

Human interaction and group behavior analysis. Existing group behavior studies mainly recognize action-oriented behaviors such as hugging, handshaking or walking, but not social relations. Often, group spatial configuration and actions are exploited for the recognition. Our study differs in that we aim to recognize abstract relation traits from faces.

Deep learning. Deep learning has achieved remarkable success in many tasks of face analysis, e.g. face parsing , face landmark detection , face attribute prediction , and face recognition . However, deep learning has not yet been adopted for face-driven social relation mining that requires joint reasoning from multiple subjects. In this work, we propose a deep model to cope with complex facial attributes from heterogeneous datasets, and joint learning from face pair.

Social Relation Prediction from Face Images

We define the social relation traits based on the interpersonal circle proposed by Kiesler , where human relations are divided into 16 segments as shown in Fig. 2. Each segment has its opposite side in the circle, such as “friendly and hostile”. Therefore, the 16 segments can be considered as eight binary relations, whose descriptions and examples are given in Table 1. More detailed descriptions are provided in the supplementary material. We also provide positive and negative visual samples for each relation in Fig. 2, showing that they are visually perceptible. For instance, “friendly” and “competitive” are easily separable because of the conflicting meanings. However, some relations are close such as “friendly” and “trusting”, implying that a pair of faces can have more than one social relation.

2 Social Relation Dataset

To investigate the detectability of social relations from a pair of face images, we build a new datasethttp://mmlab.ie.cuhk.edu.hk/projects/socialrelation/index.html, containing 8,3068,306 images chosen from web and movies. Each image is labelled with faces’ bounding boxes and their pairwise relations. This is the first face dataset measuring social relation traits and it is challenging because of large face variations including poses, occlusions, and illuminations.

We carefully built this dataset. Five performing arts students were asked to label each relation for each face image independently. Thus, each label has five annotations. A label is accepted if more than three annotations are consistent. The inconsistent samples were presented again to the five annotators to seek consensusThe average Fleiss’ kappa of the eight relation traits’ annotation is 0.62, indicating substantial inter-rater agreement.. To facilitate the annotation task, we also provide multiple cues to the annotators. First, to help them understand the social relations, we list ten related adjectives defined by for the positive and negative samples on each relation trait, respectively. Multiple example images are also provided. Second, for the image frames selected from the movies, the annotators were asked to get familiar with the stories. The subtitles were presented during labelling.

3 Baseline Method

To improve the baseline method, we incorporate some spatial cues to train the deep network as shown in Fig.3(a), which includes 1) two faces’ positions {xl,yl,wl,hl,xr,yr,wr,hr}\{x^{l},y^{l},w^{l},h^{l},x^{r},y^{r},w^{r},h^{r}\}, representing the xx-,yy-coordinates of the upper-left corner, width, and height of the bounding boxes; wlw^{l} and wrw^{r} are normalized by the image width. Similar for hlh^{l} and hrh^{r}; 2) the relative faces’ positions: xl−xrwl,yl−yrhl\frac{x^{l}-x^{r}}{w^{l}},\frac{y^{l}-y^{r}}{h^{l}}, and 3) the ratio between the faces’ scales: wlwr\frac{w^{l}}{w^{r}}. The above spatial cues are concatenated as a vector, xs\textbf{x}_{s}, and combined with the shared representation xt\textbf{x}_{t} for learning relation traits.

As the above description, each binary variable gig_{i} can be predicted by linear regression,

where ϵ\epsilon is an additive error random variable, which is distributed following a standard logistic distribution, ϵ∼Logistic(0,1)\epsilon\sim Logistic(0,1). [⋅;⋅][\cdot;\cdot] indicates the column-wise concatenation of two vectors. Therefore, the probability of gig_{i} given xt\textbf{x}_{t} and xs\textbf{x}_{s} can be written as a sigmoid function, p(gi=1∣xt,xs)=1/(1+exp⁡{−wgiT[xs;xt]})p(g_{i}=1|\textbf{x}_{t},\textbf{x}_{s})=1/(1+\exp\{-{\textbf{w}}^{\mathsf{T}}_{g_{i}}[\textbf{x}_{s};\textbf{x}_{t}]\}), indicating that p(gi∣xt,xs)p(g_{i}|\textbf{x}_{t},\textbf{x}_{s}) is a Bernoulli distribution, p(g_{i}|\textbf{x}_{t},\textbf{x}_{s})=p(g_{i}=1|\textbf{x}_{t},\textbf{x}_{s})^{g_{i}}\big{(}1-p(g_{i}=1|\textbf{x}_{t},\textbf{x}_{s})\big{)}^{1-g_{i}}.

Combining the above probabilistic definitions, the deep network is trained by maximising a posterior probability,

where Ω={{wgi}i=18,W,Kl,Kr}\Omega=\{\{\textbf{w}_{g_{i}}\}_{i=1}^{8},\textbf{W},\textbf{K}^{l},\textbf{K}^{r}\} and the constraint means the filters are tied. Note that xt\textbf{x}_{t} and xs\textbf{x}_{s} represent the hidden features and the spatial cues extracted from the left and right face images, respectively. Thus, the variable gig_{i} is independent with Il\textbf{I}^{l} and Ir\textbf{I}^{r}, given xt\textbf{x}_{t} and xs\textbf{x}_{s}.

By taking the negative logarithm of Eqn.(2), it is equivalent to minimising the following loss function

where the second and the third terms correspond to the traditional cross-entropy loss, while the remaining terms indicate the weight decays of the parameters. Eqn.(3) is defined over single training sample and is a highly nonlinear function because of the hidden features xt\textbf{x}_{t}. It can be efficiently solved by stochastic gradient descent .

4 A Cross-Dataset Approach

As investigated by the psychological studies , the social relations of face images are strongly related to some hidden high-level factors, such as emotion. Learning these semantic concepts implicitly from raw image pixels imposes great challenge. To explicitly learn these factors, an ideal solution is to introduce two additional loss functions on top of xl\textbf{x}^{l} and xr\textbf{x}^{r} respectively, representing that not only the concatenation of xl\textbf{x}^{l} and xr\textbf{x}^{r} learns the relation traits, but each of them also learns the high-level factors of its corresponding face image. However, this solution is impractical, because labelling both social relations and emotions of face images is too expensive.

To overcome this limitation, we extend the baseline model by pre-training the DCN with face attributes, which are borrowed from existing face databases. These attributes capture the high-level factors, guiding the predictions of relation traits. The advantages are three folds: 1) face attributes, such as age, gender, and expressions, are highly correlated with the high-level factors of social relations, as supported by the psychological studies ; 2) leveraging the existing face databases not only improves generalized capacity but also make data preparation much easier; and 3) the face representation induced by semantic attributes can bridge the gap between the high-level relation traits and low-level image pixels.

In particular, we make use of data from three public datasets, including AFLW , CelebFaces , and Kaggle . Different datasets have been labelled with different sets of face attributes. A summary is given in Table 2, where the attributes are partitioned into four groups.

It is clear that the training datasets are from multiple heterogenous sources and they have been labelled with different sets of attributes. For instance, AFLW only contains gender and poses, while Kaggle only has expressions. In addition, these datasets exhibit different statistical distributions, causing issues during pre-training. It can be shown that if we perform joint training directly, each attribute is trained by the labelled data alone, instead of benefitting from the existence of the unlabelled data. Consider a simple example of three datasets, denoted as AA, BB, and CC, where AA and BB are labelled with attribute y1y^{1} and y2y^{2} respectively, while dataset CC is labelled with y1y^{1}, y2y^{2} and y3y^{3}. Moreover, xA\textbf{x}_{A} indicates a training sample from dataset AA. Given three training samples xA\textbf{x}_{A}, xB\textbf{x}_{B}, and xC\textbf{x}_{C}, attribute classification is to maximise the joint probability p(yA1,yA2,yA3,p(y^{1}_{A},y^{2}_{A},y^{3}_{A}, yB1,yB2,yB3,y^{1}_{B},y^{2}_{B},y^{3}_{B}, yC1,yC2,yC3∣xA,xB,xC)y^{1}_{C},y^{2}_{C},y^{3}_{C}|\textbf{x}_{A},\textbf{x}_{B},\textbf{x}_{C}). Since the samples are independent and AA and BB only contain attributes y1y^{1} and y2y^{2} respectively, the joint probability can be factorized as p(yA1,yA2,yA3∣xA)p(y^{1}_{A},y^{2}_{A},y^{3}_{A}|\textbf{x}_{A}) ⋅\cdot p(yB1,yB2,yB3∣xB)p(y^{1}_{B},y^{2}_{B},y^{3}_{B}|\textbf{x}_{B}) ⋅\cdot p(yC1,yC2,yC3∣xC)p(y^{1}_{C},y^{2}_{C},y^{3}_{C}|\textbf{x}_{C}) == p(yA1∣xA)p(y^{1}_{A}|\textbf{x}_{A}) ⋅\cdot p(yB2∣xB)p(y^{2}_{B}|\textbf{x}_{B}) ⋅\cdot p(yC1,yC2,yC3∣xC)p(y^{1}_{C},y^{2}_{C},y^{3}_{C}|\textbf{x}_{C}). For example, we have ∑yA2,yA3p(yA1,yA2,yA3∣xA)\sum_{y^{2}_{A},y^{3}_{A}}p(y^{1}_{A},y^{2}_{A},y^{3}_{A}|\textbf{x}_{A}) == p(yA1∣xA)p(y^{1}_{A}|\textbf{x}_{A}). As the attributes are also independent, the joint probability can be further written as p(yA1,yC1∣xA,xC)p(yB2,yC2∣xB,xC)p(yC3∣xC)p(y^{1}_{A},y^{1}_{C}|\textbf{x}_{A},\textbf{x}_{C})p(y^{2}_{B},y^{2}_{C}|\textbf{x}_{B},\textbf{x}_{C})p(y^{3}_{C}|\textbf{x}_{C}), indicating that each attribute classifier is trained by the labelled data alone. For instance, the classifier of the first attribute is trained by data from AA and CC.

Bridging the gaps between multiple datasets. Since faces from different datasets share similar structure in local part, such as mouth and eyes, we propose a bridging layer based on the local correspondence to cope with the different dataset distributions. In particular, we establish a face descriptor hh based on the mixture of aligned facial parts. As shown in Fig. 3(b), we build a three-level hierarchy to partition the facial parts’ shape, where each child node groups the data of its parents into clusters, such as u2,11u^{1}_{2,1} and u2,101u^{1}_{2,10}. In the top layer, the faces are divided into 10 clusters by K-means using the landmark locations from the SDM face alignment algorithm . Each cluster captures the topological changes due to viewpoints. Fig. 3(b) shows the mean face of each cluster. In the second layer, for each node, we perform K-means using the locations of landmarks in the upper and lower face region, and obtain 10 clusters respectively. These clusters captures the local shape of the facial parts. Then the mean HOG feature of the faces in each cluster is regarded as the corresponding template. Given a new sample, the descriptor hh is obtained by concatenating its L2-distance to each template.

In this case, the descriptor hh serves as a correspondence label for datasets. We use it as additional input in the fully connected layer for facial feature x (see Fig.3(b)). Thus the learned face representations for samples from different datasets are driven to be close if the correspondence labels are similar. It is worth noting that this bridging layer is different from the work of , where the algorithms build some clusters from training data as an auxiliary task. Differently, the proposed method uses the aligned facial part association, which is well suited for our problem, instead of simply construct the cluster from the whole image. Moreover, since the construction of hh is unsupervised, it contains noise and may harm the training if used as targets. Instead, we use the descriptor as additional input, which shows better performance than used as output (see Table. 5). The rest of the DCN structure is described in Fig.3(b), which includes four convolutional layers, three max-pooling layers, two local response normalization layers, and two fully-connected layers. The rectified linear unit is adopted as the activation function.

Learning procedure. Similar to the relation prediction network, the training process can be done by back-propagation (BP) using stochastic gradient descent (SGD) . The difference is that we have missing attribute labels in the training set. Specifically, we use the cross-entropy loss for attribute classification, with an estimated attribute y~l\widetilde{y}_{l}, the back-propagation error ele^{l} is

Experiments

Facial attribute datasets. To enable accurate social relation prediction, we employ three datasets to cover a wide-range of facial attributes: Annotated Facial Landmarks in the Wild (AFLW) (24,386 faces), CelebFaces (87,628 faces) and a facial expression dataset on Kaggle contest (35,887 faces). Table 2 summarises the data. All the attributes are binary and labelled manually. To evaluate the performance of the cross dataset approach, we randomly select 2,000 testing faces from AFLW and CelebFaces, respectively. For the Kaggle dataset, we follow the protocol of the expression contest by using the 7,178 testing faces.

Social relation dataset. We build the social relation dataset as described in Sec. 3.2. Table 3 presents the statistics of this dataset. Specially, to reduce the potential effect from annotators’ subjectivity, we select a subset (522 cases) from the testing images and build an additional testing set. The images in this subset are all from movies. As the annotators know the movies’ story, they can give objective annotation assisted by the subtitle.

Baseline algorithm. In addition to the strong baseline method in Sec. 3.3, we train an additional baseline classifier by extracting the HOG features from the given face images. The features from the two faces are then concatenated and we use a linear support vector machine (SVM) to train a binary classifier for each relation trait. For simplicity, we call this method “HOG+SVM”, and the baseline method in Sec. 3.3 “Baseline DCN”.

Performance evaluation. We divide the relation dataset into training and testing partitions of 7,459 and 847 images, respectively. The face pairs in these two partitions are mutually exclusive. To account for the imbalance positive and negative samples, a balanced accuracy is adopted:

where NpN_{p} and NnN_{n} are the numbers of positive and negative samples, whilst npn_{p} and nnn_{n} are the numbers of true positive and true negative. We first train the network as Sec. 3.3 (i.e., Baseline DCN). After that, to examine the influences of different attribute groups, we pre-train four DCN variants using only one group of attribute (expression, age, gender, and pose). In addition, we compare the effectiveness between the full model with and without spatial cue.

Fig. 4 shows the accuracies of the different variants. All variants of our deep model outperform the baseline HOG+SVM. We observe that the cross dataset pre-training is beneficial, since pre-training with any of the attribute groups improves the overall performance. In particular, pre-training with expression attributes outperforms other groups of attributes (improving from 64.0% to 70.6%). This is not surprising since social relation is largely manifested from expression. The pose attributes come next in terms of influence to relation prediction. The result is also expected since when people are in a close or friendly relation, they tend to look at the same direction or face each other. Finally, the spatial cue is shown to be useful for relation prediction. However, we also observe that not every trait is improved by the spatial cue and some are degraded. This is because currently we simply use the face scale and location directly, of which the distribution is inconsistent in images from different sources. As for the relation traits, “dominant” is the most difficult trait to predict as it needs to be determined by more complicated factors, such as the social role and environmental context. The trait of “assured” is also difficult since it is visually subtle compared to other traits such as “competitive” and “friendly”. In addition, we conduct analysis on the movie testing subset. Table 4 shows the average accuracy on the eight relation traits of the two baseline algorithms and the proposed method. The results correspond to that of the whole testing set. This supports the reliability of the proposed dataset.

Some qualitative results are presented in Fig. 5. Positive relation traits, such as “trusting”, “warm”, “friendly” are inferred between the US President Barack Obama and his family members. Interestingly, “dominant” trait is predicted between him and his daughter (Fig. 5(a)). The upper image in Fig. 5(b) was taken in his election celebration party with the US Vice President Joe Biden. We can see the relation is quite different from that of the lower image, in which Obama was in the presidential election debate. Fig. 5(c) includes the images for Angela Merkel, Chancellor of Germany and David Cameron, Prime Minister of UK. The upper image is usually used in the news articles on US spying scandal, showing low probability on the “trusting” trait. More positive and negative results on different relation traits are shown in Fig. 6 (a). In addition, we show some false positives in Fig. 6 (b), which are mainly caused by faces with large occlusions.

2 Further Analyses

Facial expression recognition. Given the essential role of expression attributes, we further evaluate our cross dataset approach on the challenging Kaggle facial expression dataset. Following the protocol in , we classify each face into one of the seven expressions, (i.e. angry, disgust, fear, happy, sad, surprise, and neutral). The Kaggle winning method reports an accuracy of 71.2% by applying a CNN with SVM loss function. Our method achieves a better performance of 75.10%, through fusing data from multiple sources with the proposed bridging layer.

The effectiveness of bridging layer. We examine the effectiveness of the bridging layer from two perspectives. First, we show some clusters discovered by using the face descriptor (Sec. 3.4). It is observed that the proposed approach successfully divides samples from different datasets into coherent clusters of similar face patterns. Second, we examine the balanced accuracy (Eqn. (5)) of attribute classification with and without the bridging layer (Table 5). It is observed that bridging layer benefits the recognition of most attributes, especially the expression attributes. The results suggest the bringing layer an effective way to combine heterogeneous datasets for visual learning by deep network. Moreover, treating bridging layer as input provides higher accuracy than as output.

3 Application: Character Relation Profiling

We show an example of application on using our method to profile the relations among the characters in a movie automatically. Here we choose the movie Iron Man. We focus on different interaction patterns, such as conversation and conflict, of the main roles “Tony Stark” and “Pepper Potts”. Firstly, we apply a face detector to the movie and select the frames capturing the two roles. Then, we apply our algorithm on each frame to infer their relation traits. The predicted probabilities are averaged across 5 neighbouring frames to obtain a smooth profile. Fig. 7 shows a video segment with the traits of “friendly” and “competitive”. Our method accurately captures the friendly talking scene and the moment when Tony and Pepper were in a conflict (where the “competitive” trait is assigned with a high probability while the “friendly” trait is low).

Conclusion

In this paper we investigate a new problem of predicting social relation traits from face images. This problem is challenging in that accurate prediction relies on recognition of complex facial attributes. We have shown that deep model with bridging layer is essential to exploit multiple datasets with potential missing attribute labels. Future work will integrate face cues with other information such as environment context and body gesture for relation prediction. We will also investigate other interesting applications such as relation mining from image collection in social network. Moreover, we can also explore modelling relations of more than two people, which can be implemented by voting or graphical model, where each node is a face and edge is relations between faces.

References