Mutual Information Regularized Identity-aware Facial ExpressionRecognition in Compressed Video

Xiaofeng Liu, Linghao Jin, Xu Han, Jane You

Introduction

The video modality is increasingly important in many computer vision applications . Considering the natural dynamic property of the human face expression , many works propose to explore spatio-temporal features of facial expression recognition (FER) from the videos. Recently, deep neural networks (DNN) have achieved significant progress for image-based facial expression recognition , while the processing of expression video is still challenging.

Although the multi-frame series can inherit richer information and the temporal-correlation between consecutive frames can typically be useful for FER, the video also introduced a lot of redundancy. The signal-to-noise ratio (SNR) in FER videos is exceptionally low due to the slight muscle activity .

The typical spatio-temporal FER DNNs explore the spatial and temporal clues in a series of frames to extract the facial expression-related information . The recursive neural network (RNN), 3D convolutional, and non-local networks are the most frequently used network backbones . Unfortunately, using DNN to process many consecutive frames can be computationally costly. It can be hard to scalable for the long videos . Furthermore, for some networks, such as RNNs, modeling long-term dependence can be difficult . Many FER techniques, in particular, are able to achieve good performance using image-based FER with decision-level fusion, which completely lost the temporal dependence. This suggests the temporal cues can be hard to explore in the low SNR FER videos with the conventional solutions .

We argue that the compressed domain can be suitable for the FER task for four reasons. 1) the consecutive frames in video modality have many uninformative and repeating patterns, which may drown the “interesting" and “true" signal . With the standard video compression algorithms, the compression ratio can usually be hundred times . Manipulating on the compressed domain can significantly reduce the cost of computation and memory. 2) the typical compression methods (e.g., MPEG-4, H.264, and HEVC) break the video to the I frame (intracoded frames) with the first image, and follows several P frames (predictive frames) which encoded as the “change" or “movement" , as shown in Figure. 1. The fundamental of expression is the action of the face muscle.. Many FER systems, in reality, are based on the action unit framework .

As a result, compressed P frames will inherit off-the-shelf but useful expression-related factors, and their pattern is substantially simpler than raw images. 3) our compressed domain exploration can also be effective since it focuses on the “true" signals rather than processing the repeatedly near-duplicates . 4) Since the to-be-processed data is transmitted in compressed format, the decoding procedure is not needed in the real-world mission.

Furthermore, the FER task has long suffered from the high inter-subject variation caused by identity discrepancies in facial attributes . The learned features may capture more identity-related information than expression-related information, and are not purely related to the FER task. Noticing that the P frames may contain the relative location of face key points, which can be related to the identity . Metric learning is a standard approach for extracting identity from the expression representation . Inspired by the adversarial disentanglement, propose to render the identity removed face, which is inspired by adversarial disentanglement (GAN). These researchs, on the other hand, concentrate on image-based FER. extend the metric learning for video data by explicitly substitute the image with the video features, without taking into account the video’s characteristics. Furthermore, these methods necessitate the use of the identity label and multiple expressions of the same person, which significantly restricts their applicability to the in-the-wild FER task .

In this paper, we target to exploit the identity information from the I frame using a pre-trained face recognition network, e.g., FaceNet . Their identity embeddings are remarkably reliable, since they achieve high accuracy over millions of identities , and robust to a broad range of nuisance factors such as expression, pose, illumination and occlusion variations.

Using the identity feature as the anchor, we can explicitly enforce the marginal independence of our identity and expression feature. Instead of the complicated adversarial training for disentanglement, we adopt the mutual information (MI) as the statistical measure of the independence of these two representations . The MI of two random variables can usually be intractable to directly and precisely measure in a high-dimensional space . Recently, some of the works illustrate that mutual information can be differentiable approximated . We propose to minimize the differentiable MI measure as the objective. Practically, it can be a latent space min-min game of an encoder-discriminator framework, which follows a collaborative fashion rather than adversarial competition. We note that GAN is notorious for its unstable model collapse , while our MI regularization is concise and efficient.

This work extends our previous work in the following significant ways:

∙\bullet A novel expression and identity disentanglement framework based on practical mutual information minimization, which follows a min-min game with the joint and marginal distribution sampling.

∙\bullet We demonstrate the generality of our framework in more FER dataset, i.e., MMI, carry out all experiments using the novel MIC framework.

∙\bullet Moreover, the systematical cross-dataset evaluation, sensitivity analysis, and identity feature extraction analysis are provided.

The main contributions of this paper are summarized as follows:

∙\bullet We propose to inference expression from the residual frames, which explores the off-the-shelf yet valuable expression related muscle movement in the up to two orders of magnitude compressed domain.

∙\bullet Targeting for the identity-aware video-based FER, the independence of expression and identity representations from P frames and I frame are enforced with the differentiable MI measure.

∙\bullet The separability of expression and identity representations is maximized by a min-min game with the joint and marginal distribution sampling, which does not rely on the unstable adversarial game to achieve identity elimination.

We evidenced its effectiveness on several video-based FER benchmarks with much faster inference. The promising performance evidenced its generality and scalability.

Related Works

Video-based FER has been thoroughly researched, since facial expression is a natural and universal means for human communication . Considering that the expression is essentially a dynamic action which should take minute muscle movements through time into account . Traditionally, the handcrafted features are utilized to represent the spatio-temporal cues and for FER. Frame aggregation and spatiotemporal FER networks are being developed in parallel with the exponential growth of deep learning. Frame aggregation approaches may make use of image-based FER networks by conducting frame-wise aggregation at the decision-level or feature-level .

The essential temporal correlation, on the other hand, is not investigated. Instead, the spatio-temporal FER networks use sequential frames to utilize both the spatial/textural and temporal information . Furthermore, cascaded networks suggest integrating CNN-learned perceptual vision representations with RNNs for variable-length video . Moreover, by using a non-local network for video processing, the number of potential connections and the corresponding computation costs grow exponentially to the number of frames. The 3D convolutional has shared kernel-weights along the time axis, which has also been widely used for video-based FER .

However, recent studies have only looked at the image domain, and the spatiotemporal FER does not outperform aggregation methods substantially . To the best of our knowledge, this is the first effort to investigate the compressed video FER, which is orthogonal to these advantages and can be conveniently added to each other.

Video compression convert the digital video into a specific format that is suitable for recording and distribution of this video . Conventionally, the H.264/MPEG-4 and the Advanced Video Coding are the typical algorithms . Typically, the video sequence is divided into several Group Of Pictures (GOP) by the video codecs. In a GOP, there is an I frame and follow by several P frames. Specifically, the I frame is a self-contained RGB frame with full visual representation, while the P frame is the inter frames that hold motion vectors and residuals w.r.t. the previous frame . The motion vector can be used as an alternative to the optical flow , which needs to decode the RGB images.

The recent action recognition method proposes to aggregate I frame, residuals, and motion vectors in the compressed domain, without the RGB image decoding. Although FER shares some similarities with action recognition, the movement range and use of I frame can be vastly different. In action recognition , the I frame is directly used to predict the action and combine with the result of P frames. In contrast, the I frame in FER usually be a neutral face (different from the video label). Furthermore, the low-resolution motion vector can not well encode the expression.

Mutual-information has a long history in unsupervised learning. The infomax principle , as prescribed for neural networks, advocates maximizing the MI between network input and output. This can be the fundamental of many ICA algorithms, which can be nonlinear but are often hard to adapt for use with deep networks. Recently, some of the works proposed to achieve unsupervised learning with MI. proposes a generative adversarial network to minimize MI with positive and negative samples for Independent Component Analysis (ICA). It introduces a strategy to draw samples from the joint distribution and the product of marginal distributions and proposed to train an encoder and a discriminator to minimize the Jansen-Shannon divergence. Moreover, it has recently been shown that the GAN framework can be extended not only to maximize or minimize MI but also to explicitly compute it using the Mutual Information Neural Estimation (MINE) proposed in . In , the DeepInfoMax (DIM) is proposed to learn the representations based on both local and global information. In , Deep Graph Infomax (DGI) extends this approach to graph-structured data.

Inspired by these works, we are targeting to utilize the MI as the justified measure of independence, and minimize it directly as our disentanglement objective.

Eliminating identity can benefit to extract more “pure" expression feature . We note that previous identity-aware FER methods usually explicitly require the identity labels of FER datasets to sample the triplets , while the identity label is not common in FER. In contrast, our solution does not relies on the identity label of FER samples, but utilizes the easily available face recognition dataset.

The typical solution for FER is metric learning. Our MI regularization is also related to the triplet loss , which maximizes the Euclidean or cosine distance between two identities. With the development of GAN, adversarial training also can be utilized for disentanglement . Instead, we consider the mutual information to be a more meaningful divergence to capture complex non-linear relationships, between the identity and expression representations. Besides, the identity label is required in these methods . We can also choose adversarial training as a baseline to achieve identity elimination in our framework. We note that the adversarial game is notorious for its unstable model collapse , while our solution follows a collaborative way.

Several adversarial disentanglement works demonstrate that simply separate the input may result in the extracted feature has no meaningful information . The reconstruction of input can explicitly enforce the disentangled factors to be complementary to each other. However, reconstructing the video can be hugely underconstrained.

Methodology

We propose to develop an efficient video-based FER framework that operates directly on the compressed domain. The overall framework is shown in Figure. 2, which is consisted of four core modules. The pre-trained identity branch and FER branch (frame embedding network fEf_{E}, aggregation module, and Classifier) work on the I frame and undecoded P frames, respectively. The dependence of identity and expression representations is then measured using a differentiable mutual information regularization module. Furthermore, when the apex frame is annotated, the complementary constraint can be applied to stabilize the early stage training.

Noticing that the motion vectors Tt\mathcal{T}^{t} has a much lower resolution, since their values within the same macroblock are identical. Considering the micro-movements of facial expression in each frame, the low resolution Tt\mathcal{T}^{t} usually not helpful for the FER. For P frame reconstruction Pit=Pi−Titt−1+ΔitP_{i}^{t}=P_{i-\mathcal{T}_{i}^{t}}^{t-1}+\Delta_{i}^{t}, where index all the pixels and P0=IP^{0}=I. Then, Tt\mathcal{T}^{t} and Δt\Delta^{t} are processed by discrete cosine transform and entropy-encoded.

Besides, most existing action recognition methods with compressed video independently concatenate the paired Δt\Delta^{t} and Tt\mathcal{T}^{t} at each time step and predict an action score of each P-frame. The temporal cues and their development patterns are important for the FER task . We simply choose the LSTM in to model the sequential development of residual frames and summarize the information to an expression feature zEz_{E}. Since our LSTM is applied to 512-dim features, the computation burden is largely smaller than work on the raw images. Noticing that more advanced RNN, 3D CNN, or attention networks can potentially be utilized to replace our LSTM model to further boost the performance .

2 MI regularization

To eliminate the identity-related factors in our FER representation, we propose to utilize the identity feature from pre-trained face recognizer zIz_{I} as anchor, and explicitly inspect the information w.r.t. zIz_{I} in zEz_{E}.

Achieving the disentanglement of different factors requires two major objectives, i.e., 1) each factor has its specific information, and 2) does not incorporate the information of the other factors . For example, zEz_{E} can achieve 1) using the conventional C.E. loss minimization w.r.t. the expression label yy, and zIz_{I} is from the pre-trained identity extractor, which inherently has identity information. However, how to explicitly measure the dependency between these factors and minimize this metric to achieve the latter objective can be challenging.

Actually, the mutual information (MI) is the exact metric to measure the amount of information obtained about one random variable through observing another random variable.

MI minimization explicitly enforces the joint distribution to be equal to the product of marginals, which leads to the statistical independence of two vectors. Instead, the MI maximization can result in two vectors have the same information, and the MI is simply equal to the entropy of a variable.

Besides, we leverage Monte-Carlo integration to avoid computing the integrals to compute MI(zE;zI)^nMI{\widehat{(z_{E};z_{I})}}_{n} as

The estimated MI(zE;zI)MI(z_{E};z_{I}) is used as the supervision to update the FER branch. By utilizing MI regularization, the adversarial discriminator is no longer needed in our new framework, which makes the balance of each module easier. Note that we need the additional neural network TθT_{\theta} to measure the MI, but it is collaboratively trained with the FER branch to maximize the discrepancy between the two features. Essentially, we are playing a min-min game instead of a min-max game. Therefore, it is easier to stabilize the training (compared to adversarial training).

Besides, MI is a symmetric measure, while the conditional entropy H(zI∣zE)=H(zI)−MI(zI;zE)H(z_{I}|z_{E})=H(z_{I})-MI(z_{I};z_{E}) optimized in conventional disentanglement works is asymmetric and essentially we should calculate both H(zI∣zE)H(z_{I}|z_{E}) and H(zI∣zE)H(z_{I}|z_{E}) as supervision signal .

To maximize the discrepancy of zEz_{E} and zIz_{I}, we can also apply the adversarial disentanglement solutions . Nevertheless, with the above-mentioned limitations, such methods can be hard to optimize and lead to inferior performance.

3 Complementary constraint

Many FER datasets follow a well-defined collection protocol, which usually starts from the neutral face and then develops to an expression. Specifically, the video in CK+ consists of a sequence that shifts from the neutral expression to an apex facial expression. The last frame usually is the apex frame, which has the most strong expression intensity. Actually, the image-based FER methods select the last three frames to construct their training and testing datasets. Similarly, in MMI , the video frames usually start from the neutral face and develop to the apex around the middle of the video, and returning back to the neutral at the end of the video. Noticing that the apex frame (i.e., last frame in CK+ or middle frame in MMI) can clearly incorporate both the identity and expression information. Therefore, we are possible to utilize the apex frame as a reference of reconstruction, and simply apply the L2\mathcal{L}_{2} loss.

where I^Apex=Dec(zI,zE)\hat{I}_{Apex}=Dec(z_{I},z_{E}). The complementary restriction is not necessary for our system since the FER loss is heavily weighted in the FER branch. It requires to maintain sufficient information w.r.t. expression and not easy to have nothing meaningful. However, the complementary constraint does helpful for the convergence in the early stage. When the apex is annotated, we only need to decode the apex frame in the decoded image domain at the start of a few training epochs.

4 Overall objectives

We have three to be minimized objectives, i.e., cross-entropy loss, mutual information and L1\mathcal{L}_{1} loss, which works collaboratively to update each module. The expression classification is the main task of the FER. We choose the typical cross-entropy loss LCE=−∑c=1Cyclog⁡(Cls(zE)c)\mathcal{L}_{CE}=-\sum_{c=1}^{C}y_{c}\log(Cls(z_{E})_{c}) to ensure zEz_{E} contains sufficient expression information and finally have a good performance on CC-class expression classification. We use ycy_{c} and Cls(zE)cCls(z_{E})_{c}indicate the cthc^{th} class probability of the label and classifier softmax predictions respectively. Since the FER branch can be updated with all of the losses, we assign the balance parameter α∈\alpha\in and β∈\beta\in to mutual information and L1\mathcal{L}_{1} loss minimization objectives, respectively. Specifically, we update our fEf_{E} and LSTM modules with

For the MI estimator TθT_{\theta}, we update it with MI(zE;zI)^nMI{\widehat{(z_{E};z_{I})}}_{n}. Moreover, the decoder module DecDec is updated with L1\mathcal{L}_{1}. The detailed training flow is shown in Algorithm 1. We note that only the FER branch, i.e., fEf_{E}, LSTM and Cls, is used for testing.

Since we are using the neural network TθT_{\theta} for mutual information estimation, it is scalable, flexible, and completely trainable via back-propagation. Moreover, the decoder for reconstruction is used for complementary constraints. In contrast, FLF uses a discriminator and a decoder for adversarial disentanglement and complementary constraint. Considering the TθT_{\theta} in MIC and discriminator in FLF has a similar structure, there is no significant difference for the network complexity. The mutual information calculation with Eq. 4 has the complexity of O(n)\mathcal{O}(n), where nn is the number of sampled data. The linear complexity can be simple for implementation. Conventionally, the fast computation of MI is limited to discrete variables . For the continuous random variables, its complexity is quadratic to the number of samples, which is not desired for a loss function.

Experiments

In this section, we first detail our experimental setup, present a quantitative analysis of our model, and finally compare it with state-of-the-art methods. The good FER accuracy and high inference speed in testing demonstrate its effectiveness.

CK+ Dataset is referring to the Cohn-Kanade AU-Coded Expression dataset, which is a widely accepted FER benchmark . The video is collected in a restricted environment, in which the participate subjects are facing the recorder with an empty background. The video in CK+ consists of a sequence that shifts from the neutral expression to an apex facial expression. The last frame usually is the apex frame, which has the most strong expression intensity. The expression included in this dataset is anger, contempt, disgust, fear, happiness, sadness, and surprise. There are 327 facial expression videos collected from 118 subjects. Following the previous works, we use subject independent 10-folds cross-validation .

Many FER datasets follow a well-defined collection protocol, which usually starts from the neutral face and develops to an expression. Specifically, the image-based FER methods select the last three frames to construct their training and testing datasets. In Fig. 3, we show the compressed and decoded frames in CK+ dataset. Similarly, in MMI , the video frames usually start from the neutral face and develop to the apex around the middle of the video, and returning back to the neutral at the end of the video. Noticing that the apex frame (i.e., last frame in CK+ or middle frame in MMI) can clearly incorporate both the identity and expression information. Therefore, we are possible to utilize the apex frame as a reference of reconstruction, and simply apply the L1\mathcal{L}_{1} loss.

MMI Dataset consist of a total of 326 facial expression videos from 32 participants. There are 213 labeled videos with the expression label angry, disgust, fear, happy, sad, and surprise. The video frames start from the neutral face. Then the expression is developed to the apex in the middle of video, and returning back to the neutral at the end of the video. In our experiments, we follow the previous works to use subject independent 10-folds cross-validation . In Fig. 4, we show the compressed and decoded frames in the MMI dataset.

AFEW Dataset is more close to the uncontrolled real-world environment. It is consists of video clips of movies . The video in AFEW has a spontaneous facial expression. The AFEW has seven expressions: anger, disgust, fear, happiness, sadness, surprise, and neutral. Following the evaluation protocol in EmotiW , there are training, validation, and testing sets. Since its testing label is not available, we follow the previous work to use the validation set for comparison . Noticing that the validation set is not used in the training stage for the parameter or hyper-parameter tuning. In Fig. 5, we show the compressed and decoded frames in the AFEW dataset.

2 Implementation details

We preprocess video frames and augment the data according to for fair comparison. For these three datasets, the videos only have one GOP and do not need to segment the video. We utilize the Pytorch deep learning platform for our framework. In the training stage, on all datasets, we set the batch size to 48. All of the modules use the Adam optimizer with momentum 0.9, and a weight decay of 1e-5 for 100 training epochs. On the CK+ and MMI datasets, the learning rate is initialized to 1e-1, and be modified to 1e-2 for the 30th epoch. For the AFEW dataset, we initialize the learning rate to 1e-4, and modify it to 8e-6 for the 30th epoch and 1e-7 for the 60th epochs.

All of our training/testings use an NVIDIA Titan X GPU. We note that the calculation of the accumulated residuals to recover the apex frame is measured on Intel E5-2698 v4 CPUs, but we do not need this operation in testing. For the testing speed, we measure the average frame per second (fps) according to the average running time, which is the sum of the data pre-processing time and the FER branch forward pass time.

FAN shows the ResNet can be a powerful backbone. To further demonstrate the generality of our framework, we use the first five convolutional layers (i.e., before the second residual block) and the first fully connected layer in ResNet18 as our feature extractor backbone. We denote this ResNet backbone as ResNet6.

For the CK+ and MMI dataset, we choose the last or middle frame as the apex reference image, respectively. Since the complementary constraint is only used to stabilize the initial training of disentanglement, we uniformly decrease β\beta from 1 to 0 until the 30th30^{th} epoch. Practically, we use the grid search to find the optimal α\alpha and set it to 0.1, 0.1, 0.2 on CK+, MMI, and AFEW datasets, respectively.

3 Evaluation and ablation study

The 10-fold cross-validation performance of our proposed method is shown in Table 1. For a fair comparison, the image-based experiment settings are not incorporated in the tables. Besides, only the state-of-the-art (SOTA) accuracy obtained by the single-models (non-ensemble model) is listed.

Many models, e.g., PHRNN-MSCNN , CTSLSTM and SC , achieved the SOTA performance by utilizing the facial landmarks. However, this operation highly relies on fine-grained landmark detection, which itself is a challenging task , and unavoidably introduced additional computation.

Based on the mode variational LSTM , our proposed MIC achieves the SOTA result without the facial landmarks, 3D face models, or optical flow. It worth noticing that the much simpler CNN encoding network makes more residual frames that can be processed parallel than . Moreover, our efficiency also benefits from avoiding the decompress of the video. Since the videos are typically stored and transmitted with the compressed version, and the residuals are off-the-shelf. As a result, the proposed compressed domain MIC can speed up the testing about 3 times over , and achieve better accuracy.

Besides, is a typical metric-learning-based identity removing method. Our solution can significantly outperform it with respect to both speed and accuracy. Actually, the sampling of tuplets usually makes the training not scalable , while our identity eliminating scheme is concise and effective.

When we remove some modules from our framework, the performances have different degrees of decline. We use -MI and -I^\hat{I} to denote the MIC without MI regularizer or complementary constraint, respectively. The performance drop of MIC is significant when we remove the MI regularization module, which further evidenced that the identity can be a notorious factor for FER. The result also implies that the identity can be well encoded by the face recognition network and the disentanglement with mutual information is feasible. Compared with using the conventional adversarial training based disentanglement as an alternative (i.e., MIC-MI+Adv), our MI regularizer is easy to train and can converge 1.8 times faster in training.

We can also follow the action recognition method to concatenate the motion vector and residual as the input, and denote as MIC+Tt\mathcal{T}^{t}. However, we do not achieve significant performance on all datasets, but the inference speed in testing can be slower. This may be related to the coarse resolution of the motion vector can not well describe the fine-grain muscle movement of the face. More appealingly, the performance of our MIC can be further improved With the ResNet backbone. With the simplified 6-layer ResNet, MIC outperforms the FAN w.r.t. both accuracy and processing speed. Aggregation-based methods do not explore the temporal cues and can be computational costly to compare all possible image pairs within a set . We note that the ResNet18 used in FAN is sequentially pre-trained on the additional MSCeleb-1M face recognition dataset and FER+ expression dataset.

The confusion matrix of our proposed MIC method on the CK+ is reported in Figure 7 (left). The accuracy for the expression of class happiness, disgust, anger, surprise, and contempt are almost perfect.

In Table 2, we investigate the identity eliminating performance. The first metric is following to use recognize identity with zEz_{E}. Besides, we can directly use mutual information as the metric of independence. We can see that the residual frame itself can incorporate much less identity information than the decoded images as in , while it is still possible to detect identity with the facial contours. The MI regularization can explicitly remove the identity factors and outperforms the adversarial training .

The evaluation results on the MMI dataset are shown in Table 3. The performance is also consistent with the CK+ dataset, which evidenced its effectiveness and generality. All of our methods achieve comparable performance to the landmark-based STOA methods. It is more promising that MIC can be significantly better than the methods without the landmarks. Our MIC is efficient, since we explore the correlation in video frames in the compressed domain.

In Figure 6, we give a comparison of accuracy w.r.t. each emotion among five MIC baselines, and the confusion matrix of our MIC is reported in Figure 7 (middle). There is a good performance for the expression class of fear, happiness, sadness, and surprise. In contrast, the accuracy of expression class anger and disgust is relatively limited. Especially, there is a high degree of confusion between anger and disguise. This may be related to the subtle movements between these expressions are relatively in the residual frames.

The evaluation of the proposed MIC on the AFEW dataset is shown in Table 4. We note that only the SOTA accuracy obtained by the single-models (non-ensemble model) are listed for a fair comparison. Besides, the audio modality in AFEW can be used to boost the recognition performance . We note that we only focus on the image compression in this paper, and the audio/video data are stored in separate tracks, but the additional modality can also potentially to be added on our framework following the multi-modal methods .

With the simplified mode variational LSTM-based backbone, the exploration in the compressed domain can achieve comparable or even better recognition performance. More promisingly, our MIC can also achieve real-time processing for the uncontrolled environment, which evidenced its generality. We note that the typical time resolution in FER is 24fps .

In addition, we note that the complementary constraint requires the apex frame in training, which is not applicable for the AFEW dataset. We do not apply the reconstruction loss in the AFEW task. Although the requirement of the apex frame imposes some limitations on the training, it does not affect the generality of the testing or implementation of the trained model with CK+ and MMI.

Some of the works propose to improve the image-based FER networks and combine the frame-wise scores for video-based FER . The image-based FER methods achieves high performance, but uses a very deep network DenseNet-161 and pretrains it on the private Situ dataset. Moreover, utilize the sophisticated post-processing. Actually, an intuition of a statistic-based solution is to avoid LSTM and speed up the processing. However, with the super deep and complicated structure, their processing can be much slower than our solution.

uses VGGFace as the backbone of fEf_{E} and an RNN model with LSTM units to capture the temporal dynamic cues of the videos. Moreover, also propose to modify the LSTM model for the spatial-temporal modeling. However, all of the above solutions are applied to the decoded space, which requires decoding processing and needs to handle much more complicated data. With the 6-layer ResNet backbone as an expression feature extractor, the performance of our MIC model can be further improved without an additional FER dataset for pre-training. Overall, the proposed MIC can improve the testing speed by a large margin and can achieve the SOTA accuracy as the previous models.

4 Identity feature extraction

We adopt the pre-trained face recognizer FaceNet to extract the identity factor from the I frame. The feature embedded with the convolutional layers and the first fully connected layer can be robust to a broad range of nuisance factors such as expression, pose, illumination, and occlusion variations, since it achieves high accuracy over millions of identities . To check the expression information in the extracted identity feature zIz_{I}, we use the feature for expression classification as . The results are shown in 9. We note that achieve zero FER accuracy with zIz_{I} does not mean zIz_{I} has no information about expression. Instead, approaching the chance probability, i.e., uniform distribution w.r.t. expression classes, indicates the expression is well disentangled from the identity feature zIz_{I} with FaceNet.

5 Sensitivity analysis

We use α\alpha and β\beta to balance the MI and complementary constraint terms and choose the best value with grid searching. In Tab. 10, we provide the sensitivity analysis of using different α\alpha in three datasets. We can see that we can achieve the best performance on CK+ and MMI with α=0.1\alpha=0.1. For the AFEW dataset, the performance is relatively stable for α\alpha from 0.15 to 0.25. We simply use 0.2 for all of our MIC models on the AFEW dataset.

β\beta is used to balance the complementary constraint term, which can be helpful for stabilizing the training. In Tab. 11, both the fixed β\beta and linear changing β\beta are compared. Decreasing β\beta from 1 to 0 for 30 or 50 epochs can usually achieve the best performance.

6 Cross-database validation

Since the subjects are different across CK+ and MMI datasets, the cross-dataset evaluation can be used to evidence if the trained model is affected by identity . In Tab. 5, the FER model trained CK+ training sets in the previous experiment is directly implemented to the MMI testing set. We can see that our MIC can outperform the other methods with the same backbone. In Tab. 6, we use MMI as training data and test on CK+ dataset. The proposed MIC outperforms the other methods with the same backbone consistently.

We note that FAN uses ResNet18 as a feature extractor and sequentially pre-trained on MSCeleb-1M face recognition dataset and FER+ expression dataset.

In Tab. 7 and Tab. 8, we use the training set of AFEW for training and test on MMI and CK+, respectively. Considering the large domain shift between AFEW and MMI/CK+, there is a significant performance drop. We note that the proposed method can also achieve better performance than its backbones .

7 Critical discussion and future work

The proposed MIC framework has demonstrated its effectiveness w.r.t. accuracy and testing speed. However, MIC highly relies on the separation of I and P frames in compressed video. Although MPEG-4, H.264, and HEVC formats are widely used, some of the compression solutions do not follow the motion compensation with I and P frames. For example, the YUV format compresses the video by considering the different changes of brightness and chromaticity. Moreover, frame loss can be common in real-world video transfer. How will the frame loss affect the FER performance is underexplored.

Identity can be the most challenging variation for FER , while the FER performance can also be affected by pose and illuminations. It can be promising to take the other variations into account.

The recent work proposes to utilize the additional unlabeled face dataset to boost the performance, which can be a powerful means for alleviating the scale issue of video-based FER.

For the large domain gap, e.g., AFEW to MMI/CK+, the domain adaptation methods can be used to achieve better cross-dataset performance.

Conclusion

In this paper, we target to explore the facial expression cues directly on the compressed video domain. We are motivated by our practical observation that facial muscle movements can be well encoded in the residual frames, which can be informative and free of cost. Besides, the video compression can reduce the repeating boring patterns in the videos, which rendering the representation to be robust. The increased relevance and reduced complexity or redundancy in FER videos make computation much more effective. We extract the identity and expression factor from the I frame and P frame, respectively, and explicitly enforce their independence with concise and effective mutual information regularization. When the apex frame label is available in training, the complementary constraint can further stabilize the training. In three video-based FER benchmarks, our MIC can improve the performance without the additional identity, face model, or facial landmarks labels. The processing speed of the test stage is promising for real-time FER. Moreover, our mutual information regularization can potentially be a good alternative to adversarial training for many disentanglement tasks.

References