LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood

Abhinav Kumar, Tim K. Marks, Wenxuan Mou, Ye Wang, Michael Jones, Anoop Cherian, Toshiaki Koike-Akino, Xiaoming Liu, Chen Feng

Introduction

Modern methods for face alignment (facial landmark localization) perform quite well most of the time, but all of them fail some percentage of the time. Unfortunately, almost all of the state-of-the-art (SOTA) methods simply output predicted landmark locations, with no assessment of whether (or how much) downstream tasks should trust these landmark locations. This is concerning, as face alignment is a key pre-processing step in numerous safety-critical applications, including advanced driver assistance systems (ADAS), driver monitoring, and remote measurement of vital signs . As deep neural networks are notorious for producing overconfident predictions , similar concerns have been raised for other neural network technologies , and they become even more acute in the era of adversarial machine learning where adversarial images may pose a great threat to a system . However, previous work in face alignment (and landmark localization in general) has largely ignored the area of uncertainty estimation.

To address this need, we propose a method to jointly estimate facial landmark locations and a parametric probability distribution representing the uncertainty of each estimated location. Our model also jointly estimates the visibility of landmarks, which predicts whether each landmark is occluded due to extreme head pose.

We find that the choice of methods for calculating mean and covariance is crucial. Landmark locations are best obtained using heatmaps, rather than by direct regression. To estimate landmark locations in a differentiable manner using heatmaps, we do not select the location of the maximum (argmax) of each landmark’s heatmap, but instead propose to use the spatial mean of the positive elements of each heatmap. Unlike landmark locations, uncertainty distribution parameters are best obtained by direct regression rather than from heatmaps. To estimate the uncertainty of the predicted locations, we add a Cholesky Estimator Network (CEN) branch to estimate the covariance matrix of a multivariate Gaussian or Laplacian probability distribution. To estimate visibility of each landmark, we add a Visibility Estimator Network (VEN). We combine these estimates using a joint loss function that we call the Location, Uncertainty and Visibility Likelihood (LUVLi) loss. Our primary goal in designing this model was to estimate uncertainty in landmark localization. In the process, not only does our method yields accurate uncertainty estimation, but it also produces SOTA landmark localization results on several face alignment datasets.

Uncertainty can be broadly classified into two categories : epistemic uncertainty is related to a lack of knowledge about the model that generated the observed data, and aleatoric uncertainty is related to the noise inherent in the observations, e.g., sensor or labelling noise. The ground-truth landmark locations marked on an image by human labelers would vary across multiple labelings of an image by different human labelers (or even by the same human labeler). Furthermore, this variation will itself vary across different images and landmarks (e.g., it will vary more for occluded landmarks and poorly lit images). The goal of our method is to estimate this aleatoric uncertainty.

The fact that each image only has one ground-truth labeled location per landmark makes estimating this uncertainty distribution difficult, but not impossible. To do so, we use a parametric model for the uncertainty distribution. We train a neural network to estimate the parameters of the model for each landmark of each input face image so as to maximize the likelihood under the model of the ground-truth location of that landmark (summed across all landmarks of all training faces).

The main contributions of this work are as follows:

This is the first work to introduce the concept of parametric uncertainty estimation for face alignment.

We propose an end-to-end trainable model for the joint estimation of landmark location, uncertainty, and visibility likelihood (LUVLi), modeled as a mixed random variable.

We compare our model using multivariate Gaussian and multivariate Laplacian probability distributions.

Our algorithm yields accurate uncertainty estimation and state-of-the-art landmark localization results on several face alignment datasets.

We are releasing a new dataset with manual labels of the locations of 6868 landmarks on over 19, ⁣00019,\!000 face images in a wide variety of poses, where each landmark is also labeled with one of three visibility categories.

Related Work

Early methods for face alignment were based on Active Shape Models (ASM) and Active Appearance Models (AAM) as well as their variations . Subsequently, direct regression methods became popular due to their excellent performance. Of these, tree-based regression methods proved particularly fast, and the subsequent cascaded regression methods improved accuracy.

Disadvantages of Heatmap-Based Approaches. These heatmap-based methods have at least two disadvantages. First, since the goal of training is to mimic a proxy ground-truth heatmap containing a fixed symmetric Gaussian, the predicted heatmaps are poorly suited to uncertainty prediction . Second, they suffer from quantization errors since the heatmap’s argmax is only determined to the nearest pixel . To achieve sub-pixel localization for body pose estimation, replaces the argmax with a spatial mean over the softmax. Alternatively, for sub-pixel localization in videos, samples two additional points adjacent to the max of the heatmap to estimate a local peak.

Landmark Regression with Uncertainty. We have only found two other methods that estimate uncertainty of landmark regression, both developed concurrently with our approach. The first method estimates face alignment uncertainty using a non-parametric approach: a kernel density network obtained by convolving the heatmaps with a fixed symmetric Gaussian kernel. The second performs body pose estimation with uncertainty using direct regression method (no heatmaps) to directly predict the mean and precision matrix of a Gaussian distribution.

2 Uncertainty Estimation in Neural Networks

Uncertainty estimation broadly uses two types of approaches : sampling-based and sampling-free. Sampling-based methods include Bayesian neural networks , Monte Carlo dropout , and bootstrap ensembles . They rely on multiple evaluations of the input to estimate uncertainty , and bootstrap ensembles also need to store several sets of weights . Thus, sampling-based methods work for small 11D regression problems but might not be feasible for higher-dimensional problems .

Sampling-free methods produce two outputs, one for the estimate and the other for the uncertainty, and optimize Gaussian log-likelihood (GLL) instead of classification and regression losses . combines the benefits of sampling-free and sampling-based methods.

Recent object detection methods have used uncertainty estimation . Sampling-free methods jointly estimate the four parameters of the bounding box using Gaussian log-likelihood , Laplacian log-likelihood , or both . However, these methods assume the four parameters of the bounding box are independent (assume a diagonal covariance matrix). Sampling-based approaches use Monte Carlo dropout and network ensembles for object detection. Uncertainty estimation has also been applied to pixelwise depth regression , optical flow , pedestrian detection and 33D vehicle detection .

Proposed Method

Figure 2 shows an overview of our LUVLi Face Alignment. The input RGB face image is passed through a DU-Net architecture, to which we add three additional components branching from each U-net. The first new component is a mean estimator, which computes the estimated location of each landmark as the weighted spatial mean of the positive elements of the corresponding heatmap. The second and the third new component, the Cholesky Estimator Network (CEN) and the Visibility Estimator Network (VEN), emerge from the bottleneck layer of each U-net. CEN and VEN weights are shared across all U-nets. The CEN estimates the Cholesky coefficients of the covariance matrix for each landmark location. The VEN estimates the probability of visibility of each landmark in the image, 11 meaning visible and meaning not visible. For each U-net ii and each landmark jj, the landmark’s location estimate μij\bm{\mu}_{ij}, estimated covariance matrix Σij\bm{\Sigma}_{ij}, and estimated visibility v^ij\widehat{v}_{ij} are tied together by the LUVLi loss function Lij\mathcal{L}_{ij}, which enables end-to-end optimization of the entire framework.

Rather than the argmax of the heatmap, we choose a mean estimator for the heatmap that is differentiable and enables sub-pixel accuracy: the weighted spatial mean of the heatmap’s positive elements. Unlike the non-parametric model of , our uncertainty prediction method is parametric: we directly estimate the parameters of a single multivariate Laplacian or Gaussian distribution. Furthermore, our method does not constrain the Laplacian or Gaussian covariance matrix to be diagonal.

Let Hij(x,y)\bm{H}_{ij}(x,y) denote the value at pixel location (x,y)(x,y) of the jthj{\text{th}} landmark’s heatmap from the ithi{\text{th}} U-net. The landmark’s location estimate μij=[μijx,μijy]T\bm{\mu}_{ij}=[\mu_{ijx},\mu_{ijy}]^{T} is given by first post-processing the pixels of the heatmap Hij\bm{H}_{ij} with a function σ\sigma, then taking the weighted spatial mean of the result (See (16) in the supplementary material). We considered three different functions for σ\sigma: the ReLU function (eliminates the negative values), the softmax function (makes the mean estimator a soft-argmax of the heatmap ), and a temperature-controlled softmax function (which, depending on the temperature setting, provides a continuum of softmax functions that range from a “hard” argmax to the uniform distribution). The ablation studies (Section 5.5) show that choosing σ\sigma to be the ReLU function yields the simplest and best mean estimator.

2 LUVLi Loss

Occluded landmarks, e.g., landmarks on the far side of a profile-pose face, are common in real data. To explicitly represent visibility, we model the probability distributions of landmark locations using mixed random variables. For each landmark jj in an image, we denote the ground-truth (labeled) visibility by the binary variable vj∈{0,1}v_{j}\in\{0,1\}, where 11 denotes visible, and the ground-truth location by pj\mathbf{p}_{j}. By convention, if the landmark is not visible (vj=0v_{j}=0), then pj=∅\mathbf{p}_{j}=\emptyset, a special symbol indicating non-existence. Together, these variables are distributed according to an unknown distribution p(vj,pj)p(v_{j},\mathbf{p}_{j}). The marginal Bernoulli distribution p(vj)p(v_{j}) captures the probability of visibility, p(pj∣vj ⁣= ⁣1)p(\mathbf{p}_{j}|v_{j}\!=\!1) denotes the distribution of the landmark location when it is visible, and p(pj∣vj ⁣= ⁣0) ⁣= ⁣1∅(pj)p(\mathbf{p}_{j}|v_{j}\!=\!0)\!=\!\mathbf{1}_{\emptyset}(\mathbf{p}_{j}), where 1∅\mathbf{1}_{\emptyset} denotes the PMF that assigns probability one to the symbol ∅\emptyset.

After each U-net ii, we estimate the joint distribution of the visibility vv and location z{\bf z} of each landmark jj via

where qv(v)q_{v}(v) is a Bernoulli distribution with

where v^ij\widehat{v}_{ij} is the predicted probability of visibility, and

where P{\mathcal{P}} denotes the likelihood of the landmark being at location z{\bf z} given the estimated mean μij\bm{\mu}_{ij} and covariance Σij\bm{\Sigma}_{ij}.

The LUVLi loss is the negative log-likelihood with respect to q(v,z)q(v,{\bf z}), as given by

and thus minimizing the loss is equivalent to maximum likelihood estimation.

The terms of (5) are a binary cross entropy plus vjv_{j} times the negative log-likelihood of pj\mathbf{p}_{j} with respect to P{\mathcal{P}}. This can be seen as an instance of multi-task learning , since we are predicting three things about each landmark: its location, uncertainty, and visibility. The first two terms on the right hand side of (5) can be seen as a classification loss for visibility, while the last term corresponds to a regression loss of location estimation. The sum of classification and regression losses is also widely used in object detection .

Minimization of negative log-likelihood also corresponds to minimizing KL-divergence, since

where pv:=p(vj=1)p_{v}:=p(v_{j}=1) for brevity, minimizing the negative log-likelihood (LUVLi loss) is also equivalent to minimizing the combination of KL-divergences given by

For the multivariate location distribution P{\mathcal{P}}, we consider two different models: Gaussian and Laplacian.

Gaussian Likelihood. The 22D Gaussian likelihood is:

In (11), T2T_{2} is the squared Mahalanobis distance, while T1T_{1} serves as a regularization or prior term that ensures that the Gaussian uncertainty distribution does not get too large.

Laplacian Likelihood. We use a 22D Laplacian likelihood given by:

In (13), T2T_{2} is a scaled Mahalanobis distance, while T1T_{1} serves as a regularization or prior term that ensures that the Laplacian uncertainty distribution does not get too large.

3 Uncertainty and Visibility Estimation

Our proposed method uses heatmaps for estimating landmarks’ locations, but not for estimating their uncertainty and visibility. We experimented with several methods for computing a covariance matrix directly from a heatmap, but none were accurate enough. We discuss this in Section 5.1.

Cholesky Estimator Network (CEN). We represent the uncertainty of each landmark location using a 2 ⁣× ⁣22\!\times\!2 covariance matrix Σij\bm{\Sigma}_{ij}, which is symmetric positive definite. The three degrees of freedom of Σij\bm{\Sigma}_{ij} are captured by its Cholesky decomposition: a lower-triangular matrix Lij\bm{L}_{ij} such that Σij ⁣= ⁣LijLijT\bm{\Sigma}_{ij}\!=\!\bm{L}_{ij}\bm{L}_{ij}^{T}. To estimate the elements of Lij\bm{L}_{ij}, we append a Cholesky Estimator Network (CEN) to the bottleneck of each U-net. The CEN is a fully connected linear layer whose input is the bottleneck of the U-net (128 ⁣× ⁣4 ⁣× ⁣4 ⁣= ⁣2, ⁣048(128\!\times\!4\!\times\!4\!=\!2,\!048 dimensions)) and output is an Np ⁣× ⁣3N_{p}\!\times\!3-dimensional vector, where NpN_{p} is the number of landmarks (e.g., 6868). As the Cholesky decomposition Lij\bm{L}_{ij} of a covariance matrix must have positive diagonal elements, we pass the corresponding entries of the output through an ELU activation function , to which we add a constant to ensure the output is always positive (asymptote is negative xx-axis).

Visibility Estimator Network (VEN). To estimate the visibility of the landmark vev_{e}, we add another fully connected linear layer whose input is the bottleneck of the U-net (128 ⁣× ⁣4 ⁣× ⁣4 ⁣= ⁣2, ⁣048(128\!\times\!4\!\times\!4\!=\!2,\!048 dimensions)) and output is an NpN_{p}-dimensional vector. This is passed through a sigmoid activation so the predicted visibility v^ij\widehat{v}_{ij} is between and 11.

The addition of these two fully connected layers only slightly increases the size of the original model. The loss for a single U-net is the averaged Lij\mathcal{L}_{ij} across all the landmarks j=1,...,Npj=1,...,N_{p} , and the total loss L\mathcal{L} for each input image is a weighted sum of the losses of all KK of the U-nets:

At test time, each landmark’s mean and Cholesky coefficients are derived from the KthK{\text{th}} (final) U-net. The covariance matrix is calculated from the Cholesky coefficients.

New Dataset: MERL-RAV

To promote future research in face alignment with uncertainty, we now introduce a new dataset with entirely new, manual labels of over 19, ⁣00019,\!000 face images from the AFLW dataset. In addition to landmark locations, every landmark is labeled with one of three visibility classes. We call the new dataset MERL Reannotation of AFLW with Visibility (MERL-RAV).

Visibility Classification. Each landmark of every face is classified as either unoccluded, self-occluded, or externally occluded, as illustrated in Figure 3. Unoccluded denotes landmarks that can be seen directly in the image, with no obstructions. Self-occluded denotes landmarks that are occluded because of extreme head pose—they are occluded by another part of the face (e.g., landmarks on the far side of a profile-view face). Externally occluded denotes landmarks that are occluded by hair or an intervening object such as a cap, hand, microphone, or goggles. Human labelers are generally very bad at localizing self-occluded landmarks, so we do not provide ground-truth locations for these. We do provide ground-truth (labeled) locations for both unoccluded and externally occluded landmarks.

Relationship to Visibility in LUVLi. In Section 3, visible landmarks (vj=1v_{j}=1) are landmarks for which ground-truth location information is available, while invisible landmarks (vj=0v_{j}=0) are landmarks for which no ground-truth location information is available (pj=∅\mathbf{p}_{j}=\emptyset). Thus, invisible (vj=0v_{j}=0) in the model is equivalent to the self-occluded landmarks in our dataset. In contrast, both unoccluded and externally occluded landmarks are considered visible (vj=1v_{j}=1) in our model. We choose this because human labelers are generally good at estimating the locations of externally occluded landmarks but poor at estimating the locations of self-occluded landmarks.

Existing Datasets. The most commonly used publicly available datasets for evaluation of 22D face alignment are summarized in Table 1. The 300300-W dataset uses a 6868-landmark system that was originally used for Multi-PIE . Menpo 22D makes a hard distinction (denoted F/P) between nearly frontal faces (F) and profile faces (P). Menpo 22D uses the same landmarks as 300300-W for frontal faces, but for profile faces it uses a different set of 3939 landmarks that do not all correspond to the 6868 landmarks in the frontal images. 300300W-LP-22D is a synthetic dataset created by automatically reposing 300300-W faces, so it has a large number of labels, but they are noisy. The 33D model locations of self-occluded landmarks are projected onto the visible part of the face as if the face were transparent (denoted by T). The WFLW and AFLW-6868 datasets do not identify which landmarks are self-occluded, but instead label self-occluded landmarks as if they were located on the visible boundary of the noseless face.

Differences from Existing Datasets. Our MERL-RAV dataset is the only one that labels every landmark using both types of occlusion (self-occlusion and external occlusion). Only one other dataset, AFLW, indicates which individual landmarks are self-occluded, but it has far fewer landmarks and does not label external occlusions. COFW and COFW-6868 indicate which landmarks are externally occluded but do not have self-occlusions. Menpo 22D categorizes faces as frontal or profile, but landmarks of the two classes are incompatible. Unlike Menpo 22D, our dataset smoothly transitions from frontal to profile, with gradually more and more landmarks labeled as self-occluded.

Our dataset uses the widely adopted 6868 landmarks used by 300300-W, to allow for evaluation and cross-dataset comparison. Since it uses images from AFLW, our dataset has pose variation up to ±120∘\pm 120^{\circ} yaw and ±90∘\pm 90^{\circ} pitch. Focusing on yaw, we group the images into five pose classes: frontal, left and right half-profile, and left and right profile. The train/test split is in the ratio of 4 ⁣: ⁣14\!:\!1. Table 2 provides the statistics of our MERL-RAV dataset. A sample image from the dataset is shown in Figure 3. In the figure, unoccluded landmarks are green, externally occluded landmarks are red, and self-occluded landmarks are indicated by black circles in the face schematic on the right.

Experiments

Our experiments use the datasets 300300-W , 300300W-LP-22D , Menpo 22D , COFW-6868 , AFLW-1919 , WFLW , and our MERL-RAV dataset. Training and testing protocols are described in the supplementary material. On a 1212 GB GeForce GTX Titan-X GPU, the inference time per image is 1717 ms.

Evaluation Metrics. We use the standard metrics NME, AUC, and FR . In each table, we report results using the same metric adopted in respective baselines.

Normalized Mean Error (NME). The NME is defined as:

Failure Rate (FR). FR refers to the percentage of images in the test set whose NME is larger than a certain threshold.

We train on the 300300-W , and test on 300300-W, Menpo 22D , and COFW-6868 . Some of the models are pre-trained on the 300300W-LP-22D .

Results: Localization and Cross-Dataset Evaluation. The face alignment results for 300300-W Split 11 and Split 22 are summarized in Table 3 and 4, respectively. Table 4 also shows the results of our model (trained on Split 22) on the Menpo and COFW-6868 datasets, as in . The results in Table 3 show that our LUVLi landmark localization is competitive with the SOTA methods on Split 11, usually one of the best two. Table 4 shows that LUVLi significantly outperforms the SOTA on Split 22, performing best on 55 out of the 66 cases (33 datasets ×\times 22 metrics). This is particularly impressive on 300300-W Split 22, because even though most of the other methods are pre-trained on the 300300W-LP-22D dataset (as was our best method, LUVLi*), our method without pre-training still outperforms the SOTA in 22 of 66 cases. Our method performs particularly well in the cross-dataset evaluation on the more challenging COFW-6868 dataset, which has multiple externally occluded landmarks.

Figure 7 shows that all three terms of our method’s predicted covariance matrices are highly predictive of the actual uncertainty: the mean squared residuals (error) are strongly proportional to the predicted covariance values, as evidenced by Pearson correlation coefficients of 0.980.98 and 0.990.99. However, decreasing NbinN_{\text{bin}} from 734 (plotted in Figure 7) to just 36 makes the correlation coefficients decrease to 0.84,0.80,0.720.84,0.80,0.72. Thus, the predicted uncertainties are excellent after averaging but may yet have room to improve.

Heatmaps vs. Direct Regression for Uncertainty. We tried multiple approaches to estimate the uncertainty distribution from heatmaps, but none of these worked nearly as well as our direct regression using the CEN. We believe this is because in current heatmap-based networks, the resolution of the heatmap (64×6464\times 64) is too low for accurate uncertainty estimation. This is demonstrated in Figure 8, which shows a histogram over all landmarks in 300300-W Test (Split 22) of LUVLi’s predicted covariance in the narrowest direction of the covariance ellipse (the smallest eigenvalue of the predicted covariance matrix). The figure shows that in most cases, the uncertainty ellipses are less wide than one heatmap pixel, which explains why heatmap-based methods are not able to accurately capture such small uncertainties.

2 AFLW-19 Face Alignment

On AFLW-1919, we train on 20, ⁣00020,\!000 images, and test on two sets: the AFLW-Full set (4, ⁣3864,\!386 test images) and the AFLW-Frontal set (1, ⁣3141,\!314 test images), as in . Table 6 compares our method’s localization performance with other methods that only train on AFLW-1919 (without training on any 6868-landmark dataset). Our proposed method outperforms not only the other uncertainty-based method KDN , but also all previous SOTA methods, by a significant margin on both AFLW-Full and AFLW-Frontal.

3 WFLW Face Alignment

Landmark localization results for WFLW are shown in Table 7. More detailed results on WFLW are in the supplementary material. Compared to the SOTA methods, LUVLi yields the second best performance on all metrics. Furthermore, while the other methods only predict landmark locations, LUVLi also estimates the prediction uncertainties.

4 MERL-RAV Face Alignment

Results of Landmark Localization. Results for all head poses on our MERL-RAV dataset are shown in Table 8.

5 Ablation Studies

Table 10 compares modifications of our approach on Split 22. Table 10 shows that computing the loss only on the last U-net performs worse than computing loss on all U-nets, perhaps because of the vanishing gradient problem . Moreover, LUVLi’s log-likelihood loss without visibility outperforms using MSE loss on the landmark locations (which is equivalent to setting all Σij=I\bm{\Sigma}_{ij}=\mathbf{I}). We also find that the loss with Laplacian likelihood (13) outperforms the one with Gaussian likelihood (11). Training from scratch is slightly inferior to first training the base DU-Net architecture before fine-tuning the full LUVLi network, consistent with previous observations that the model does not have strongly supervised pixel-wise gradients through the heatmap during training . Regarding the method for estimating the mean, using heatmaps is more effective than direct regression (Direct) from each U-net bottleneck, consistent with previous observations that neural networks have difficulty predicting continuous real values . As described in Section 3.1, in addition to ReLU, we compared two other functions for σ\sigma: softmax, and a temperature-scaled softmax (τ\tau-softmax). Results for temperature-scaled softmax and ReLU are essentially tied, but the former is more complicated and requires tuning a temperature parameter, so we chose ReLU for our LUVLi model. Finally, reducing the number of U-nets from 8 to 4 increases test speed by about 2 ⁣×2\!\times with minimal decrease in performance.

Conclusions

In this paper, we present LUVLi, a novel end-to-end trainable framework for jointly estimating facial landmark locations, uncertainty, and visibility. This joint estimation not only provides accurate uncertainty predictions, but also yields state-of-the-art estimates of the landmark locations on several datasets. We show that the predicted uncertainty distinguishes between unoccluded and externally occluded landmarks without any supervision for that task. In addition, the model achieves sub-pixel accuracy by taking the spatial mean of the ReLU’ed heatmap, rather than the arg max. We also introduce a new dataset containing manual labels of over 19, ⁣00019,\!000 face images with 6868 landmarks, which also labels every landmark with one of three visibility classes. Although our implementation is based on the DU-Net architecture, our framework is general enough to be applied to a variety of architectures for simultaneous estimation of landmark location, uncertainty, and visibility.

References

LUVLi Face Alignment: Estimating Landmarks’ Location, Uncertainty, and Visibility Likelihood Supplementary Material

Appendix A1 Implementation Details

Images are cropped using the detector bounding boxes provided by the dataset and resized to 256×256256\times 256. Images with no detector bounding box are initialized by adding 55% uniform noise to the location of each edge of the tight bounding box around the landmarks, as in .

Training. We modified the PyTorch code for DU-Net , keeping the number of U-nets K=8K=8 as in . Unless otherwise stated, we use the 22D Laplacian likelihood (12) as our landmark location likelihood, and therefore we use (13) as our final loss function. All U-nets have equal weights λi=1\lambda_{i}=1 in (14). For all datasets, visibility vj=1v_{j}=1 is assigned to unoccluded landmarks (those that are not labeled as occluded) and to landmarks that are labeled as externally occluded. Visibility vj ⁣= ⁣0v_{j}\!=\!0 is assigned to landmarks that are labeled as self-occluded and landmarks whose labels are missing.

Training images for 300300-W Split 11 are augmented randomly using scaling (0.75−1.25)(0.75-1.25), rotation (−<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><moseparator="true">,</mo><mo>−</mo></mrow><annotationencoding="application/x−tex">,−</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.7778em;vertical−align:−0.1944em;"></span><spanclass="mpunct">,</span><spanclass="mspace"style="margin−right:0.1667em;"></span><spanclass="mord">−</span></span></span></span></span>)(-<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo separator="true">,</mo><mo>−</mo></mrow><annotation encoding="application/x-tex">,-</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7778em;vertical-align:-0.1944em;"></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">−</span></span></span></span></span>) and color jittering (0.6,1.4)(0.6,1.4) as in , while those from 300300-W Split 22, AFLW-1919, WFLW-9898 and MERL-RAV datasets are augmented randomly using scaling (0.8−1.2)(0.8-1.2), rotation (−<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><moseparator="true">,</mo></mrow><annotationencoding="application/x−tex">,</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3em;vertical−align:−0.1944em;"></span><spanclass="mpunct">,</span></span></span></span></span>)(-<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo separator="true">,</mo></mrow><annotation encoding="application/x-tex">,</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3em;vertical-align:-0.1944em;"></span><span class="mpunct">,</span></span></span></span></span>), color jittering (0.6,1.4)(0.6,1.4), and random occlusion, as in .

The RMSprop optimizer is used as in , with batch size 2424. Training from scratch takes 100100 epochs and starts with learning rate 2.5×10−42.5\times 10^{-4}, which is divided by 55, 22, and 22 at epochs 3030, 6060, and 9090 respectively . When we initialize from pretrained weights, we finetune for 5050 epochs using the LUVLi loss: 2020 with learning rate 10−410^{-4}, followed by 3030 with learning rate 2 ⁣× ⁣10−52\!\times\!10^{-5}. We consider the model saved in the last epoch as our final model.

Testing. Whereas heatmap based methods adjust their pixel output with a quarter-pixel offset in the direction from the highest response to the second highest response, we use the spatial mean as the landmark location without carrying out any adjustment nor shifting the heatmap even by a quarter of a pixel. We do not need to implement a sub-pixel shift, because our spatial mean over the ReLUed heatmaps already performs sub-pixel location prediction.

Spatial Mean The spatial mean μij\bm{\mu}_{ij} of each of the heatmap is defined as

where \sigma\bigl{(}\bm{H}_{ij}(x,y)\bigr{)} denotes the output of post-processing the heatmap pixel with a function σ\sigma.

Appendix A2 Additional Experiments and Results

We now provide additional results evaluating our system’s performance in terms of both localization and uncertainty estimation.

For Split 11, we initialized using the pre-trained DU-Net model available from the authors of , then fine-tuned on the 300300-W training set (Split 11) using our proposed architecture and LUVLi loss. For Split 22, for the experiments in which we pre-trained on 300300W-LP-22D, we pre-trained from scratch on 300300W-LP-22D using heatmaps (using the original DU-Net architecture and loss). We then fine-tuned on the 300300-W training set (Split 22) using our proposed architecture and LUVLi loss.

A2.1.2 Comparison with KDN [13]

To compare directly with Chen et al. , in Figure 9 we plot normalized mean error (NME) vs. predicted uncertainty (rank, from smallest to largest), as in Figure 1 of . (We obtained the predicted uncertainty and NME data of from the authors.) The figure shows that for our method as well as for , there is a strong trend that higher predicted uncertainties correspond to larger location errors. However, the errors of our method are significantly smaller than the errors produced by .

A2.1.3 Verifying Predicted Uncertainty Distributions

For every image, for each landmark jj, our network predicts a mean μKj\bm{\mu}_{Kj} and a covariance matrix ΣKj\bm{\Sigma}_{Kj}. We can view this as our network predicting that a human labeler of that image will effectively select the landmark location pj\mathbf{p}_{j} for that image from the Laplacian distribution from (12) with mean μKj\bm{\mu}_{Kj} and covariance ΣKj\bm{\Sigma}_{Kj}:

If we had multiple labels (e.g., ground-truth landmark locations from multiple human labelers) for a single landmark in one image, then it would be straightforward to evaluate how well our method’s predicted probability distribution matches the distribution of labeled landmark locations. Unfortunately, face alignment datasets only have a single ground-truth location for each landmark in each image. This makes it difficult, but not impossible, to evaluate how well the human labels for images in the test set fit our method’s predicted uncertainty distributions. We propose the following method for verifying the predicted probability distributions.

Suppose we transform the ground-truth location of a landmark, pj\mathbf{p}_{j}, using the predicted mean and covariance for that landmark as follows:

If our method’s predictions are correct, then from (17), pj∼P(z∣μKj, ⁣ΣKj)\mathbf{p}_{j}\sim P({\bf z}|\bm{\mu}_{Kj},\!\bm{\Sigma}_{Kj}). Hence, pj′\mathbf{p}_{j}^{\prime} is drawn from the transformed distribution P(z′)P({\bf z}^{\prime}), where z′=ΣKj−0.5(z−μKj){\bf z}^{\prime}=\bm{\Sigma}_{Kj}^{-0.5}({\bf z}-\bm{\mu}_{Kj}):

After this simple transformation (transforming the labeled ground-truth location pj\mathbf{p}_{j} of each landmark using its predicted mean and covariance), we have transformed our network’s prediction about pj\mathbf{p}_{j} into a prediction about pj′\mathbf{p}_{j}^{\prime} that is much easier to evaluate, because the distribution in (19) is simply a standard 22D Laplacian distribution—it no longer depends on the predicted mean and covariance.

Thus, our method predicts that after the transformation (18), every ground-truth landmark location pj′\mathbf{p}_{j}^{\prime} is drawn from the same standard 22D Laplacian distribution (19). Now that we have an entire population of transformed labels that our model predicts are all drawn from the same distribution, it is easy to verify whether the labels fit our model’s predictions. Figure 10 shows a scatter plot of the transformed locations, pj′\mathbf{p}_{j}^{\prime}, for all landmarks in all test images of 300300-W (Split 22). We plot the histogram of the marginalized landmark locations (xx- or yy-coordinate of pj′\mathbf{p}_{j}^{\prime}) in orange above and to the right of the plot, and overlay the marginal pdf of the standard Laplacian (19) in black. The excellent match between the transformed landmark locations and the standard Laplacian distribution indicates that our model’s predicted uncertainty distributions are quite accurate. Since Kullback-Leibler (KL) divergence is invariant to affine transformations like the one in (18), we can evaluate the KL-divergence (printed at the top of the scatterplot) between the standard 22D Laplacian distribution (19) and the distribution of the transformed landmark locations (using their 22D histograms) as a numerical measure of how well the predictions of our model fit the distribution of labeled locations.

A2.1.4 Relationship to Variation Among Human Labelers on Multi-PIE

We test our Split 22 model on 812812 frontal face images of all subjects from the Multi-PIE dataset , then compute the mean of the uncertainty ellipses predicted by our model across all 812812 images. To compute the mean, we first normalize the location of each landmark using the inter-ocular distance, as in , and also normalize the covariance matrix by the square of the inter-ocular distance. We then take the average of the normalized locations across all faces to obtain the mean landmark location. The covariance matrices are averaged across all faces using the log-mean-exponential technique. The mean location and covariance matrix of each landmark (averaged across all faces) is then used to plot the results which are shown on the right in Figure 11.

We compare our model predictions with Figure 5 of , shown on the left of Figure 11. To create that figure, tasked three different human labelers with annotating the same frontal face images from the Multi-PIE database of 80 different subjects in frontal pose with neutral expression. For each landmark, they plotted the the covariance of the label locations across the three labelers using an ellipse. Note the similarity between our model’s predicted uncertainties (on the right of Figure 11 and the covariance across human labelers (on the left of Figure 11), especially around the eyes, nose, and mouth. Around the outside edge of the face, note that our model predicts that label locations will vary primarily in the direction parallel to the edge, which is precisely the pattern observed across human labelers.

A2.1.5 Sample Uncertainty Ellipses on Multi-PIE

To illustrate how the predicted uncertainties output by our method vary across different subjects from Multi-PIE, in Figure 12 we overlay our model’s mean uncertainty predictions (in blue, copied from right side of Figure 11) with our model’s predicted uncertainties of some of the individual Multi-PIE face images (in various colors). To simplify the figure, we plot all landmarks except for the eyes, nose, and mouth.

A2.1.6 Laplacian vs. Gaussian Likelihood

We have described two versions of our model: one whose loss function (13) uses a 22D Laplacian probability distribution (12), and another whose loss function (11) uses a 22D Gaussian probability distribution (10). We now discuss the question of which of these two models performs better.

The numerical comparisons are shown in Table 11. The numbers in the first two columns of the table were also presented in the ablation studies table, Table 10.

Comparing the Predicted Uncertainties. To compare the two models’ predicted uncertainties as well as their predicted locations, we consider the probability distributions over landmark locations that are predicted by each model. We want to know which model’s predicted probability distributions better explain the ground-truth locations of the landmarks in the test images. In other words, we want to know which model assigns a higher likelihood to the ground-truth landmark locations (i.e., which model yields a lower negative log-likelihood on the test data). We compute the negative log-likelihood of the ground-truth locations pj\mathbf{p}_{j} from the last hourglass using (13) for the Laplacian model and (11) for the Gaussian model. The results, in the last column of Table 11, show that the Laplacian model gives a lower negative log-likelihood. In other words, the ground-truth landmark locations have a higher likelihood under our Laplacian model. We conclude that the learned Laplacian model explains the human labels better than the learned Gaussian model.

A2.2 WFLW Face Alignment

Data Splits and Implementation Details. The training set consists of 7, ⁣5007,\!500 images, while the test set consists of 2, ⁣5002,\!500 images. In Table 13, we report results on the entire test set (All), which we also reported in Table 7. In Table 13, we additionally report results on several subsets of the test set: large head pose (326326 images), facial expression (314314 images), illumination (698698 images), make-up (206206 images), occlusion (736736 images), and blur (773773 images). The images are cropped using the detector bounding boxes provided by and resized to 256×256256\times 256.

Results of Facial Landmark Localization. Table 13 compares our method’s landmark localization results with those of other state-of-the-art methods on the WFLW dataset. Our method performs performs in the top two methods on all the metrics. Importantly, all of the other methods only predict landmark locations–they do not predict the uncertainty of their estimated landmark locations. Not only does our method place in the top two on all three landmark localization metrics, but our method also accurately predicts its own uncertainty of landmark localization.

A2.3 MERL-RAV Face Alignment

If all of the facial landmarks are visible, then this reduces to our previous definition of NME (15).

A2.4 Additional Qualitative Results

In Figure 13, we show example results on images from four datasets on which we tested.

A2.5 Video Demo of LUVLi

We include a short demo video of our LUVLi model that was trained on our new MERL-RAV dataset. The video demonstrates our method’s ability to predict landmarks’ visibility (i.e., whether they are self-occluded) as well as their locations and uncertainty. We take a simple face video of a person turning his head from frontal to profile pose and run our method on each frame independently. Overlaid on each frame of video, we plot each estimated landmark location in yellow, and plot the predicted uncertainty as a blue ellipse. To indicate the predicted visibility of each landmark, we modulate the transparency of the landmark (of the yellow dot and blue ellipse). Landmarks whose predicted visibility is close to 1 are shown as fully opaque, while landmarks whose predicted visibility is close to zero are fully transparent (are not shown). Landmarks with intermediate predicted visibilities are shown as partially transparent.

In the video, notice that as the face approaches the profile pose, points on the far edge of the face begin to disappear, because the method correctly predicts that they are not visible (are self-occluded) when the face is in profile pose.

A2.6 Examples from our MERL-RAV Dataset

Figure 14 shows several sample images from our MERL-RAV dataset. The ground-truth labels are overlaid on the images. On each image, unoccluded landmarks are shown in green, externally occluded landmarks are shown in red, and self-occluded landmarks are indicated by black circles in the face schematic to the right of the image.

Acknowledgements

We would like to thank Lisha Chen from RPI for providing results from their method and Zhiqiang Tang and Shijie Geng from Rutgers University for providing their pre-trained models on 300300-W (Split 11). We would also like to thank Adrian Bulat and Georgios Tzimiropoulos from the University of Nottingham for detailed discussions on getting bounding boxes for 300300-W (Split 22). We also had very useful discussions with Peng Gao from Chinese University of Hong Kong on the loss functions and Moitreya Chatterjee from University of Illinois Urbana-Champaign. We are also grateful to Maitrey Mehta from the University of Utah who volunteered for the demo. We also acknowledge anonymous reviewers for their feedback that helped in shaping the final manuscript.