Heatmap Regression via Randomized Rounding

Baosheng Yu, Dacheng Tao

Introduction

Semantic landmarks are sets of points or pixels in images containing rich semantic information. They reflect the intrinsic structure or shape of objects such as human faces , hands , bodies , and household objects . Semantic landmark localization is fundamental in computer and robot vision . For example, semantic landmark localization can be used to register correspondences between spatial positions and semantics (semantic alignment), which is extremely useful in visual recognition tasks such as face recognition and person re-identification . Therefore, robust and efficient semantic landmark localization is extremely important in applications requiring accurate semantic landmarks including robotic grasping and facial analysis applications such as face makeup , animation , and reenactment .

Coordinate regression and heatmap regression are two widely-used methods for deep learning-based semantic landmark localization . Rather than directly regressing the numerical coordinate with a fully-connected layer, heatmap-based methods aim to predict the heatmap where the maximum activation point corresponds to the semantic landmark in the input image. An intuitive example of heatmap representation is shown in Fig. 1. Due to the effective spatial generalization of heatmap representation, heatmap regression method is robust to large variations in pose, illumination, and occlusion in unconstrained settings . Heatmap regression has performed particularly well in semantic landmark localization tasks including facial landmark detection and human pose estimation . Despite this promise, heatmap regression method suffers from an inherent drawback, namely that the indices of the activation points in heatmaps are always integers. Vanilla heatmap-based methods therefore fail to predict the numerical coordinates in sub-pixel precision. Sub-pixel localization is nevertheless important in real-world scenarios with the fractional part of numerical coordinates originating from: 1) the input image being captured either by a low-resolution camera and/or at a relatively large distance; and 2) the heatmap is usually at a much lower resolution than the input image due to the downsampling stride of convolutional neural networks. As a result, low-resolution heatmaps significantly degrade heatmap regression performance. Considering that the computational cost of convolutional neural networks usually depends quadratically on the resolution of the input image or the feature map, there is a trade-off between the localization accuracy and the computational cost for heatmap regression . Furthermore, the downsampling stride of heatmap is not always equal to the downsampling stride of feature map: given an original image of 512×512512\times 512 pixels, a heatmap regression model with the input size 128×128128\times 128 pixels, and the feature map with a downsampling stride 44 pixels, we then have the size of heatmap 32×3232\times 32 pixels, i.e., the downsampling stride of heatmap s=16s=16 pixels. For simplicity, we do not distinguish between the above two settings to address the quantization error in a unified manner. Unless otherwise mentioned, we refer to s>1s>1 as the downsampling stride of the heatmap.

In vanilla heatmap regression, 1) during training, the ground truth numerical coordinates are first quantized to generate the ground truth heatmap; and 2) during testing, the predicted numerical coordinates can be decoded from the maximum activation point in the predicted heatmap. However, typical quantization operations such as floor, round, and ceil discard the fractional part of the ground truth numerical coordinates, making it difficult to reconstruct the fractional part even from the optimal predicted heatmap. This error induced by the transformation between numerical coordinates and heatmap is known as the quantization error. To address the problem of quantization error, here we introduce a new quantization system to form a lossless transformation between the numerical coordinates and the heatmap. In our approach, during training, the proposed quantization system uses a set of activation points, and the fractional part of the numerical coordinate is encoded as the activation probabilities of different activation points. During testing, the fractional part can then be reconstructed according to the activation probabilities of the top kk maximum activation points in the predicted heatmap. To achieve this, we introduce a new quantization operation called randomized rounding, or random-round, which is widely used in combinatorial optimization to convert fractional solutions into integer solutions with provable approximation guarantees . Furthermore, the proposed method can easily be implemented using a few lines of source code, making it a plug-and-play replacement for the quantization system of existing heatmap regression methods.

In this paper, we address the problem of quantization error in heatmap regression. The remainder of the paper is structured as follows. In the preliminaries, we briefly review two typical semantic landmark localization methods, coordinate regression and heatmap regression. In the methods, we first formally introduce the problem of quantization error by decomposing the prediction error into the heatmap error and the quantization error. We then discuss quantization bias in vanilla heatmap regression and prove a tight upper bound on the quantization error in vanilla heatmap regression. To address quantization error, we devise a new quantization system and theoretically prove that the proposed quantization system is unbiased and lossless. We also discuss uncertainty in heatmap prediction as well as the unbiased annotation when forming a robust semantic landmark localization system. In the experimental section, we demonstrate the effectiveness of our proposed method on popular facial landmark detection datasets (WFLW, 300W, COFW, and AFLW) and human pose estimation datasets (MPII and COCO).

Related Work

Semantic landmark localization, which aims to predict the numerical coordinates for a set of pre-defined semantic landmarks in a given image or video, has a variety of applications in computer and robot vision including facial landmark detection , hand landmark detection , human pose estimation , and household object pose estimation . In this section, we briefly review recent works on coordinate regression and heatmap regression for semantic landmark localization, especially in facial landmark localization applications.

Coordinate regression has been widely and successfully used in semantic landmark localization under constrained settings, where it usually relies on simple yet effective features . To improve the performance of coordinate regression for semantic landmark localization in the wild, several methods have been proposed by using cascade refinement , parametric/non-parametric shape models , multi-task learning , and novel loss functions .

2 Heatmap Regression

The success of deep learning has prompted the use of heatmap regression for semantic landmark localization, especially for robust and accurate facial landmark localization and human pose estimation . Existing heatmap regression methods either rely on large input images or empirical compensations during inference to mitigate the problem of quantization error . For example, a simple yet effective compensation method known as “shift a quarter to the second maximum activation point” has been widely used in many state-of-the-art heatmap regression methods .

Several methods have been developed to address the problem of quantization error in three aspects: 1) jointly predicting the heatmap and the offset in a multi-task manner ; 2) encoding and decoding the fractional part of numerical coordinates via a modulated 2D Gaussian distribution ; and 3) exploring differentiable transformations between the heatmap and the numerical coordinates . Specifically, generates the fractional part sensitive ground truth heatmap for video-based face alignment, which is known as fractional heatmap regression. Under the assumption that the predicted heatmap follows a 2D Gaussian distribution, decodes the fractional part of numerical coordinates from the modulated predicted heatmap. The soft-argmax operation is differentiable , and has been intensively explored in human pose estimation .

Preliminaries

In this section, we introduce two widely-used semantic landmark localization methods, coordinate regression and heatmap regression. For simplicity, we use facial landmark detection as an intuitive example.

Coordinate Regression. Given a face image, semantic landmark detection aims to find the numerical coordinates of a set of pre-defined facial landmarks xi=(xi,yi)\boldsymbol{x}_{i}=\left(x_{i},y_{i}\right), where i=1,2,…,Ki=1,2,\dots,K, indicate the indices of different facial landmarks (e.g., a set of five pre-defined facial landmarks can be left eye, right eye, nose, left mouth corner, and right mouth corner). It is natural to train a model (e.g., deep neural networks) to directly regress the numerical coordinates of all facial landmarks. The coordinate regression model then can be optimized via a typical regression criterion such as mean squared error (MSE) and mean absolute error (MAE). For the MSE criterion (also the L2 loss), we have

where xip\boldsymbol{x}^{p}_{i} and xig\boldsymbol{x}^{g}_{i} indicate the predicted and the ground truth numerical coordinates, respectively. When using the MAE criterion (also the L1 loss), the loss function L\mathcal{L} can be defined in a similar way to (1).

Heatmap Regression. Heatmaps (also known as confidence maps) are simple yet effective representations of semantic landmark locations. Given the numerical coordinate xi\boldsymbol{x}_{i} for the ii-th semantic landmark, it then corresponds to a specific heatmap hi\boldsymbol{h}_{i} as shown in Fig. 1. For simplicity, we assume hi\boldsymbol{h}_{i} is the same size as the input image in this section and leave the problem of quantization error to the next section. With the heatmap representation, the problem of semantic landmark localization can be translated into heatmap regression via two heatmap subroutines: 1) encode (from the ground truth numerical coordinate xig\boldsymbol{x}_{i}^{g} to the ground truth heatmap hig\boldsymbol{h}_{i}^{g}); and 2) decode (from the predicted heatmap hip\boldsymbol{h}_{i}^{p} to the predicted numerical coordinate xip\boldsymbol{x}_{i}^{p}). The main deep learning-based heatmap regression for semantic landmark localization framework is shown in Fig. 2.

Therefore, with the decode operation in (2), the problem of semantic landmark localization can be solved by training a deep model to predict heatmap hip\boldsymbol{h}_{i}^{p}.

where Σ\Sigma is the covariance matrix (a positive semi-definite matrix) and σ>0\sigma>0 is the standard deviation in both directions, i.e.,

When σ → 0\sigma~{}\to~{}0, the ground truth heatmap can be generated by assigning a positive value at the ground truth numerical coordinate xig\boldsymbol{x}_{i}^{g}, i.e.,

Specifically, when σ→0\sigma\to 0, the ground truth heatmap defined in (5) is also known as the binary heatmap.

Given the ground truth heatmap, the heatmap regression model then can be optimized using typical pixel-wise regression criteria such as MSE, MAE, or Smooth-L1 . Specifically, for Gaussian heatmaps, the heatmap regression model is usually optimized with the pixel-wise MSE criterion, i.e.,

When using the MAE/Smooth-L1 criteria, the loss function can be defined in a similar way to (6). For binary heatmap, the heatmap regression model can also be optimized with the pixel-wise cross-entropy criterion, i.e.,

where LCE\mathcal{L}_{\text{CE}} indicates the cross-entropy criterion with a softmax function as the activation/normalization function. A comprehensive review of different loss functions for semantic landmark localization is beyond the scope of this paper, but we refer interested readers to for descriptions of coordinate regression and for heatmap regression. Unless otherwise mentioned, we use the MSE criterion for Gaussian heatmap and the cross-entropy criterion for binary heatmap in this paper.

Method

In this section, we first introduce the quantization system in heatmap regression and then formulate the quantization error in a unified way by correcting the quantization bias in a vanilla quantization system. Lastly, we devise a new quantization system via randomized rounding to address the problem of quantization error.

Heatmap regression for semantic landmark localization usually contains two key components: 1) heatmap prediction; and 2) transformation between the heatmap and the numerical coordinates. The quantization system in heatmap regression is a combination of the encode and decode operations. During training, when the ground truth numerical coordinates xig\boldsymbol{x}_{i}^{g} are floating-point numbers, we then need to calculate a specific Gaussian kernel matrix using (3) for each landmark, since different numerical coordinates usually have different fractional parts. As a result, it will significantly increase the training loads of the heatmap regression model. For example, given 9898 landmarks per face image, the kernel size 11×1111\times 11, and a mini-batch of 1616 training samples, we then have to run (3) for 98×16×11×11=189,72898\times 16\times 11\times 11=189,728 times in each training iteration. To address this issue, existing heatmap regression methods usually first quantize numerical coordinates into integers, where a standard kernel matrix can then be shared for efficient ground truth heatmap generation . However, the above-mentioned existing heatmap regression methods usually suffer from the inherent drawback of failing to encode the fractional part of numerical coordinates. Therefore, how to efficiently encode the fractional information in numerical coordinates still remains challenging. Furthermore, during the inference stage, the predicted numerical coordinates xip\boldsymbol{x}_{i}^{p} obtained by a decode operation in (2) are also integers. As a result, typical heatmap regression methods usually fail to efficiently address the fractional part of the numerical coordinate during both training and inference, resulting in localization error.

To analyze the localization error caused by the quantization system in heatmap regression, we formulate the localization error as the sum of heatmap error and quantization error as follows:

where xiopt\boldsymbol{x}_{i}^{opt} indicates the numerical coordinate decoded from the optimal predicted heatmap. Generally, the heatmap error corresponds to the error in heatmap prediction, i.e., ∥hip−hig∥2\|\boldsymbol{h}_{i}^{p}-\boldsymbol{h}_{i}^{g}\|_{2}, and the quantization error indicates the error caused by both the encode and decode operations. If there is no heatmap error, the localization error then all originates from the error of the quantization system, i.e.,

The generalizability of deep neural networks for heatmap prediction, i.e., the heatmap error, is beyond the scope of this paper. We do not consider the heatmap error during the analysis of quantization error in this paper.

That is, for integer quantization operations floor, round, and ceil, we have t=1.0t=1.0, t=0.5t=0.5, and t=0t=0, respectively. Furthermore, when the downsampling stride s>1s>1, the decode operation in (2) then becomes

A vanilla quantization system for heatmap regression can then be formed by the encode operation in (10) and the decode operation in (11). When applied to a vector or a matrix, the integer quantization operation defined in (10) is an element-wise operation.

2 Quantization Error

In this subsection, we first correct the bias in a vanilla quantization system to form an unbiased vanilla quantization system. With the unbiased quantization system, we then provide a tight upper bound on the quantization error for vanilla heatmap regression.

Let ϵx\epsilon_{x} denote the fractional part of xig/sx_{i}^{g}/s, and ϵy\epsilon_{y} denote the fractional part of yig/sy_{i}^{g}/s. Given the downsampling stride of the heatmap s>1s>1, we then have

Given the assumption of a “perfect” heatmap prediction model or no heatmap error, i.e., hip(x)=hig(x)\boldsymbol{h}_{i}^{p}(\boldsymbol{x})=\boldsymbol{h}_{i}^{g}(\boldsymbol{x}), we then have the predicted numerical coordinates

Considering that xig,yigx_{i}^{g},y_{i}^{g} are independent variables, we thus have the quantization bias in the vanilla quantization system as follows:

Therefore, only the encode operation in (10), i.e., the round operation, is unbiased. Furthermore, given ∀t ∈ \forall t~{}\in~{} for the encode operation in (10), we can correct the bias of the encode operation with a shift on the decode operation, i.e.,

For simplicity, we use the round operation, i.e., t=0.5t=0.5, to form an unbiased quantization system as our baseline. Though the vanilla quantization system defined by (10) and (13) is unbiased, it causes non-invertible localization error. An intuitive explanation for this is that the encode operation in (10) directly discards the fractional part of the ground truth numerical coordinates, thus making it impossible for the decode operation to accurately reconstruct the numerical coordinates.

Given an unbiased quantization system defined by the encode operation in (10) and the decode operation in (13), we then have that the quantization error tightly upper bounded, i.e.,

where s>1s>1 indicates the downsampling stride of the heatmap.

From Theorem 1, we know that the vanilla quantization system defined by (10) and (13) will cause non-invertible quantization error and that the upper bound of the quantization error linearly depends on the downsampling stride of the heatmap. As a result, given the heatmap regression model, it will cause extremely large localization error for large faces in the original input image, making it a significant problem in many important face-related applications such as face makeup, face swapping, and face reenactment.

3 Randomized Rounding

In vanilla heatmap regression, each numerical coordinate corresponds to a single activation point in the heatmap, while the indices of the activation point are all integers. As a result, the fractional part of the numerical coordinate is usually ignored during the encode process, making it an inherent drawback of heatmap regression for sub-pixel localization. To retain the fractional information when using heatmap representations, we utilize multiple activation points around the ground truth activation point. Inspired by the randomized rounding method , we address the quantization error in vanilla heatmap regression by using a probabilistic approach. Specifically, we encode the fractional part of the numerical coordinate to different activation points with different activation probabilities. An intuitive example is shown in Fig. 3.

We describe the proposed quantization system as follows. Given the ground truth numerical coordinate xig=(xig,yig)\boldsymbol{x}_{i}^{g}=(x_{i}^{g},y_{i}^{g}) and a downsampling stride of the heatmap s>1s>1, the ground truth activation point in the heatmap is (xig/s,yig/s)(x_{i}^{g}/s,y_{i}^{g}/s), which are usually floating-point numbers, and we are unable to find the corresponding pixel in the heatmap. If we ignore the fractional part (ϵx,ϵy)(\epsilon_{x},\epsilon_{y}) using a typical integer quantization operation, e.g., round, the ground truth activation point will be approximated by one of the activation points around the ground truth activation point, i.e., (⌊xig/s⌋,⌊yig/s⌋)\left(\lfloor x_{i}^{g}/s\rfloor,\lfloor y_{i}^{g}/s\rfloor\right), (⌊xig/s⌋+1,⌊yig/s⌋)\left(\lfloor x_{i}^{g}/s\rfloor+1,\lfloor y_{i}^{g}/s\rfloor\right), (⌊xig/s⌋,⌊yig/s⌋+1)\left(\lfloor x_{i}^{g}/s\rfloor,\lfloor y_{i}^{g}/s\rfloor+1\right), and (⌊xig/s⌋+1,⌊yig/s⌋+1)\left(\lfloor x_{i}^{g}/s\rfloor+1,\lfloor y_{i}^{g}/s\rfloor+1\right). However, the above process is not invertible. To address this, we randomly assign the ground truth activation point to one of the alternative activation points around the ground truth activation point, and the activation probability is determined by the fractional part of the ground truth activation point as follows:

To achieve the encode scheme in (14) in conjunction with current minibatch stochastic gradient descent training algorithms for deep learning models, we introduce a new integer quantization operation via randomized rounding, i.e., random-round:

Given the encode operation in (15), if we do not consider the heatmap error, we then have the activation probability at x\boldsymbol{x}:

As a result, the fractional part of the ground truth numerical coordinate (ϵx,ϵy)(\epsilon_{x},\epsilon_{y}) can be reconstructed from the predicted heatmap via the activation probabilities of all activation points, i.e.,

where Xig\mathcal{X}_{i}^{g} indicates the set of activation points around the ground truth activation point, i.e.,

Given the encode operation in (15) and the decode operation in (17), we then have that the 1) encode operation is unbiased; and 2) quantization system is lossless, i.e., there is no quantization error.

From Theorem 2, we know that the quantization system defined by the encode operation in (15) and the decode operation in (17) is unbiased and lossless.

4 Activation Points Selection

The fractional information of the numerical coordinate (ϵx,ϵy)(\epsilon_{x},\epsilon_{y}) is well-captured by the randomized rounding operation, allowing us to reconstruct the ground truth numerical coordinate xig\boldsymbol{x}_{i}^{g} without the quantization error. However, during the inference phase, the ground truth numerical coordinate xig\boldsymbol{x}_{i}^{g} is unavailable and heatmap error always exists in practice, making it difficult to identify the proper set of ground truth activation points Xig\mathcal{X}_{i}^{g}. In this section, we describe a way to form a set of alternative activation points in practice.

We introduce two activation point selection methods as follows. The first solution is to estimate all activation points via the points around the maximum activation point. As shown in Fig. 4, given the maximum activation point, we then have four different sets of alternative activation points, Xig1,Xig2,Xig3,and Xig4\mathcal{X}_{i}^{g_{1}},\mathcal{X}_{i}^{g_{2}},\mathcal{X}_{i}^{g_{3}},\text{and}~{}\mathcal{X}_{i}^{g_{4}}. Therefore, given the predicted heatmap in practice, we then take a risk of choosing an incorrect set of alternative activation points. To find a robust set of alternative activation points, we may use all nine activation points around the maximum activation point, i.e.,

Another solution of alternative activation points Xig\mathcal{X}_{i}^{g} is to generalize the argmax operation with the argtopk operation, i.e., we decode the predicted heatmap hip\boldsymbol{h}_{i}^{p} to obtain the numerical coordinate xip\boldsymbol{x}_{i}^{p} according to the top kk largest activation points,

If there is no heatmap error, the two alternative activation points solutions presented above, i.e., the alternative activation points in (18) and (19), are equal to each other when using the decode operation in (17). Specifically, we find that the activation points in (19) achieve comparable performance to the activation points in (20) when k=9k=9. For simplicity, unless otherwise mentioned, we use the set of alternative activation points defined by (20) in this paper. Furthermore, when we take the heatmap error into consideration, the values of different kk then forms a trade-off on the selection of activation points, i.e., a larger kk will be robust to activation point selection whilst also increasing the risk of noise from the heatmap error. See more discussion in Section 4.5 and the experimental results in Section 5.5.

5 Discussion

In this subsection, we provide some insights into the proposed quantization system with respect to: 1) the influence of human annotators on the proposed quantization system in practice; and 2) the underlying explanation behind the widely used empirical compensation method “shift a quarter to the second maximum activation point”.

Unbiased Annotation. We assume the “ground truth numerical coordinates” are always accurate, while the ground truth numerical coordinates are usually labelled by human annotators at the risk of the annotation bias. Given an input image, the ground truth numerical coordinates xig\boldsymbol{x}_{i}^{g} can be obtained by clicking a specific pixel in the image, which is a simple but effective annotation pipeline provided by most image annotation tools. For sub-pixel numerical coordinates, especially in low-resolution input images, the annotators may click either one of all possible pixels around the ground truth numerical coordinates due to human visual uncertainty. As shown in Fig. 5, clicking any one of the four possible pixels causes annotation error, which corresponds to the fractional part of the ground truth numerical coordinate (ϵx′,ϵy′)=(xig−⌊xig⌋, yig−⌊yig⌋)(\epsilon_{x}^{\prime},\epsilon_{y}^{\prime})=\left(x_{i}^{g}-\lfloor x_{i}^{g}\rfloor,~{}y_{i}^{g}-\lfloor y_{i}^{g}\rfloor\right). Given enough data samples, if the annotators click the pixel according to the following distribution, i.e.,

the fractional part then can be well captured by the heatmap regression model and we refer to it as an unbiased annotation.

If we take the downsampling stride into consideration, (ϵx,ϵy)(\epsilon_{x},\epsilon_{y}) is then a joint result of both the downsampling of the heatmap and the annotation process, i.e.,

On the one hand, if the heatmap regression model uses a low input resolution (or a large downsampling stride s≫1s\gg 1), the fractional part (ϵx,ϵy)(\epsilon_{x},\epsilon_{y}) then mainly comes from the downsampling of the heatmap; on the other hand, if the heatmap regression model uses a high input resolution, the annotation process will also have a significant influence on the heatmap regression. Therefore, when using a high input resolution model in practice, a diverse set of human annotators help reduce the bias in annotation process.

Empirical Compensation. “Shift a quarter to the second maximum activation point” has become an effective and widely used empirical compensation method for heatmap regression , but it still lacks a proper explanation. We thus provide an intuitive explanation according to the proposed quantization system. The proposed quantization system encodes the ground truth numerical coordinates into multiple activation points, and the activation probability of each activation point is decided by the fractional part, i.e., the activation probability indicates the distance between the activation point and the ground truth activation point. Therefore, the ground truth activation point is closer to the ii-th maximum activation point than the (i+1)(i+1)-th maximum activation point. We demonstrate the averaged activation probabilities for the top kk activation points on the WFLW dataset in Table I.

We find that the marginal improvement decreases as the number of activation points increases, i.e., the second maximum activation points provides the maximum improvement to the reconstruction of the fractional part. This observation partially explains the effectiveness of the compensation method “shift a quarter to the second maximum activation point”, which can be seen as a special case of the proposed method (20) with k=2k=2.

Furthermore, the proposed quantization system shares the same motivation with the bilinear interpolation. Specifically, the bilinear interpolation usually aims to find the value of the unknown function f(x,y)f(x,y) given its neighbors f(x1,y1)f(x_{1},y_{1}), f(x1,y2)f(x_{1},y_{2}), f(x2,y1)f(x_{2},y_{1}), and f(x2,y2)f(x_{2},y_{2}). For the proposed quantization system, we have f(x,y)=(x,y)f(x,y)=(x,y), which indicates the location of landmarks. Specifically, if there is no heatmap error, we then have x1=⌊xig/s⌋x_{1}=\lfloor x_{i}^{g}/s\rfloor, x2=⌊xig/s⌋+1x_{2}=\lfloor x_{i}^{g}/s\rfloor+1, y1=⌊yig/s⌋y_{1}=\lfloor y_{i}^{g}/s\rfloor, and y2=⌊yig/s⌋+1y_{2}=\lfloor y_{i}^{g}/s\rfloor+1. If we take heatmap error into consideration, the ground truth activation points are usually unknown. Therefore, the number of alternative activation points also controls the trade-off between the robustness of the quantization system and the risk of noise from the heatmap error (Also see details in Section 4.4).

Facial Landmark Detection

In this section, we perform facial landmark detection experiments. We first introduce widely used facial landmark detection datasets. We then describe the implementation details of our proposed method. Finally, we present our experimental results on different datasets and perform comprehensive ablation studies on the most challenging dataset.

We use four widely used facial landmark detection datasets:

WFLW . WFLW contains 10,00010,000 face images, including 7,5007,500 training images and 2,5002,500 testing images, with 9898 manually annotated facial landmarks. All face images are selected from the WIDER Face dataset , which contains face images with large variations in scale, expression, pose, and occlusion.

300W . 300W contains 3,1483,148 training images, including 337337 images from AFW , 2,0002,000 images from the training set of HELEN , and 811811 images from the training set of LFPW . For testing, there are four different settings: 1) common: 554554 images, including 330330 and 224224 images from the testsets of HELEN and LFPW, respectively; 2) challenge: 135135 images from IBUG; 3) full: 689689 images as a combination of common and challenge; and 4) private: 600600 indoor/outdoor images. All images are manually annotated with 6868 facial landmarks.

COFW . COFW contains 1,8521,852 images, including 1,3451,345 training and 507507 testing images. All images are manually annotated with 2929 facial landmarks.

AFLW . AFLW contains 24,38624,386 face images, including 20,00020,000 images for training and 4,8364,836 images for testing. For testing, there are two settings: 1) full: all 4,8364,836 images for testing; and 2) front: 1,3141,314 frontal images selected from the full set. All images are manually annotated with 2121 facial landmarks. For fair comparison, we use 1919 facial landmarks, i.e., the landmarks on two ears are ignored.

2 Evaluation Metrics

We use the normalized mean error (NME) as the evaluation metric in this paper, i.e.,

where dd indicates the normalization distance. For fair comparison, we report the performances on WFLW, 300W, and COFW using two the normalization methods, inter-pupil distance (the distance between the eye centers) and inter-ocular distance (the distance between the outer eye corners). We report the performance on AFLW using the size of the face bounding box as the normalization distance, i.e., the normalization distance can be evaluated by d=w∗hd=\sqrt{w*h}, where ww and hh indicate the width and height of the face bounding box, respectively.

3 Implementation Details

We implement the proposed heatmap regression method for facial landmark detection using PyTorch . Following the practice in , we use HRNet as our backbone network, which is an efficient counterpart of ResNet , U-Net , and Hourglass for semantic landmark localization. Unless otherwise mentioned, we use HRNet-W18 as the backbone network in our experiments. All face images are cropped and resized to 256×256256\times 256 pixels and the downsampling stride of the feature map is 44 pixels. For training, we perform widely-used data augmentation for facial landmark detection as follows We horizontally flip all training images with probability 0.50.5 and randomly change the brightness (±0.125\pm 0.125), contrast (±0.5\pm 0.5), and saturation (±0.5\pm 0.5) of each image. We then randomly rotate the image (±30∘\pm 30^{\circ}), rescale the image (±0.25\pm 0.25), and translate the image (±16\pm 16 pixels). We also randomly erase a rectangular region in the training image . All our models are initialized from the weights pretrained on ImageNet . We use the Adam optimizer with batch size 1616. The learning rate starts from 0.0010.001 and is divided by 1010 for every 6060 epochs, with 150150 training epochs in total. During the testing phase, we horizontally flip testing images as the data augmentation and average the predictions.

4 Comparison with Current State-of-the-Art

To demonstrate the effectiveness of the proposed method, we compare it with recent state-of-the-art facial landmark detection methods as follows. As shown in Table II, the proposed method outperforms recent state-of-the-art methods on the most challenging dataset, WFLW, with a clear margin for all different settings. For the 300W dataset, we try to report the performances under different settings for fair comparison. As shown in Table III, the proposed method achieves comparable performances for all different settings. Specifically, LAB uses the boundary information as the auxiliary supervision; compared to Wing , which uses the coordinate regression for semantic landmark localization, the heatmap-based methods usually achieve better performance on the challenge subset. In Table IV, we see that the proposed method outperforms recent state-of-the-art methods with a clear margin on COFW dataset. AFLW captures a wide range of different face poses, including both frontal faces and non-frontal faces. As shown in Table V, the proposed method achieves consistent improvements for both frontal faces and non-frontal faces, suggesting robustness across different face poses.

5 Ablation Studies

To better understand the proposed quantization system in different settings, we perform ablation studies the most challenging dataset, WFLW .

The influence of different input resolutions. Heatmap regression models use a fixed input resolution, e.g., 256×256256\times 256 pixels, but training and testing images usually represent a wide range of resolutions, e.g., most faces in the WFLW dataset have an inter-ocular distance of between 3030 and 120120 pixels. Therefore, we compare the proposed method with the baseline using an input resolution from 64×6464\times 64 pixels to 512×512512\times 512 pixels, i.e., a heatmap resolution from 16×1616\times 16 pixels to 128×128128\times 128 pixels. In Fig. 6, the proposed method significantly improves heatmap regression performance when using a low input resolution. The increasing number of high-resolution images/videos in real-world applications is a challenge with respect to the computational cost and device memory needed to overcome the problem of sub-pixel localization by increasing the input resolution of deep learning-based heatmap regression models. For example, in the film industry, it has sometimes become necessary to swap the appearance of a target actor and a source actor to generate higher fidelity video frames in visual effects, especially when an actor is unavailable for some scenes . The manipulation of actor faces in video frames relies on accurate localization of different facial landmarks and is performed at megapixel resolution, inducing a huge computational cost for extensive frame-by-frame animation. Therefore, instead of using high-resolution input images, our proposed method delivers another efficient solution for dealing with accurate semantic landmark localization.

The influence of different numbers of alternative activation points. In the proposed quantization system, the activation probability indicates the distance between activation point to the ground truth activation point (xig/s,yig/s)(x_{i}^{g}/s,y_{i}^{g}/s). If there is no heatmap error, the alternative activation points in (20) then give the same result as in (18). If the heatmap error cannot be ignored, there will be a trade-off on the number of alternative activation points: 1) a small kk increases the risk of missing the ground truth alternative activation points; 2) a large kk introduces the noise from irrelevant activation points, especially for large heatmap error. We demonstrate the performance of the proposed method by using different numbers of alternative activation points in Table VI. Specifically, we see that 1) when using a high input resolution, the best performance is achieved with a relatively large kk; and 2) the performance is smooth near the optimal number of alternative activation points, making it easy to find a proper kk for validation data. As introduced in Section 3, binary heatmaps can be seen as a special case of Gaussian heatmaps with standard deviation σ=0\sigma=0. Considering that Gaussian heatmaps have been widely used in semantic landmark localization applications, we generalize the proposed quantization system to Gaussian heatmaps and demonstrate the influence of different numbers of alternative activation points in Table VII. Specifically, we see that 1) when applying the proposed quantization system to the model using Gaussian heatmap, it achieves comparable performance to the model using binary heatmap; and 2) the optimal number of alternative activation points increases with the standard deviation σ\sigma.

The influence of different “bounding box” annotation policies. For facial landmark detection, a reference bounding box is required to indicate the position of the facial area. However, there is a performance gap when using different reference bounding boxes . A comparison between two widely used “bounding box” annotation policies is shown in Fig. 7, and we introduce two different bounding box annotation policies as follows:

P1: This annotation policy is usually used in semantic landmark localization tasks, especially in facial landmark localization. Specifically, the rectangular area of the bounding box tightly encloses a set of pre-defined facial landmarks.

P2: This annotation policy has been widely used in face detection datasets . The bounding box contains the areas of the forehead, chin, and cheek. For the occluded face, the bounding box is estimated by the human annotator based on the scale of occlusion.

We demonstrate the experimental results using different annotation policies in Table VIII. Specifically, we find that the policy P1 usually achieves better results, possibly because the occluded forehead (e.g., hair) introduces additional variations to the face bounding boxes when using the policy P2. Furthermore, the model trained using the policy P2 is more robust to different bounding box policy during testing, suggesting its robustness to inaccurate bounding boxes from face detection algorithms.

Qualitative Analysis. As shown in Fig. 8, we present some “good” and “bad” facial landmark detection examples according to NME. Specifically, for the good cases presented in the first row, most images are of high quality; For the bad cases in the second row, most images contain heavy blurring and/or occlusion, making it difficult to accurately identify the contours of different facial parts.

Human Pose Estimation

In this section, we perform human pose estimation experiments to further demonstrate the effectiveness of the proposed quantization system for accurate semantic landmark localization.

We perform experiments on two popular human pose estimation datasets,

MPII : The MPII Human Pose dataset contains around 28,821 images with 40,522 person instances, in which 11,701 images for testing and the remaining 17,120 images for training. Following the experimental setup in , we use 22,246 person instances for training and evaluate the performance on the MPII validation set with 2,958 person instances, which is a heldout subset of MPII training set.

COCO : The COCO dataset contains over 200,000 images and 250,000 person instances, in which each person instance is labeled with 17 keypoints. Following the experimental setup in , we evaluate the proposed method on the validation set with 5,000 images.

2 Implementation Details

We utilize recent state-of-the-art heatmap regression method for human pose estimation, HRNet , as our baseline. Specifically, the proposed quantization system can be easily integrated into most heatmap regression models and we have made the source code of human pose estimation based on the HRNet baseline publicly available. For the MPII Human Pose dataset, we use the standard evaluation metric, head-normalized probability of correct keypoint or PCKh . Specifically, a correct keypoint should fall within α∗l\alpha*l pixels of the ground truth position, where ll indicates the normalization distance and α∈\alpha\in is the matching threshold. For fair comparison, we apply two different matching thresholds, PCKh@0.5 and PCKh@0.1, where a smaller matching threshold, α=0.1\alpha=0.1, indicates a more strict evaluation metric for accurate semantic landmark localization . For the COCO dataset, we use the standard evaluation metric, averaged precision (AP) and averaged recall (AR), where the object keypoint similarity or OKS is used as the similarity measure between the ground truth objects and the predicted objects .

3 Results

The experimental results on the MPII dataset are shown in Table IX. Specifically, when using a coarse evaluation metric, PCKh@0.5, both the proposed method and the compensation method achieve comparable performance to the baseline method, suggesting that the quantization error is trivial in coarse semantic landmark localization; When using a more strict evaluation metric, PCKh@0.1, the compensation method, which can be seen as a special case of H3R with k=2k=2, significantly improves the baseline, e.g., from 32.832.8 to 37.737.7, while the proposed method H3R further improves the performance from 37.737.7 to 39.339.3. The experimental results on the COCO dataset are shown in Table X. Specifically, the proposed method clearly improves the averaged precision (AP) in different settings and the major improvements on AP come from: 1) a strict evaluation metric, e.g., AP0.75; and 2) large/medium person instances, i.e., AP(M) and AP(L). Furthermore, we also find that the improvement decreases when increasing the input resolution, e.g., from 0.6740.674 to 0.7200.720 for 192×128 (0.046↑)192\times 128~{}(0.046{\uparrow}), from 0.7230.723 to 0.7500.750 for 256×192 (0.027↑)256\times 192~{}(0.027{\uparrow}), and from 0.7480.748 to 0.7620.762 for 384×288 (0.014↑)384\times 288~{}(0.014{\uparrow}).

Conclusion

In this paper, we address the problem of sub-pixel localization for heatmap-based semantic landmark localization. We formally analyze quantization error in vanilla heatmap regression and propose a new quantization system via randomized rounding operation, which we prove is unbiased and lossless. Experiments on facial landmark localization and human pose estimation datasets demonstrate the effectiveness of the proposed quantization system for efficient and accurate sub-pixel localization.

Acknowledgement

Dr. Baosheng Yu is supported by ARC project FL-170100117.

References

Appendix A Proofs of Theorem 1 and Theorem 2

Given an unbiased quantization system defined by the encode operation in (10) and the decode operation in (13), we then have that the quantization error tightly upper bounded, i.e.,

where s>1s>1 indicates the downsampling stride of the heatmap.

Given the ground truth numerical coordinate xig=(xig,yig)\boldsymbol{x}_{i}^{g}=\left(x_{i}^{g},y_{i}^{g}\right), the predicted numerical coordinate xip=(xip,yip)\boldsymbol{x}_{i}^{p}=\left(x_{i}^{p},y_{i}^{p}\right), and the downsampling stride of the heatmap s>1s>1, if there is no heatmap error, we then have

where hip(x)\boldsymbol{h}_{i}^{p}(\boldsymbol{x}) and hig(x)\boldsymbol{h}_{i}^{g}(\boldsymbol{x}) indicate the ground truth heatmap and the predicted heatmap, respectively. Therefore, according to the decode operation in (13), we have the predicted numerical coordinate as

where ϵx=xig/s−⌊xig/s⌋\epsilon_{x}=x_{i}^{g}/s-\lfloor x_{i}^{g}/s\rfloor and ϵy=yig/s−⌊yig/s⌋\epsilon_{y}=y_{i}^{g}/s-\lfloor y_{i}^{g}/s\rfloor. The quantization error of vanilla quantization system then can be evaluated as follows:

The maximum quantization error ∣xip−xig∣=s/2|x_{i}^{p}-x_{i}^{g}|=s/2 is achieved when ϵx=t\epsilon_{x}=t. Similarly, we have the maximum quantization error ∣yip−yig∣=s/2|y_{i}^{p}-y_{i}^{g}|=s/2 is achieved with ϵy=t\epsilon_{y}=t. Considering that xipx_{i}^{p} and yipy_{i}^{p} are linearly independent variables, we thus have

The maximum quantization error is achieved with ϵx=ϵy=t\epsilon_{x}=\epsilon_{y}=t. That is, the quantization error in vanilla quantization system is tightly upper bounded by 2s/2\sqrt{2}s/2. ∎

Given the encode operation in (15) and the decode operation in (17), we then have that the 1) encode operation is unbiased; and 2) quantization system is lossless, i.e., there is no quantization error.

Given the ground truth numerical coordinate xig=(xig,yig)\boldsymbol{x}_{i}^{g}=\left(x_{i}^{g},y_{i}^{g}\right), the predicted numerical coordinate xip=(xip,yip)\boldsymbol{x}_{i}^{p}=\left(x_{i}^{p},y_{i}^{p}\right), and the downsampling stride of the heatmap s>1s>1, we then have

Therefore, the encode operation in (15), i.e., random-round, is an unbiased encode operation for heatmap regression.

We then prove that the quantization system is losses as follows. For the decode operation in (17), if there is no heatmap error, we then have

We can reconstruct the fractional part of xig\boldsymbol{x}_{i}^{g}, i.e.,

That is, (xip,yip)=(xig,yig)\left(x_{i}^{p},y_{i}^{p}\right)=\left(x_{i}^{g},y_{i}^{g}\right), i.e., there is no quantization error. ∎

Appendix B Experiments

In this section, we provide additional experimental results on facial landmark detection and human pose estimation.

The influence of different numbers of training samples. The proposed quantization system does not rely on any assumption about the number of training samples, and is lossless for heatmap regression if there is no heatmap error. However, heatmap prediction performance will be influenced by the number of training samples: increasing the number of training samples improves the model generalizability from the learning theory perspective. Therefore, we perform experiments to evaluate the influence of the proposed method when using different numbers of training samples in practice. As shown in Table XI, we find that 1) the proposed method delivers consistent improvements when using different numbers of training samples; and 2) increasing the number of training samples significantly improves the performance of heatmap regression models with low-resolution input images.

The influence of different backbone networks. If we do not take the heatmap prediction error into consideration, the quantization error in heatmap regression is then caused by the downsampling of heamaps: 1) the downsampling of input images and 2) the downsampling of CNN feature maps. Though the analysis of heamap prediction error is out the scope of this paper, we perform some experiments to demonstrate the influence of different feature maps from the backbone networks in practice. Specifically, we perform experiments using the following two settings: 1) upsampling the feature maps from HRNet ; or 2) using the feature maps from U-shape backbone networks, i.e., U-Net . As shown in Table XII, we see that 1) directly upsampling the feature maps achieves comparable performance with the baseline method; 2) U-Net performs better than HRNet-W18 when using a small input resolution (e.g., 64×6464\times 64 pixels), while is significantly worse than HRNet when using a large input resolution (e.g., 256×256256\times 256 pixels); and 3) U-Net contains more parameters and requires much more computations than HRNet when using the same input resolution. It would be interesting to further explore more efficient U-shape networks for low-resolution heatmap-based semantic landmark localization.

The influence of different types of heatmap. We perform some experiments to demonstrate the influence when using different types of heatmap, Gaussian heatmap and binary heatmap. As shown in Table XIII, 1) when using a large input resolution, the heatmap regression model using either Gaussian heatmap or binary heatmap achieves comparable performance; and 2) when using a low input resolution, the heatmap regression model achieves better performance with the binary heatmap. We demonstrate the differences between binary heatmap and Gaussian heatmap in Figure 9. Specifically, the Gaussian heatmap improves the robustness of heatmap prediction, while at the risk of increasing the uncertainty on the maximum activation point in the predicted heatmap. Therefore, when training very efficient heatmap regression models using a low input resolution, we recommend the binary heatmap.

The qualitative comparison between the vanilla quantization system and the proposed quantization system. We provide some demo images for facial landmark detection using both the baseline method (i.e., k=1k=1) and the proposed quantization method (i.e., k=9k=9) to demonstrate the effectiveness of the proposed method for accurate semantic landmark localization. As shown in Fig. 10, we see that the black landmarks are closer to the blue landmarks than the yellow landmarks, especially when using low resolution models (e.g., 64×6464\times 64 pixels).

B.2 Human Pose Estimation

To utilize the proposed method for accurate semantic landmark localization, it contains only one hyper-parameter kk, i.e., the number of activation points. To better understand the effectiveness of the proposed method for human pose estimation, we perform ablation studies using different numbers of alternative activation points on both MPII and COCO datasets. As shown in Table XIV and Table XV, we find that the proposed method achieves comparable performance when kk is between 1010 and 2525, making it easy to choose a proper kk on the validation data for human pose estimation applications.