Real-Time Hand Tracking Using a Sum of Anisotropic Gaussians Model

Srinath Sridhar, Helge Rhodin, Hans-Peter Seidel, Antti Oulasvirta, Christian Theobalt

Introduction

Marker-less articulated hand motion tracking has important applications in human–computer interaction (HCI). Tracking all the degrees of freedom (DOF) of the hand for such applications is hard because of frequent self-occlusions, fast motions, limited field-of-view, uniform skin color, and noisy data. In addition, these applications impose constraints on tracking, including the need for real-time performance, high accuracy, robustness, and low latency. Most approaches from the literature thus frequently fail on even moderately fast and complex hand motion.

Previous methods for hand tracking can be broadly classified into either generative methods or discriminative methods . Generative methods usually employ a dedicated model of hand shape and articulation whose pose parameters are optimized to fit image data. While this yields temporally smooth solutions, real-time performance necessitates fast local optimization strategies which may converge to erroneous local pose optima. In contrast, discriminative methods detect hand pose from image features, e.g., by retrieving a plausible hand configuration from a learned space of poses, but the results are usually temporally less stable.

Recently, a promising hybrid method has been proposed that combines generative and discriminative pose estimation for hand tracking from multi-view video and a single depth camera . In their work, generative tracking is based on an implicit Sum of Gaussians (SoG) representation of the hand, and discriminative tracking uses a linear SVM classifier to detect fingertip locations. This approach showed increased tracking robustness compared to prior work but was limited to using isotropic Gaussian primitives to model the hand.

In this paper, we build on this previous method and further develop it to enable fast, more accurate, and robust articulated hand tracking at real-time rates of 2525 fps. We contribute a fundamentally extended generative tracking algorithm based on an augmented implicit shape representation.

The original SoG model is based on the simplifying assumption that all Gaussians in 3D have isotropic covariance, facilitating simpler projection and energy computation. However, in the case of hand tracking this isotropic 3D SoG model reveals several disadvantages. Therefore we introduce a new 3D Sum of Anisotropc Gaussians (SAG) representation (Figure 3) that uses anisotropic 3D Gaussian primitives attached to a kinematic skeleton to approximate the volumetric extent and motion of the hand. This step towards a more general class of 3D functions complicates the projection from 3D to 2D and thus the computation of the pose fitting energy. However, it maintains important smoothness properties and enables a better approximation of the hand shape with less primitives (visualized as ellipsoids in Figure 3). Our approach, in contrast to previous methods , models the full perspective projection of 3D Gaussians. To summarize, the primary contributions of our paper are:

An advancement of that generalizes the SoG-based tracking to one based on a new 3D Sum of Anisotropic Gaussians (SAG) model, thus enabling tracking using fewer primitives.

Utilization of a full perspective projection model for projection of 3D Gaussians to 2D in matrix-vector form.

Analytic derivation of the gradient of our pose fitting energy, which is smooth and differentiable, to enable real-time optimization.

We evaluate the improvements enabled by SAG-based generative tracking over previous work. Our contributions not only lead to more accurate and robust real-time tracking but also allow tracking of objects in addition to the hand.

Previous Work

Following the survey of Erol et al. we review previous work by categorizing them into either model-based tracking methods or single frame pose estimation methods. Model-based tracking methods use a hand model, usually a kinematic skeleton with additional surface modeling, to estimate the parameters that best explain temporal image observations. Single frame methods are more diverse in their algorithmic recipes, they make fewer assumptions about temporal coherence and often use non-parametric models of the hand. Hand poses are inferred by exploiting some form of inverse mapping from image features to a space of hand configurations.

Model-based Tracking: Rehg and Kanade were one of the first to present a kinematic model-based hand tracking method. Lin et al. studied the constraints of hand motion and proposed feasible base states to reduce the search space size. Oikonomidis et al. presented a method based on particle swarm optimization for full DoF hand tracking using a depth sensor and achieved a frame rate of 1515 fps with GPU acceleration. Other model-based methods using global optimization for pose inference fail to perform at real-time frame rates .

Primitive shapes such as spheres and (super-)quadrics have been explored for tracking objects , and, recently, for tracking hands . However, perspective projection of complex shapes is hard to represent analytically and therefore fast optimization is hard. In this work we use anisotropic Gaussian primitives with analytical expression for perspective projection. An overview of perspective projection of spheroids, which are conceptually similar to anisotropic Gaussians, can be found in .

Tracking hands with objects imposes additional constraints on hand motion. Methods proposed by Hamer et al. , and others model these constraints. However, these methods require offline computation and are unsuitable for interaction applications.

Single Frame Pose Estimation: Single frame methods estimate hand pose in each frame of the input sequence without taking temporal information into account. Some methods build an exemplar pose database and formulate pose estimation as a database indexing problem . The retrieval of the whole hand pose was explored by Wang and Popović . However, the hand pose space is large and it is difficult to sample it with sufficient granularity for jitter-free pose estimation. Sridhar et al. proposed a part-based pose retrieval method to reduce the search space. Decision and regression forests have been successfully used in full body pose estimation to learn human pose from a large synthetic dataset . This approach has been recently adopted for hands . These methods generally lack temporal stability and recover only joint positions or part labels instead of a full kinematic skeleton.

Hybrid Tracking: Hybrid frameworks that combine the advantages of model-based tracking and single frame pose estimation can be found in full body pose estimation and early hand tracking . A hybrid method that uses color and depth data for hand tracking was proposed . However, this method is limited to studio conditions and uses isotropic Gaussian primitives. In this paper, we extend their method by introducing an improved (model-based) generative tracker. This new tracker alone leads to higher tracking accuracy and robustness than the baseline method it extends.

Tracking Overview

Figure 2 shows an overview of our tracking framework. The goal is to estimate hand pose robustly by maximizing the similarity between the hand model and the input images. The tracker developed in this paper extends the generative pose estimation method of the tracking algorithm in .

The input to our method are a set of RGB images from 55 video cameras (Point Grey Flea 3) of resolution 320×240320\times 240 (see Figure 2). The cameras are calibrated and run at 6060 fps. The hand is modeled as a full kinematic skeleton with 2626 degrees-of-freedom (DOF), and unlike other methods that deliver only joint locations or part labels in the images, our approach computes full kinematic joint angles Θ∗={θj∗}\Theta^{*}=\{\theta_{j}^{*}\}.

Our method maximizes the similarity between the 3D hand model projected into all RGB images, and the images themselves, by means of a fast iterative local optimization algorithm. The main novelty and contribution of this work is the new shape representation and pose optimization framework used in the tracker (Section 4). Our pose fitting energy is smooth, differentiable and allows fast gradient-based optimization. This enables real-time hand tracking with higher accuracy and robustness while using fewer Gaussian primitives.

SAG-based Generative Tracking

The generative tracker from is based on a representation called a Sum of Gaussians (SoG) model, that was originally proposed in for full body tracking. The basic concept of the SoG model is to approximate the 3D volumetric extent of the hand by isotropic Gaussians attached to the bones of the skeleton, with a color associated to each Gaussian (Figure 3 (c)). Similarly, input images are segmented to regions of coherent color, and each region is approximated by a 2D SoG (Figure 4 (b-c)). A SoG-based pose fitting energy was defined by measuring the overlap (in terms of spatial support and color) between the projected 3D Gaussians and all the image Gaussians. This energy is maximized with respect to the degrees of freedom to find the correct pose.

Unfortunately, a faithful approximation of the hand volume with a collection of isotropic 3D Gaussians often requires many Gaussian primitives with small standard deviation, a problem akin to packing a volume with spherical primitives. With SoG, this leads to sub-optimal hand shape approximation and increased computational complexity due to a high number of primitives in 3D (Figure 3). In this paper, we extend the SoG model and represent the hand shape in 3D with anisotropic Gaussians, yielding a Sum of Anisotropic Gaussians model (see Figure 1). This not only enables a better approximation of the hand shape with less 3D primitives (Figure 3), but also leads to higher pose estimation accuracy and robustness. The move to anisotropic 3D Gaussians complicates their projection into 2D where scaled orthographic projection cannot be used. But we show that the numerical benefits of the SoG representation hold equally for the SAG model: 1) We derive a pose fitting energy that is smooth and analytically differentiable for the SAG model under perspective projection that allows efficient optimization with a gradient-based iterative solver; 2) We show that occlusions can be efficiently approximated with the SAG model within our energy formulation. This is in contrast to many other generative trackers where occlusion handling leads to discontinuous pose fitting energies.

We represent both the volume of the hand in 3D, as well as the RGB images with a collection of anisotropic Gaussian functions. A Sum of Anisotropic Gaussians (SAG) model thus takes the form:

where Gi(.)\mathcal{G}_{i}(.) denotes a un-normalized, anisotropic Gaussian

with mean μi\bm{\mu}_{i} and covariance matrix is Σi\bm{\Sigma}_{i} for the ithi^{th} Gaussian. Each Gaussian also has an associated average color vector ci\mathbf{c}_{i} in HSV color space.

The Gaussians in 3D have infinite spatial support, which is an advantageous property for pose fitting, as explained later, but also means that the SAG does not represent a finite volume (C(x)>0\mathcal{C}(\mathbf{x})>0 everywhere). We therefore assume that the hand is well modeled by a 3D SAG if the surface passes through each Gaussian at a distance of 11 standard deviation from the mean.

Hand Model Initialization: The hand model for tracking requires initialization of the skeleton dimensions, Gaussian covariances that control their shapes, and associated colors for an actor before it can be used for tracking. Our method accepts manually created hand models which could be obtained from a laser scan. Alternatively, we also provide a fully automatic procedure to obtain a hand model to fit each person. This method uses a greedy optimization strategy to optimize for a total of 33 global hand shape parameters and 33 independent scaling parameters (along the local principal axes) for each of the 1717 Gaussians in the hand model. We observed that starting with a manual model and then using the greedy fitting algorithm works best.

2D RGB Image Modeling: We approximate the input RGB images using 2D SoG, CI\mathcal{C}_{I}, by quad-tree clustering of regions of similar color. While it would also be possible to approximate the image as 2D SAG, the computational expense of the non-uniform region segmentation would prohibit realtime performance. We found in our experiments that around 500500 2D image Gaussians were generated for each camera image.

2 Projection of 3D SAG to 2D SAG

Pose optimization (Section 4.3) requires the comparison of the projections of the 3D SAG into all camera views, with the 2D SoG of each RGB image. Intuitively (for a moment ignoring infinite support), SAG and SoG can be visualized as ellipsoids and spheres, respectively.

The perspective projections of spheres and ellipsoids both yield ellipses in 2D . For the case of isotropic Gaussians in 3D, like in the earlier SoG model, projection of a 3D Gaussian can be approximated as a 2D Gaussian with a standard deviation that is a scaled orthographic projection of the 3D standard deviation . This simple approximation does not hold for our anisotropic Gaussians.

∣M∣|\mathbf{M}| is the determinant of M\mathbf{M}, Aij\mathbf{A}_{ij} is a matrix A\mathbf{A} with its ithi^{th} row and jthj^{th} column removed, and kijk_{ij} is the element at the ithi^{th} row and jthj^{th} column of K\mathbf{K}. Please see the supplementary material for the derivation of M\mathbf{M} and the projection with arbitrary camera matrices. This more general projection also leads to a more involved pose fitting energy with more involved derivatives than for the SoG model, as explained in the next section.

3 Pose Fitting Energy

We now define an energy that measures the quality of overlap between the projected 3D SAG CP=Π(CH)\mathcal{C}_{P}=\Pi(\mathcal{C}_{H}), and the image SoG CI\mathcal{C}_{I}, and that is optimized with respect to the pose parameters Θ\Theta of the hand model. Our overlap measure is an extension of the SoG overlap measure to a SAG. Intuitively, we assume that two Gaussians in 2D match well if their spatial support aligns, and their color matches. This criterion can be expressed by the spatial integral over their product, weighted by a color similarity term. The similarity of any two sets, Ca\mathcal{C}_{a} and Cb\mathcal{C}_{b}, of SAG or SoG in 2D (including combined models with both isotropic and anisotropic Gaussians) can thus be defined as

where EpqE_{pq} is the integral overlap measure mentioned earlier, d(cp,cq)d(\mathbf{c}_{p},\mathbf{c}_{q}) measures color similarity using the Wendland function , and Epq=d(cp,cq) DpqE_{pq}=d(\mathbf{c}_{p},\mathbf{c}_{q})\,D_{pq}. Unlike the SoG model, for the general case of potentially anisotropic Gaussians, the term DpqD_{pq} evaluates to

Using this Gaussian similarity formulation allows us to compute the similarity between the image SoG CI\mathcal{C}_{I} and the projected hand SAG CP\mathcal{C}_{P}.

We also need to consider occlusions of Gaussians from a camera view. Computing a function that indicates occlusion analytically independent of pose parameters is generally difficult and may lead to a discontinuous similarity function. Thus, we use a heuristic approximation of occlusion that yields a continuous fitting energy defined as follows

where wphw_{p}^{h} is a weighting factor for each projected 3D Gaussian of the hand model. EqqE_{qq} is the overlap of an image Gaussian with itself. With this formulation, an image Gaussian cannot contribute more to the overlap similarity than by its own footprint in the image. To find the hand pose, EsimE_{sim} is optimized with respect to Θ\Theta, as described in the following section. Note that the infinite support of the Gaussians is advantageous as it leads to an attracting force between the projected model and the image of the hand, even if they do not overlap in a camera view.

4 Pose Optimization

The final energy that we maximize to find the hand pose takes the form

where Elim(Θ)E_{lim}(\Theta) penalizes motions outside of parameter limits quadratically, and weight wlw_{l} is empirically set to 0.10.1. With the SoG formulation, it was possible to express the energy function (with a scaled orthographic projection) in a closed form analytic expression, and to derive the analytic gradient. We have found that Esim(Θ)E_{sim}(\Theta) in our SAG-based, even with its full perspective projection model, can still be written in closed form with analytic gradient.

We derive the analytical gradient of Esim{E_{sim}} with respect to the degrees of freedom Θ\Theta in three steps. For each Gaussian pair (h,q)(h,q) and parameter θj\theta_{j} we compute

We exemplify the computation at hand of step a); the input to a) is the change of the ellipsoid covariance matrix ∂Σh−1\partial\bm{\Sigma}_{h}^{-1} and the change of position ∂μh\partial\bm{\mu}_{h} with respect to the DOF θj\theta_{j}. In this step we are interested in the total derivative

Following matrix calculus, the partial derivatives of the cone matrix M\bm{M} with respect to μh,Σh\bm{\mu}_{h},\bm{\Sigma}_{h} are

with eie^{i} the ithi^{th} unit vector, μhi{\bm{\mu}_{h}}_{i} the ithi^{th} entry of μh\bm{\mu}_{h}, H=Skμhμh⊤Σh−⊤\bm{H}=\mathbf{S}^{k}\bm{\mu}_{h}\bm{\mu}_{h}^{\top}\bm{\Sigma}_{h}^{-\top}, where kk indexes the unique elements of the symmetric matrix Σh−1\bm{\Sigma}_{h}^{-1}, and Sk\mathbf{S}^{k} is the symmetric structure matrix with the kthk^{th} elements equal to one, and zero otherwise. Steps b) and c) are derived in a similar manner and can be found in the supplementary document. The total similarity energy Esim{E_{sim}} is the weighted sum over all θj\theta_{j} and DpqD_{pq} according to equation 12. Combined with the analytical gradient of Elim(Θ)E_{lim}(\Theta) we obtain an analytic formulation for ∂∂ΘE(Θ)\frac{\partial}{\partial\Theta}\mathcal{E}(\Theta). As sums of independent terms, both E\mathcal{E} and ∂∂ΘE\frac{\partial}{\partial\Theta}\mathcal{E} lend themselves to parallel implementation.

Even though evaluation of fitting energy and gradient is much more involved than for the SoG model, both share the same smoothness properties and can be evaluated efficiently, and thus an optimal pose estimate can be computed effectively using a standard gradient-based optimizer. The optimizer is initiliazed with an extrapolation of the pose parameters from the two previous time steps. The SAG framework leads to much better accuracy and robustness and requires far fewer shape primitives to be compared, as validated in Section 5.

Experiments

We conducted extensive experiments to show that our SAG-based tracker outperforms the SoG based method it extends. We also compare with another state-of-the-art method that uses a single depth camera . We ran our experiments on the publicly available Dexter 1 dataset which has ground truth annotations. This dataset contains challenging, slow and fast motions. We processed all 77 sequences in the dataset and, while evaluated their algorithm only on the slow motions, we evaluated our method on fast motions as well.

For all results we used 1010 gradient ascent iterations. Our method runs at a framerate of 2525 fps on an Intel Xeon E5-1620 running at 3.60 GHz with 16 GB RAM. Our implementation of the SoG-based tracker of runs slightly faster at 4040 fps.

Accuracy: Figure 6 shows a plot of the average error for each sequence in our dataset. Over all sequences, SAG had an error of 24.124.1 mm, SoG had an error of 31.831.8 mm, and had an error of 42.442.4 mm (only 3 sequences). The mean standard deviations were 11.211.2 mm for SAG, 13.913.9 mm for SoG, and 8.98.9 mm for (3 sequences only). Our errors are higher than those reported by because we performed our experiments on both the slow and fast motions as opposed to slow motions only. Additionally, we discarded the palm center used by since this is not clearly defined. We would like to note that perform their tracking on the depth data in Dexter 1 , and use no temporal information. In summary, SAG achieves the lowest error and is 7.77.7 mm better than SoG. This improvement is nearly the width of a finger thus making it a significant gain in accuracy.

Error Frequency: Table 1 shows an alternative view of the accuracy and robustness improvement of SAG. We calculated the percentage of frames of each sequence in which the tracking error is less than xx mm where x∈{15,20,25,30,45}x\in\{15,20,25,30,45\}. This experiment shows clearly that SAG outperforms SoG in almost all sequences and error bounds. In particular the improvement in accuracy is measured by the increased number of frames with error smaller than 1515 mm, and the robustness to fast motions by the smaller number of dramatic failures of errors larger than 3030 mm. For example, in the adbadd sequence 70.7%70.7\% of frames are better than 1515 mm for SAG while only 34.5%34.5\% of frames for SoG. Note that when x=100x=100 mm, the percentage of frames <x<x mm is 100%100\% for SAG.

Effect of Number of Cameras: To evaluate the scalability of our method to the number of cameras we conducted an experiment where each camera was progressively disabled with total active cameras ranging from 22 to 55. This leads to 2626 possible camera combinations for each sequence and a total of 156156 runs for both the SAG and SoG methods. We excluded the random sequence as it was too challenging for tracking with 3 or less cameras.

Figure 5 shows the average error over all runs for varying cameras. Clearly, SAG produces lower errors and standard deviations for all camera combinations. We also observe a diverging trend and hypothesize that as the number of cameras is increased the gap between SAG and SoG will also increase. This may be important for applications requiring very precise tracking such as motion capture for movies. We associate the improvements in accuracy of SAG with its ability to approximate the users’ hand better than SoG. Figure 3 (b, d) visualizes the projected model density and reveals a better approximation for SAG.

Qualitative Tracking Results: Finally, we show several qualitative results of tracking in Figure 7 comparing SAG and SoG. Since our tracking approach is flexible we are also able to track additional simple objects such as a plate using only a few primitives. These additional results, qualitative comparison to , and failure cases can be found in the supplementary video.

Discussion and Future Work

As demonstrated in the above experiments our method advances state of the art methods in accuracy and is suitable for real-time applications. However, the generative method can lose tracking because of fast hand motions. Like other hybrid methods, we could augment our method with a discriminative tracking strategy. The generality of our method allows easy integration into such a hybrid framework.

In terms of utility, we require the user to wear a black sleeve and we use multiple calibrated cameras. These limitations could be overcome if we would only rely the depth data for our tracking. Since the SAG representation is data agnostic, we could model the depth image as a SAG as well. We intend to explore these improvements in the future.

Conclusion

We presented a method for articulated hand tracking that uses a novel Sum of Anisotropic Gaussians (SAG) representation to track hand motion. Our SAG formulation uses a full perspective projection model and uses only a few Gaussians to model the hand. Because of our smooth and differentiable pose fitting energy, we are able to perform fast gradient-based pose optimization to achieve real-time frame rates. Our approach produces more robust and accurate tracking than previous methods while featuring advantageous numerical properties and comparable runtime. We demonstrated our accuracy and robustness on standard datasets by comparing with relevant work from literature.

Acknowledgments: This work was supported by the ERC Starting Grant CapReal. We would like to thank James Tompkin and Danhang Tang.

References