A Skeleton-Driven Neural Occupancy Representation for Articulated Hands

Korrawe Karunratanakul, Adrian Spurr, Zicong Fan, Otmar Hilliges, Siyu Tang

Introduction

Humans grasp and manipulate objects with their hands. Modeling 3D poses and surfaces of human hands is important for numerous applications such as animation, games, augmented and virtual reality. Existing hand representations in the literature can be categorized into two paradigms: skeleton representations and mesh-based representations . Even though 3D skeletons are defined in the Euclidean space and are easy to interface with deep neural networks, they lack surface information and therefore not suitable for reasoning about hand-object interaction. In contrast, mesh-based hand models provide surfaces and can thus explicitly reason about physical interactions, such as hand-object manipulation . However, their pose and shape parametrizations are often hard to directly interpret and more difficult to learn in an end-to-end fashion than 3D keypoints.

In this work, we aim to bridge the gap between 3D keypoints and dense surface models. To this end, we propose Hand ArticuLated Occupancy (HALO), a novel hand representation that is driven by keypoint-based skeleton articulation and provides a high-fidelity, resolution-independent neural implicit surface. Importantly, the proposed representation is fully differentiable and can be trained end to end such that volume-based losses can back-propagate gradients to the 3D keypoints. Specifically, HALO makes two key innovations: (1) a fully differentiable skeleton canonicalization algorithm and (2) a shape-aware neural articulated hand representation that is driven directly by skeleton and naturally affords differentiable reasoning about surface and volumetric occupancy.

Differentiable skeleton canonicalization. The core advantage of the HALO model is that its surface representation is fully differentiable and driven by the 3D joint skeleton. Even though a naive iterative fitting procedure such as inverse kinematics could provide the transformations needed for changing from the canonical pose to the target skeleton, it prevents gradient flow from the surface to the keypoints. In addition, due to ambiguities of the twist angle around the finger bones, the optimization procedure may lead to unnatural surfaces. By leveraging bio-mechanical constraints that ensure plausible poses, we propose a novel differentiable layer that converts 3D keypoint locations to the corresponding bone rotations and translations with respect to a canonicalized pose space in a single forward pass. We call this operation a canonicalization layer. Importantly, this layer does not rely on iterative fitting and can therefore be effectively used in back-propagation-based learning. Furthermore, this layer enables the learning of implicit hand shape representations in the canonical space. This significantly improves the generalization capability of the learned representations across different hand poses and shapes, as shown in Sec. 5.

Skeleton-driven neural articulated hand. Recently, several neural occupancy networks for human body modelling, e.g. , have been proposed. Despite demonstrating the feasibility of parameterizing articulated deformations via neural implicit surfaces, all these methods require ground-truth bone transformations as input, making them unable to interface with keypoints. In addition, it is uncertain and has not been demonstrated how these models will function under noisy predictions without the ground truth transformations. Note that making implicit representations applicable to articulated hands driven by skeleton is non-trivial. The key challenge is how to infer realistic shapes of unseen hands from the skeletons with highly articulated hand poses. Our solution includes two steps. First, by leveraging bio-mechanical constraints, we ensure bijective mappings between 3D skeletons and hand surfaces using the canonicalization layer, which greatly simplifies the learning of the implicit surface for highly articulated hand skeletons; Second, in the canonical space, we learn both identity-dependent and pose-dependent deformations of the hand surface with a set of multi-layer perceptrons that generalize well across different hand shapes and poses. We systematically compare HALO with several baseline methods. The results show that HALO significantly improves the accuracy, generality and visual fidelity over the baselines.

Hand grasps synthesis. To demonstrate the utility of the keypoint-driven implicit surface representation, we deploy HALO for conditional generation of hands grabbing 3D objects. We propose a novel generation pipeline that synthesizes hand keypoints and yields the neural occupancy of the hand via HALO. By exploiting the differentiable nature of HALO, we design a hand-object interpenetration loss to guide the training of the 3D keypoint generator in an end-to-end fashion. Our experiments show that this loss leads to convincing hand-object contact of the generated hands. Furthermore, compared to a MANO-based method (GrabNet ), HALO produces more physically plausible and visually convincing grasps even before refinement, suggesting that the volume-based losses are effective for learning the 3D keypoint generator.

Overall, the contributions include (1) HALO, the first neural implicit hand model that is driven by keypoint-based skeleton articulation, provides smooth neural implicit surfaces, and enables differentiable reasoning about volumetric occupancy; (2) A differentiable skeleton canonicalization layer that maps any skeleton to the canonical pose with a unique bio-mechanically valid transformations; (3) A realistic human grasp generation framework that leverages the efficient volumetric occupancy checks enabled by HALO.

Related Work

Hand pose and mesh estimation. Hand pose estimation is a long standing question and several learning-based approaches have been introduced. These approaches generally involve predicting 3D key point locations , regressing MANO parameters , or directly predicting the full dense surface of the hand . The methods that directly predict 3D key points usually achieve better performance, however, they do not yield dense surface which is crucial for hand interaction. Iteratively fitting a templated-based model such as MANO to the key points could recover the dense surface but also make the process non-differentiable . Alternatively, the dense surface can also be estimated from 3D or 2D keypoints . However, such estimation could result in a change of hand pose from the input keypoints. In contrast, our model produce hand surface that faithfully respects the input pose and allows surface or volumetric losses to back-propagate directly to the keypoints.

Hand representation. The surface of 3D hands can be represented explicitly or implicitly. The most common template-based approaches such as MANO induce a prior of poses and shapes over its learned parameter space for regularization. However, using the learned parameters also increases the learning complexity as these features do not correspond directly to features in the inputs such as visible joints. In , the MANO parameters are predicted directly using additional weak supervision such as hand masks or 2D annotations . Another way of representing explicit 3D hand is to directly store the dense vertex locations of the MANO template . While being more generalizable by avoiding constraints on the parameter space, these approaches require the corresponding, dense 3D annotations, which might be difficult to acquire. Our work differs in that we can recover dense hand surfaces from 3D keypoints, eliminating the need for learning model parameters or predicting dense surface points.

Implicit representation. Several works represent object shapes by learning an implicit function using neural networks , which allows for the modeling of arbitrary object topologies with dynamic resolution. Many approaches for learning such implicit function from various input types were also proposed . These works focus on rigid objects and do not permit shape deformation. Recently, the interest is also on learning an articulated implicit function for human body . NASA represents human bodies using a set of implicit functions, but the model is limited to a specific body shape. LEAP proposes to learn inverse linear blend skinning functions for multiple body shapes, however, it relies on ground truth bone transformation matrices instead of 3D joint locations. To the best of our knowledge, there are no implicit hand representations that can generalize well to various shapes. Grasping Field learns an implicit function for hand and objects together to represent contact but treats every posed hand as a rigid object. As a result, the complexity of learning a wide range of poses increases significantly. In this work, we leverages bio-mechanical constraints of human hand to learn a novel hand model that only takes a skeleton as input and generalizes to different hand shapes and poses.

Hand-object interaction. There has been many studies into hand interacting with object in various settings . Recently, the community has begun exploring the task of generating plausible hand grasps given an object with notable studies including , , and . GanHand generates grasps suitable for each object in a given RGB image by predicting a grasp type from grasp taxonomy and its initial orientation, then optimize for a better contact with the object. GrabNet uses Basis Point Set to represent 3D objects as input to generate MANO parameters. The predicted hand is then fed to a refinement model to improve the contact. Grasping Field learns a signed distance field for both hand surface and object surface in one space, allowing the contact to be learned as regions where distances to both surfaces are zeros. However, the output surface cannot be articulated and requires hand model fitting. Our work differs from others in the way that we use the proposed hand representation to model the contact while keeping the synthesis task as simple as generating 3D keypoints.

HALO: Hand ArticuLated Occupancy

The HALO model is a skeleton-driven neural occupancy function, formally defined as Ow(x∣J)→{0,1}\mathcal{O}_{w}(x|\boldsymbol{J})\to\{0,1\}. Parameterized by neural network weights ww, it maps a 3D point xx to its occupancy value given the hand skeleton represented by a set of 3D keypoint locations J\boldsymbol{J}. In this section, we first describe how to convert an arbitrary 3D joint skeleton to the reference canonical pose in a differentiable and consistent manner, then we introduce our simple yet effective neural occupancy networks for hands.

Our goal is to learn a neural representation of the surface of human hands in a canonical space. Furthermore, we want to deform this shape based on the spatial configuration of the underlying skeleton, represented by 3D keypoints. To do so, we require a mechanism that allows us to convert the 3D keypoints into valid skeletons in the canonical pose in terms of joint angles. As the keypoints have no notion of the surface, naively converting them to axis-angles does not work due to the unconstrained twist of bones. While twist does not affect keypoints, they affect the surface.

We take inspiration from Spurr et al. which defines a consistent local coordinate system for each bone to measure the bone angles for semi-supervised learning. Our objective is to derive a differentiable mapping layer that i) provides means to convert predicted keypoints to the rest pose and back, and ii) ensures that the skeleton is free of implausible twist that would influence the surface.

Building on , we represent each finger bone by two rotation angles, flexion and abduction, relative to its parent bone (Fig. 2). Each bone cannot rotate about itself, thus, no twist. However, such formulation ignores the palm configuration, which is needed for defining the canonical pose. In this work, we propose a method to parameterize the pose of a palm in order to define a consistent canonical pose. We decouple the palmar bone configuration into 1) finger spreading and 2) palm arching. The spreading of fingers is captured via the angles between two adjacent palmar bones. The arching of the palm is defined by the angle between the two planes spanned by three adjacent palmar bones. The resulting palmar region then serves as a frame of reference for the remaining fingers. Please refer to Supp. Mat. Fig. 1 for better visualization.

This set of transformations is unique for each skeleton pose and only allows biomechanically valid transformation. For details, we kindly refer the reader to our Supp.Mat.

2 Neural Occupancy Networks for Hands

Here, we describe how to leverage the unique mapping {B−1}\{\boldsymbol{B}^{-1}\} between the posed skeleton and the canonical skeleton to learn the neural hand representation that generalizes to different shapes with highly articulated poses. We draw inspirations from NASA and explore similar neural network structure due to its simplicity and efficacy.

NASA . NASA learns an implicit representation of a human body Ow(x∣θ)→{0,1}\mathcal{O}_{w}(x|\theta)\to\{0,1\}, conditioned on the pose descriptor θ\theta. Specifically, it defines the implicit surface for each body part separately. Let Bb−1\boldsymbol{B}_{b}^{-1} be the transformation to the canonical pose for bone bb, NASA can be denoted by:

where the pose descriptor θ\theta is defined by a collection of transformation matrices {Bb}b=1B\{\boldsymbol{B}_{b}\}_{b=1}^{B}, and the probability of Ow(x∣θ)\mathcal{O}_{w}(x|\theta) is derived from the maximum occupancy probability across BB child occupancy functions Oˉwb(⋅,⋅)\bar{\mathcal{O}}^{b}_{w}(\cdot,\cdot), where each represents the body part of the bone bb. For a query point xx, each child function Oˉwb(⋅,⋅)\bar{\mathcal{O}}^{b}_{w}(\cdot,\cdot) maps xx to its local coordinate system by the transformation matrix Bb−1\boldsymbol{B}^{-1}_{b}, so that the local shape of each body part can be learned. The term Πwb[⋅]\mathit{\Pi}^{b}_{w}[\cdot] is used to provide global pose information to each child function. Essentially, by querying the occupancy value xx using Bb−1x\boldsymbol{B}^{-1}_{b}x, the NASA model learns a template shape and the correction based on the global pose with Πwb[⋅]\mathit{\Pi}^{b}_{w}[\cdot]. Note that, the bone transformation {Bb}b=1B\{\boldsymbol{B}_{b}\}_{b=1}^{B} is assumed to be given. For more details, we kindly refer the reader to .

where each implicit function Oˉwb\mathcal{\bar{O}}^{b}_{w} learns the corresponding part shape based on the hand pose descriptor Πwb[{B−1t0}]\mathit{\Pi}^{b}_{w}[\{\boldsymbol{B}^{-1}t_{0}\}] and our bone length descriptor fb(D)f^{b}(\boldsymbol{D}).

Shape descriptor variations. We investigate two versions of bone length encoders fb(D)f^{b}(\boldsymbol{D}): the local bone encoder flb(D)=dbf^{b}_{l}(\boldsymbol{D})=d_{b} and the global encoder fgb(D)=[db;MLP(D)]f^{b}_{g}(\boldsymbol{D})=[d_{b};MLP(\boldsymbol{D})], where the MLP for the global encoder has two linear layers. We follow a similar training strategy as , for more training details, please refer to the Supp. Mat.

3 Skeleton-driven Articulated Hand Model

To build a skeleton-driven articulated hand model, we combine the previously described canonicalization layer and the neural hand surface together. Specifically, HALO takes the input 3D keypoints to compute bone transformations {B−1}\{\boldsymbol{B}^{-1}\} for the occupancy networks using the canonicalization layer. As the canonicalization layer is differentiable, the model can be trained end to end and allows volume-based losses from the surface to back-propagate to the keypoints. The overview of HALO is shown in Fig. 3. Note that the bone lengths D\boldsymbol{D} can also be computed from the keypoints. During inference, only 3D keypoints are needed as input to reconstruct hand surface.

Human Grasps Generation

We show the applicability of the HALO model in the challenging task of grasps generation. Given an object, we aim to generate diverse grasps with natural and plausible hand-object interaction. Our grasp generation pipeline consists of two parts: a 3D keypoints generator based on a variational autoencoder (VAE) and the HALO model for obtaining the hand surface.

The advantages of using HALO are two-fold. First, we decouple the complexity of learning the pose, represented by the skeleton, from that of learning the surface that corresponds to the pose; Second, the implicit model enables fast intersection tests between hand and object, which can be used to efficiently compute an interpenetration loss. Combined with the differentiable skeleton canonicalization layer, the interpenetration loss can be used to improve the keypoints generator in both end-to-end training and post-optimization refinement.

Our grasp generation pipeline is similar to and , but with the following key differences. First, in , the output is a rigid implicit surface that cannot be articulated. To obtain an animatable hand for downstream tasks, additional MANO model fitting is required. Second, in , the grasps generator is trained to produce the MANO parameters which is not directly related to the Euclidean space where the hand and the object live in. The challenge of interfacing the MANO parameters with deep neural networks is reflected in the GrabNet (CoarseNet) results which will be discussed in the experiments section.

To train the VAE model, we use the following losses: the KL-divergence loss on the hand latent ZZ, L2 loss on the predicted key points, L1 bone lengths loss, and the bone angle losses. The bone angle losses are used to provide additional supervisions for learning the hand structure which consists of 1) flexion angles θif\theta^{f}_{i} and abduction angles θia\theta^{a}_{i}, 2) angles between adjacent palmar bones θip\theta^{p}_{i} and 3) angles between palmar planes θin\theta^{n}_{i}. The bone angles are the same as used in Sec. 3. The losses are defined as L1 angle difference between the prediction and the ground truth.

Interpenetration loss. In addition to the losses on the keypoints, we also use the interpenetration loss on the hand surface to avoid collision between hand and object. The key idea is to penalize every points inside the object that is also occupied by the hand. Concretely, for a set of points sampled inside the object PoP_{o} and the predicted key points Jˉ\boldsymbol{\bar{J}}, the interpenetration loss is defined as:

where DJˉD_{\boldsymbol{\bar{J}}} is the bone length vector for Jˉ\boldsymbol{\bar{J}} and Θ(Jˉ)\Theta(\boldsymbol{\bar{J}}) maps the predicted key points to the HALO pose vector θ\theta using the differentiable transformation matrices in Eq. 1.

2 Optimization-based Refinement

To demonstrate that the efficient intersection tests enabled by HALO can be used for optimization, we refine the sampled hands by changing the global translation tt to avoid collision with the object. The refinement is run for 10 steps with the interpenetration loss term in Eq. 4. The optimization objective is:

This simple optimization step aims at refining the contact after the initial prediction of HALO-VAE. It is analogous to the RefineNet in , but with an explicit objective to avoid collision instead of being a neural network denoiser.

Experiments

In this section, we assess our skeleton-driven hand model and the grasp synthesis pipeline. First, in Sec. 5.1, we validate the efficacy of HALO as a neural implicit hand model and compare it to the surface baseline and keypoints-to-surface baselines . Second, we show in Sec. 5.2 that HALO can be used effectively in generative tasks which require surface-based reasoning in form of grasp synthesis. For more experiments, please see supplementary materials.

We first evaluate the performance of the proposed implicit surface representation and analyze the effect of the keypoint-to-transformation mapping layer.

Training data. To train our neural occupancy hand model, we utilize MANO hand meshes. Following , for each mesh we sample points with two strategies: 1) uniformly sampling in the hand bounding box, 2) sampling on the surface with additional isotropic Gaussian noise. Only the uniformly sampled points are used for evaluation. The associated occupancy value of each query point is computed by casting a ray from the sampled point and counting the number of intersections along the ray. The ground truth bone transformation matrices are computed along the kinematic chain to transform the template MANO hand into the target pose. The skinning weights are taken from the skinning weights of MANO. We use the Youtube3D (YT3D) hands dataset in all our experiments. The YT3D training set contains 50,175 hand meshes of hundreds of subjects performing a wide variety of tasks in 102 videos. The test set covers 1,525 meshes from 7 videos.

Evaluation metrics. For 3D surface reconstruction evaluation, we compute the mean Intersection over Union (IoU), Chamfer-L1 distance, and normal consistency score .

Here we investigate the generalization ability of the proposed HALO model to represent articulated hands with various poses and shapes. The results are summarized in Tab. 1

Baseline. We use the NASA model as our baseline. The NASA model is designed to represent an implicit function of an articulated body. However, by changing the input dimension and the number of part-models to match the number of hand parts, it can also be used to represent an articulated hand. We trained the baseline model using the bone transformation matrices taken from MANO and the sampled query points. For details on implementation and network architecture we refer to the supplementary.

Surface vertex re-sampling. In , the surface vertices vv used for enforcing the part models in the skinning loss Ls\mathcal{L}_{s} are the mesh vertices of SMPL . Similarly, we use MANO surface vertices during training. However, we notice that the human-designed mesh often has many more vertices in the area around the joints which could cause the part models to bias toward the bone endpoints. Thus, we propose to re-sample the surface vertices uniformly on the mesh surface. This result is performance degradation but the bone connections are more natural with less artifact.

Local and global bone encoders. The bone lengths of a human hand greatly influence the hand shape. Therefore, for the local bone encoder, we add the bone length did_{i} to the back-projected query point as input to the part model [Bi−1x;di][\boldsymbol{B}^{-1}_{i}x;d_{i}]. As shown in Tab. 1, the local bone encoder improves the reconstruction quality both in terms of IoU and Chamfer-L1 distance. We further extend the local bone encoder by considering all the bone lengths as input. A concatenated vector of bone lengths is first fed into a small feed-forward neural network to get the global bone feature fg(D)f_{g}(\boldsymbol{D}), which is then concatenated with the query point xx and the local bone length did_{i} as input to the part model.

Results. By combining the local and global bone encoders, HALO significantly improves the 3D surface reconstruction quality compared to NASA. As shown in Tab. 1, the IoU is increased from 0.8960.896 to 0.9320.932 and the Chamfer-L1 distance is decreased from 1.057mm1.057mm to 0.719mm0.719mm.

We provide a qualitative comparison between NASA, and HALO in Fig. 5, confirming the quantitative results. The proposed HALO representation generalizes well for highly articulated poses, whereas the NASA model produces severe artifacts at the connection between parts.

1.2 3D keypoints to hand surface

Tab. 1 also shows the result from HALO that only takes 3D keypoints as input. The keypoint model achieves comparable surface reconstruction performance as when the ground truth transformation are given, showing the effectiveness of our method. We show the qualitative results in Fig. A Skeleton-Driven Neural Occupancy Representation for Articulated Hands and 5.

In addition, to evaluate the keypoint-to-surface pipeline, we then compare HALO to the equivalent component in and which estimates hand surface from 3D keypoints. The evaluation is done on same the Youtube3D test set where the ground truth 3D keypoints are given as input. As requires both 2D and 3D coordinates as inputs, the 2D keypoints is obtain by projecting the 3D keypoints perpendicular to the palm. For evaluation, we also report the 3D joint error between the predicted hand and the input joints. This metric measures if the input keyoints are faithfully respected by the models. By design, the HALO model does not change the keypoint locations from input to output, thus does not have this error. The comparison in Tab. 2 shows that and change the hand pose and shape in the prediction while HALO faithfully reconstructs the hand surface according to the given keypoints.

2 Grasp Synthesis

To assess the utility of HALO in downstream tasks we demonstrate our grasp generative model, HALO-VAE.

Dataset. We leverage the recently introduced GRAB dataset and compare our results to GrabNet . We compare both to the initial (coarse) predictions of GrabNet and the refined results which matches with our own two-stage generation process. The test set contains 6 unseen objects. For each object, we fix the object orientation and sample 20 hand proposals from each model.

Physics Metrics. Following , we evaluate the physical plausibility (interpenetration volume and contact ratio) and diversity, and provide results from a perceptual study. To evaluate the interpenetration and contact, we measure the ratio of frames in which the hand is in contact with the object and average the interpenetration volume. The volume is calculated by voxelizing hand and object mesh with 1mm cubes and counting the number of intersecting cubes.

User study. We asked 75 participants in a forced-alternative-choice perceptual study to ‘select the grasp that is more realistic’. For each question, the user is shown 4 views per grasp and forced to select one. We compare all possible combinations on the same object. Each question is assigned to at least 2 participants, totaling 4,800 data points per pair of model comparison. To ensure that the grasps from HALO-VAE and GrabNet have the exact same texture, we fit MANO to our generated key points for rendering.

Diversity. Following , we compute the diversity of the sampled grasps by performing k-means with 20 clusters on all samples, then evaluate the entropy of the cluster assignment and the average cluster size. More diversity results in higher value for both metrics. We use the flatten key point locations of the hands after aligning the root joint and the plane spanned by middle and index palmar bone as features.

We first validate the efficacy of the interpenetration loss (Eq. 4). We compare the HALO-VAE models with and without the interpenetration loss. The results show that the interpenetration loss helps in: 1) reducing the collision between the objects and the generated hands (Tab. 4, col.1-2), and 2) largely improves the user preference of the corresponding model (Tab. 4, first row), demonstrating the efficacy of the proposed neural occupancy representation of articulated hand for reasoning about hand-object interaction.

Next, we compare HALO-VAE with GrabNet . Both HALO-VAE and GrabNet-coarse are CVAE based generative models and end-to-end trainable, the key difference is that GrabNet-coarse generates MANO model parameters whereas HALO-VAE generates 3D keypoints. As shown in Tab. 4, HALO-VAE outperforms GrabNet-coarse by a large margin for interpenetration volume and sample diversity. Moreover, the HALO-VAE model without interpenetration also compares favorably to GrabNet-coarse, suggesting that the 3D keypoints based representation is well suited to interface with deep neural networks.

Finally, we compare our optimization-based refinement with GrabNet-refine. To the best of our knowledge, the RefineNet is not trained end-to-end with GrabNet-coarse and used for three steps during the inference. As shown in Tab. 4, our refined grasps attain a higher user score, suggesting they are more realistic and natural compared to the grasps refined by GrabNet-refine.

Discussion and Conclusion

In this work, we introduce HALO, a novel surface representation for articulated hands that can generalize to different hand poses and shapes. We address the issue of the transformation matrix requirement for inferring the 3D occupancy hand by proposing a skeleton canonicalization algorithm that computes valid transformations from 3D keypoints. The experiments show that our proposed hand model outperforms the baseline and can represent a wide range of hand poses and shapes. Finally, we demonstrate the HALO can be used to train an end-to-end grasp generator conditioned on an object and produces hand grasps with natural and realistic interaction. We believe that HALO can be useful in future work attempting to reconstruct the surface of articulated hands directly from images via differentiable rendering and for several downstream tasks that need to perform surface-based computation such as collision detection and response.

Acknowledgement

We sincerely acknowledge Shaofei Wang and Marko Mihajlovic insightful discussions and help with baselines.

References

Appendix A More Experimental Analysis

To demonstrate the applicability of the HALO hand model, we introduce the HALO-VAE model (Sec. 4) for the conditional human grasps generation task. Our model learns to generate the 3D keypoints of a hand that grasps a given object, whereas our baseline model GrabNet generates MANO parameters that represent the grasping hand. As shown in the Sec. 5.2, the proposed HALO-VAE model largely outperforms the GrabNet in terms of physical plausibility and naturalness of the generated human grasps.

Here, we provide more details and analysis on the object encoding schemes. The GrabNet encodes the object using the BPS features . Specifically, the BPS encoder is a 4-layers feed-forward network with residual connections between each layer. In our trials, we have experimented with this BPS encoder instead of our PointNet encoder. We used the same 4096 basis points as provided in the GRAB dataset. The rest of the architecture is the same to HALO-VAE.

However, with the near-zero keypoint reconstruction loss on the validation set, the generated grasps for a given object are always the same, suggesting that the information from the sampled Gaussian is not used. We suspect that during training, the hand keypoints can be inferred using only the object BPS, as the 3D keypoints and the BPS features are highly related. Consequently, the decoder can entirely ignore the features from the hand encoder, producing the same grasp for different samples.

In order to directly compare the keypoint-based and the MANO parameter based grasps generation frameworks, here we use the same object encoding scheme that employs the PointNet architecture . Specifically, we change the last layer of HALO-VAE (Fig. 4) from predicting the hand keypoints (6363 dimensions, including 3×213\times 21 keypoints) to predicting MANO parameters (6161 dimentions, including 3 global translation, 3 global rotation, 10 shape parameters and 45 pose parameters). Both models are trained without the interpenetration loss. The results are shown in Tab A.1. The keypoint-based generative model produces grasps with better contact and interpenetration while also being more diverse than those generated from the MANO parameters based model, demonstrating the efficacy of the proposed HALO-VAE model.

A.2 Comparison with Grasping Field [29] and GrabNet [59] on other datasets

In this section, we show the comparison between the generated grasps from the Grasping Field (GF) model, Grabnet, and HALO-VAE on the ObMan and HO3D test objects. Note that the HALO-VAE and Grasping Field are not directly comparable as the meshes produced by Grasping Field do not guarantee to be a valid human hand and require MANO fitting, while our HALO-VAE produces articulated implicit hand surfaces.

Nevertheless, we show qualitative and quantitative comparisons between the GF meshes after MANO fitting and the HALO hand surfaces in Fig. 1 and Tab A.2, A.2, respectively. Due to the artifacts in the hand-designed objects used in the ObMan dataset that interfere with the interpenetration evaluation, e.g non-watertight meshes, surface with holes, internal structure with wrong winding number and zero-volume meshes, we perform the evaluation using the object convex hull instead. The evaluation is performed by generating 5 grasps per object on 30 randomly chosen test objects from the ObMan dataset. In total, we evaluate 150 generated grasps from each model.

The results in Tab. A.2 and A.3 show that our HALO-VAE model produces comparable physically-plausible human grasps than Grasping Field and GrabNet with more diversity.

Appendix B Differentiable Bio-mechanical Canonicalization Layer

In this section, we elaborate on the method for converting 3D keypoints to bone transformation matrices. We closely follow the formulation in Spurr et al. to construct the local coordinate systems F\boldsymbol{F}. Here we provide a brief summary of the method. For more details on F\boldsymbol{F}, we refer the readers to the supplementary material of .

Recall that we seek to compute the set of matrices B−1{\boldsymbol{B}^{-1}}, which, in details, is obtained by sequentially performing the following operations: 1) normalizing palmar plane angles, 2) normalizing palmar bone angles, 3) constructing local coordinate frames Fi\boldsymbol{F}_{i} for each bone with respect to its parent along the kinematic chain, 4) undoing the rotation in the local frames Fi\boldsymbol{F}_{i}, 5) reverting back to the global coordinate frames. Formally,

In the following, we define the notations needed and describe the methods for constructing the local coordinate system Fi\boldsymbol{F}_{i} and each matrix in B−1\boldsymbol{B}^{-1}.

We define all the notations with respect to the right hand. The same procedure could also be applied to the left hand by flipping the x-axis of all joints without loss of generality. We denote 3D root-aligned joint locations of a posed hand as J\boldsymbol{J} where j0j_{0} is the root joint. A bone is defined as a vector pointing from the parent joint to its child b~i=ji−jp(i)\widetilde{b}_{i}=j_{i}-j_{p(i)} where p(i)p(i) denotes the parent of joint ii in the kinematic tree (see Fig. 1). We define bib_{i} as a normalized bone of b~i\widetilde{b}_{i} and call K\boldsymbol{K} the mapping from joints J\boldsymbol{J} to normalized bones. As a shorthand, we call the palmar bones that are connected to the root joint j0j_{0} as the level-0 bones (bones b1,⋯ ,b5b_{1},\cdots,b_{5}). We call a bone with kk bone segments in between itself and the root joint a k-level bone. The bone level from 0 to 3 are denoted by the color black, blue, dark purple, and orange respectively in Fig. 1(b).

Palmar bone rotations. Given a hand skeleton in global coordinate frame, we denote θip\theta^{p}_{i} to be the angle between a palmar bone ii and its adjacent palmar bone i+1i+1; θin\theta^{n}_{i} to be the plane angle between plane nin_{i} spanned by the palmar bone ii, i+1i+1 and plane ni+1n_{i+1} spanned by the palmar bone i+1i+1, i+2i+2. We denote the properties of the reference canonical hand with \prescriptc(.)\prescript{c}{}{(.)}

B.2 Palmar Bone Normalization (𝑷𝑷\boldsymbol{P})

Given a globally normalized hand keypoints, we first compute the transformation matrices P\boldsymbol{P} the rotate the palmar bones to match the canonical pose. The palmar bone transformation normalization P\boldsymbol{P} is a combination of the palmar plane angle normalization Pp\boldsymbol{P^{p}} and the palmar bone angle normalization Pa\boldsymbol{P^{a}}, with P=PaPp\boldsymbol{P}=\boldsymbol{P^{a}}\boldsymbol{P^{p}}.

Palmar Plane Angle (Pp\boldsymbol{P^{p}}). To change the bone angle, we rotate the outer bone (with middle finger being the center) about the shared bone until the plane angle is equal to the canonical angle \prescriptcθip\prescript{c}{}{\theta^{p}_{i}}, which we set to 0.8, 0.2, 0.2 radian for \prescriptcθ1p\prescript{c}{}{\theta^{p}_{1}}, \prescriptcθ2p\prescript{c}{}{\theta^{p}_{2}}, \prescriptcθ3p\prescript{c}{}{\theta^{p}_{3}}, respectively. The plane between the index and middle finger (n2n_{2}) is fixed as reference. The rotation applied on the ring finger bone b4b_{4} is also propagated to the pinky finger bone b5b_{5}.

Palmar Bone Angle (Pa\boldsymbol{P^{a}}). Secondly, we normalize the spread of the fingers by rotating the bones on the plane two adjacent bones. Concretely, we use the middle finger as reference then rotate b2b_{2} on plane n2n_{2}, b1b_{1} on plane n1n_{1}, b4b_{4} on plane n3n_{3}, and b5b_{5} on plane n4n_{4}. The transformation applied on b2b_{2} and b4b_{4} are also applied to b1b_{1} and b5b_{5} respectively. We set the canonical angle to to 0.4, 0.2, 0.2, 0.2 radian for \prescriptcθ1n\prescript{c}{}{\theta^{n}_{1}}, \prescriptcθ2n\prescript{c}{}{\theta^{n}_{2}}, \prescriptcθ3n\prescript{c}{}{\theta^{n}_{3}}, and \prescriptcθ4n\prescript{c}{}{\theta^{n}_{4}} respectively.

B.3 Local Coordinate System (𝑭𝑭\boldsymbol{F})

Now we define the local coordinate system Fi\boldsymbol{F}_{i} for each bone bib_{i}. Note that since in our formulation level-0 bones are always fixed, the only bones that characterize the hand pose are bones at level 1 to 3. Thus, we only describe the coordinate systems for non-zero-level bones. For each non-zero level bone bib_{i} (i>5i>5), its coordinate frame Fi\boldsymbol{F}_{i} is defined by three normalized vectors xi,yi,zix_{i},y_{i},z_{i}. To construct the coordinate system for level-1 to level-3 bones, the z-axis of Fi\boldsymbol{F}_{i} for non-zero level bones are always defined by the normalized bone vector of its parent zi=norm(bp(i))z_{i}=\text{norm}(b_{p(i)}). We then describe how to define the x-axis. Afterwards, each local coordinate system is defined as yiy_{i} can be obtained by a cross product: yi=norm(zi×xi)y_{i}=\text{norm}(z_{i}\times x_{i}). Note that the coordinate frame Fi\boldsymbol{F}_{i} does not have a position component because we obtain the bone vectors by subtracting the child joint. Thus, all bones are aligned to the origin. The translation components will be added in the final step of our formulation. However, for illustration purposes, we present each coordinate frame Fi\boldsymbol{F}_{i} with the translation in mind for our figures.

Coordinate Systems for Level-1 Bones Formally, we denote nin_{i} as the normal of a plane spanned by two adjacent level-0 bones where

For better illustration, Fig. 1(a) demonstrates the normal vectors nin_{i} defining each plane. Using the normal vectors above, we define the x-axis for level-1 frames Fi\boldsymbol{F}_{i} for i∈{6,7,8,9,10}i\in\{6,7,8,9,10\} as follows:

In other words, for bone b6b_{6} and b10b_{10}, we define the x-axis for their coordinate system F6\boldsymbol{F}_{6} and F10\boldsymbol{F}_{10} by −n1-n_{1} and −n4-n_{4} because they are on the edge of the palm. For bones b8b_{8}, and b9b_{9}, the x-axis for their coordinate systems are defined by the average normal around the bones.

Coordinate Systems for Level-2 and 3 Bones Given the rotation angles Rp(i)\boldsymbol{R}_{p(i)} of the level-1 bones bp(i)(6≤p(i)≤10)b_{p(i)}(6\leq p(i)\leq 10) ) and the corresponding coordinate systems Fp(i)\boldsymbol{F}_{p(i)}, we can construct coordinate frame Fi\boldsymbol{F}_{i} (11≤i≤1511\leq i\leq 15) for the second-level bones by rotating the coordinate frame Fp(i)\boldsymbol{F}_{p(i)} along the kinematic chain using Rp(i)\boldsymbol{R}_{p}(i). Concretely, the new coordinate frame is given by

Similarly, the coordinate systems for level-3 bones can be obtained by rotating the coordinate systems of level 2 bones using the rotation angles on level 2.

Rotations to the Canonical Pose (Rc\boldsymbol{R^{c}}) Given a local coordinate system Fi=(xi,yi,zi)\boldsymbol{F}_{i}=(x_{i},y_{i},z_{i}) where (xi,yi,zi)(x_{i},y_{i},z_{i}) are the axis of the coordinate system, for a bone bib_{i}, we can measure the flexion angle θif\theta^{f}_{i} and the abduction angle θia\theta^{a}_{i} that characterize bone ii with respect to Fi\boldsymbol{F}_{i}. Fig. 2 visualizes the local coordinate system Fi\boldsymbol{F}_{i} and how rotation angles are measured. Then, given a canonical pose bone bicb_{i}^{c}, we compute the rotation matrix Ric\boldsymbol{R}^{c}_{i} to transform bib_{i} to its canonical pose bicb_{i}^{c} based on the angles relative to the canonical pose (θˉif(\bar{\theta}^{f}_{i}, θˉia)\bar{\theta}^{a}_{i}) in Fi\boldsymbol{F}_{i}.

Since the angles are measured in a consistent coordinate frame, to obtain the angle needed to rotate to the canonical pose, we simply offset the angles of bib_{i} by the angle of bicb_{i}^{c} in the canonical pose:

where θicf\theta^{f}_{i_{c}} and θica\theta^{a}_{i_{c}} are the flexion and abduction angle of the canonical pose with respect to Fi\boldsymbol{F}_{i}. In our experiments, we use the canonical pose identical to that of MANO .

Each rotation matrix Ric\boldsymbol{R}^{c}_{i} is defined locally with respect to a coordinate frame Fi\boldsymbol{F}_{i}. The accumulated rotations Gi\boldsymbol{G}_{i} with respect to the root of the hand for a bone bib_{i} is then the product of the rotation matrices along the kinematic chain:

Where Gp(i)\boldsymbol{G}_{p(i)} are the rotations up to the parent of bib_{i}. With the accumulated rotation matrices encoding the joint-angles relative to the palm of the hand, we need to map all matrices to global coordinates by multiplying with Fi−1\boldsymbol{F}_{i}^{-1}. Recall that Fi\boldsymbol{F}_{i} encodes the mapping from the global coordinate system to the local frame. Thus, its inverse brings the angles back to the global coordinate frame. We summarize the accumulation of angles and the mapping to the global coordinate with the matrix Fi′\boldsymbol{F}_{i}^{\prime}:

With all the necessary components in place, we can compute the transformation matrices B−1\boldsymbol{B}^{-1} for all bones using only the 3D keypoints J\boldsymbol{J} in a differentiable manner. In Sec. 5.1.2 we show that the matrices derived using this formulation can be uses to recover hand surfaces that are very close to those attained by the ground truth transformation matrices from MANO.

Appendix C Implementation details

The loss used to train our articulated hand model can be written as:

where λ\lambda determines the weight of the skinning loss. It is set to 0.5 in all experiments. We turn off the skinning loss after the validation IoU reach 80% as we observe it allows smoother transition between hand parts.

Network architecture. For a fair comparison with the baseline , we use the same network architecture with 4 layers of size 40 for each part model in the ablation study. For the final model used in HALO-VAE, the size is increased to 64 as we observed a small surface quality improvement. The LeakyRelu with factor 0.1 is used as the activation function. All layers have a residual connection and a dropout of 0.2. The subspace projection layers Π\mathit{\Pi} map the input of size B×3B\times 3 to a vector of size 8. When used, the global bone length encoder is a 2-layer feed-forward network of size 40 that maps 16 bone lengths to an encoded vector of size 16.

We define the number of parts as B=16B=16, with one part responsible for the palm and three parts for each finger. When using the 20 transformation matrices obtained from our formulation as our pose descriptor, we disregard the transformations of the root bones, and add an identity transformation for the palm, resulting in 16 transformation matrices.

Training data. To train our neural occupancy hand model, we utilize MANO hand meshes. The query points for each mesh are selected using two strategies: 1) uniformly sampling points in the bounding box of the hand mesh where the root joint is at the origin, 2) sampling on the surface with additional isotropic Gaussian noise. For each strategy, we sample 100,000 points. The associated occupancy value of each query point is computed by casting a ray from the sampled point and counting the number of intersections along the ray. For evaluation, following , we use uniformly sampled points. The bone transformation matrices are computed along the kinematic chain to transform the template hand into the target pose. The shape descriptor β\beta is based on bone length, defined as the Euclidean distance between adjacent joints. The skinning weights are taken from the skinning weights for posing the template mesh in MANO. We use the Youtube3D (YT3D) hands dataset in all our experiments. The YT3D training set contains 50,175 hand meshes of hundreds of subjects performing a wide variety of tasks in 102 videos. The test set covers 1,525 meshes from 7 videos.

Training. We used the Adam optimizer with a learning rate 1e−41e-4 and a batch size of 64 in all experiments. For each mesh at each training step, we sample 2048 points from the 200K pre-sampled query points for the occupancy loss and 2,000 out of 6,000 surface points for the skinning loss. When surface point re-sampling is not used, we sampled 200 out of 778 mesh vertices for the skinning loss.

C.2 HALO-VAE

For training, we use GRAB dataset with the default train/test split. The objects are centered at the origin and 600 points are sampled from the surface. The keypoints are obtained by projecting the surface point using MANO.

Our keypoint VAE consists of an object encoder, keypoint encoder, and a decoder. The object encoder is a 4-layers PointNet encoder with a residual connection between each layer. The hand encoder and the decoder are 4-layers MLP networks with residuals connections. The hand encoder that takes hand key points and the object latent code then produces mean and standard variation of a 32-dimension Gaussian distribution. The decoder takes as inputs noise sample from the Gaussian distribution and the object latent vector to predict the hand key point locations. All layers have size 256. From key points to hand mesh, we use the final HALO model with differentiable canonicalization layer that takes 3D key points as inputs.

Appendix D Limitation

HALO relies on biomechanically plausible 3D keypoints. Training HALO with the model end-to-end with angle losses alleviate this problem and results in a more robust surface prediction, as highlighted in HALO-VAE. This indicates that the inductive bias of our model helps encourage a biomechanically plausible 3D hand surface. However, a severely physically implausible hand skeleton could still produce artifacts on the hand surface.

We further analyse the impact of the biomechanical violations on the reconstruction quality of hand surfaces. We uniformly sample noises with the amplitude [−x,+x][-x,+x] and add them to every dimension of every joints of a valid hand. As shown in Fig. 1, with the amplitudes are within 2mm2mm, HALO produces reasonable hand surface. When xx is increased to 5mm5mm, the reconstructed hand surface starts to show visible artifacts. Nevertheless, we reiterate that this problem can be mitigated by encouraging a bio-mechanically valid skeleton output from the estimator or generator, as in HALO-VAE.

Appendix E Qualitative Results

Figure 1 shows the HALO hand surfaces driven by the keypoint-based skeleton articulations.

In addition, to further demonstrate the generalisability of HALO, we show HALO surface driven by ground truth skeletons from the unseen Interhand2.6M dataset in Figure 2.

E.2 Generative Results

Figure 3 shows the grasps randomly sampled from HALO-VAE conditioned on the objects from the test set of the GRAB dataset.