LISA: Learning Implicit Shape and Appearance of Hands
Enric Corona, Tomas Hodan, Minh Vo, Francesc Moreno-Noguer, Chris Sweeney, Richard Newcombe, Lingni Ma
Introduction
Since the thumb opposition enabled grasping around 2 million years ago , humans interact with the physical world mainly with hands. The problems of modeling and tracking human hands have therefore naturally received a considerable attention in computer vision . Accurate and robust solutions to these problems would unlock a wide range of applications in, e.g., human-robot interaction, prosthetic design, or virtual and augmented reality.
Most research efforts related to modeling and tracking human hands, e.g., , rely on the MANO hand model , which is defined by a polygon mesh that can be controlled by a set of shape and pose parameters. Despite being widely used, the MANO model has a low resolution and does not come with texture coordinates, which makes representing the surface color difficult.
The related field of modeling and tracking human bodies has been relying on parametric meshes as well, with the most popular model being SMPL which suffers from similar limitations as the MANO model. Recent approaches for modeling human bodies, e.g., , rely on articulated models based on implicit representations, such as Neural Radiance Field or Signed Distance Field (SDF) . Such representations are capable of representing both shape and appearance and able to capture finer geometry compared to approaches based on parametric meshes. However, it is yet to be explored how well implicit representations apply to articulated objects such as the human hand and how they generalize to unseen poses.
We explore articulated implicit representations for modeling human hands and make the following contributions:
We introduce LISA, the first neural model of human hands that can capture accurate hand shape and appearance, generalize to arbitrary hand subjects, provide dense surface correspondences (via predicted skinning weights), be reconstructed from images in the wild, and easily animated.
We show how to train LISA by minimizing shape and appearance losses on a large set of multi-view RGB image sequences annotated with coarse 3D poses of the hand skeleton.
The shape, color and pose representations in LISA are disentangled by design, enabling fine control of selected aspects of the model.
Our experimental evaluation shows that LISA surpasses baselines in hand reconstruction from 3D point clouds and hand reconstruction from RGB images.
Related work
Parametric meshes. Thanks to their simplicity and efficiency, parametric meshes gained great popularity for modeling articulated objects such as bodies , hands , faces and animals . The MANO hand model is learned from a large set of carefully registered hand scans and captures shape-dependent and pose-dependent blend shapes for hand personalization. Despite widely adopted in hand tracking and shape estimation , the MANO mesh is limited by a low resolution rooted from solving a large optimization problem with classical techniques. To reconstruct finer hand geometry, graph convolutional networks are explored in and spiral filters in . Based on a professionally designed mesh template, DeepHandMesh learns the pose and shape corrective parameters by a neural network. Chen et. al., refined MANO by developing a UV-based representation. GHUM introduces a generative parametric mesh where the shape corrective parameters, skeleton and blend skinning weights are predicted by a neural network.
Implicit shape representations. Many works adopt neural networks to model geometry by learning an implicit function, which is continuous and differentiable, such as the signed distance field (SDF) or the occupancy field . To improve learning efficiency, studied part-based implicit templates to model mid-level object-agnostic shape features. Implicit representations were extended to articulated deformation, in LoopReg with a weakly-supervised training using cycle consistency by learning inverse skinning, which maps surface points to the SMPL human body model . Based on SMPL, NASA trains one OccNet per skeleton bone to approximate the shape blend shapes and pose blend shapes. PTF extends NASA and registers point clouds to SMPL. In a similar spirit, imGHUM trains four DeepSDF networks whose predictions are fused by an additional lightweight network. To eliminate the need of having the ground-truth SMPL in NASA training, SNARF utilizes an iterative root finding technique to link every query point in the posed space to the corresponding point in the canonical space, which enables differentiable forward skinning. LEAP and SCANimate additionally model both forward and inverse skinning by a neural network, and use cycle consistency to supervise training of transformation to the canonical space. LEAP also extends the framework to multi-subject learning by mapping the bone transformation to a shape feature, and SCANimate builds animatable customized clothed avatars. We take inspiration from NASA to constrain hand deformation, but explicitly model the skinning weights for blending shape and color.
Implicit appearance representations. A number of approaches have been proposed to learn appearance of a scene from multi-view images. The idea is to model the image formation process by rendering a neural volume with ray-casting . Particularly, NeRF gains popularity with an efficient formulation of modeling the radiance field. Follow-up studies show that geometry can be improved if the density is regulated by occupancy or SDF . In this work, we use VolSDF as a backbone renderer. For dynamic scenes, combine NeRF with learning a deformation field. For modeling dynamic human bodies, Neural Body attaches learnable vertex features to SMPL, and diffuses the features with sparse convolution for volumetric rendering. A-NeRF conditions NeRF with SMPL bone transformations to learn an animatable avatar. Similar ideas are proposed in and NARF . H-NeRF combines imGHUM with NeRF to enable appearance learning and train a separate network to predict SDF. In our work, the prediction of appearance and SDF is independent within each bone and later weighted by the corresponding skinning weights.
Disentangled representations. Disentangling parameters of certain properties such as pose, shape or color is desireable as it allows treating (e.g., estimating or animating) these properties independently. Inspired by the parametric mesh models, Zhou et al., trained a mesh auto-encoder to disentangle shape and pose of humans and animals. They developed an unsupervised learning technique based on a cross-consistency loss. DiForm adopted a decoder network to disentangle identity and deformation in learning an SDF-based shape embedding. A-SDF factored out shape embedding and joint angles to model articulated objects. NPM proposed to train shape embedding on canonically posed scans, followed by another network to learn the deformation field with dense supervision. A similar idea to the deformation field was adopted by i3DMM to learn a human head model. The method disentangles identity, hairstyle, and expression and is trained with dense colored SDF supervision. In this work, we propose a generative hand representation with disentangled shape, pose and appearance parameters.
Background
MANO represents the human hand as a function of the pose parameters and shape parameters :
where the hand is defined by a skeleton rig with joints, and the pose parameters represent the axis-angle representation of the relative rotation between bones of the skeleton. is a 10-dimensional vector and are the vertices of a triangular mesh. The mapping is estimated by deforming a canonical hand by a Linear Blend Skinning (LBS) transformation, with weights , where is the number of bones. Concretely, given a vertex on the canonical shape, LBS transforms the vertex as follows:
where is the rigid transformation applied on the rest pose of bone , is the entry of and denotes the homogeneous coordinates of . LISA builds on MANO’s definition of the skeleton by using the same pose parameters and bone transformations.
NeRF/VolSDF. NeRF is a state-of-the-art rendering algorithm for novel view synthesis. The algorithm models the continuous radiance field of a static scene by learning the following function:
LISA: The proposed hand model
This section provides a detailed description of the proposed hand model, which we dub LISA for Learning Implicit Shape and Appearance model.
Problem settings. Consider a dataset of multi-view RGB video sequences with known camera calibration. Each sequence captures a single hand from a random person posing random motion. The objective is to learn a hand model, which reconstructs the hand geometry, the deformation and the appearance, while also generalizes to reconstruct unseen hands and motion from test images. In contrast to prior hand modeling works, which often require a large collection of high-quality 3D hand scans, we consider a setup that lowers the requirement for data collection but adds the challenge to the algorithm. Inspired by classical hand modeling approaches, we assume that a kinematic skeleton is associated with the hand, where the coarse 3D poses are produced by pre-processing the training sequences with the state-of-art hand tracking. The motivation of using a skeleton is to regulate the hand deformation with articulation and to enable animation for the obtained model. To focus the deep network on the hands, we further simplify the input by assuming the foreground masks are known.
Our goal is to learn a mapping function from the parametric skeleton to a full hand model of the shape and appearance. In this work, we choose the skeleton to be parameterized by MANO and formulate the learning as:
In the remainder of the section, we explain how to model Eq. 6 with network training.
Independent per-bone predictions with skinning. Following , we approximate the overall hand shape by a collection of rigid parts, which are in our case defined by bones. Specifically, the network is split into MLPs predicting the signed distance, , and MLPs predicting the color, , with each MLP making an independent prediction with respect to one bone. As the input images correspond to posed hands, the point is first unposed (i.e., transformed to the coordinate space of the hand in the rest pose) using the kinematic transformations of bones, : , where and are the rotation and translation components of the transformation . With this formulation, we collect a set of independent SDF predictions and color predictions for each query point , where:
To combine the per-bone output into a single SDF and color addition, we introduce an additional MLP to learn the weights. The weight MLP takes the input as the concatenation of unposed and the predicted SDF per MLP , to output the weighting vector . A softmax layer is used to constrain the value of to be probability-alike, i.e., and . The final output for a query point is then computed by:
Note the weight vector is an analogy to skinning weights in classic LBS-based models. The similar design has also been explored by NASA and NARF . The difference is that NASA selects one MLP output, which is determined by the maximum of the predicted occupancies. NARF proposes to learn the weights with an MLP, but only uses the canonicalized points to train this module. In our design, the network sees both canonicalized points and the per-bone SDF. The SDF serves as a valuable guide in learning skinning weight. More importantly, the gradients can now back-propagate via the weights to train the per-bone MLPs. This means MLPs can leverage to avoid learning SDF for far-away points. We show in experiments that this design greatly improves geometry.
Model rendering. As in , we first need to obtain the volume densities from the predicted signed distance field before rendering. We infer it indirectly from the predicted signed distances:
where is the signed distance of , is the CDF of the Laplace distribution, and and are two learnable parameters (see for further details).
The color of a specific image pixel is then estimated via the volume rendering integral, by accumulating colors and volume densities along its corresponding camera ray . In particular, the color of the pixel is approximated by a discrete integration between near and far bounds and of a camera ray with origin :
2 Training
As shown in Fig. 2, the parameters of LISA that need to be learned are: (1) the MLPs for predicting signed distance and color for the bones, (2) the MLP that estimates the skinning weights, and (3) the shape and color latent codes to control the generation process. Note that the pose is not learned and assumed given during training. We next explain how we learn these parameters from the multi-view image sequences from the InterHand2.6M dataset .
Disentangling shape, color and pose. LISA is designed to completely disentangle the representations of pose, shape and color. The shape and color parameters are fully learnable latent vectors. Since both are user specific, we assign the same latent code for all images of the same person. In both cases, they are represented as 128-dimensional vectors, initialized from a zero-mean multivariate-Gaussian distribution with a spherical covariance, and optimized during training following the auto-decoder formulation of .
The pose parameters are defined by the 48-dimensional representation of MANO. When training on InterHand2.6M, we kept the provided ground-truth pose parameters fixed for the initial of training steps, then we started optimizing the parameters to account for errors in the ground-truth annotations.
Color calibration. In order to allow for slight differences in the intensity of the training images, we follow Neural Volumes and introduce a per-camera and per-channel gain and bias that is applied to the rendered images at training time. At inference, we use the average of these calibration parameters.
Loss functions. To learn LISA, we minimize a combination of losses that aim to ensure accurate representation of the hand color while properly regularizing the learned geometry. Specifically, we optimize the network by randomly sampling a batch of viewing directions and estimating the corresponding pixel color via volume rendering. Let be the estimated pixel color and the ground truth value. The first loss we consider is:
where denotes the -norm. We also regularize the SDF of with the Eikonal loss to ensure it approximates a signed distance function:
where is a set of points sampled both on the surface and uniformly taken from the entire scene. In order to prevent local minima in regions relying only on one or a few bones, We use the pseudo-ground truth pose and shape parameters to obtain an approximate 3D mesh and its corresponding skinning weights , which we use to supervise the predicted skinning weights :
Finally, we also regularize the latent vectors and :
The full loss is a linear combination of the four previous loss terms (with hyperparameters , , \lambda_{\textrm{\mathbf{w}}} and ):
Learning a prior for human hand SDFs. When minimizing Eq. 17, we face two main challenges. First, since we only supervise on images, the simultaneous optimization of shape and texture parameters may lead to local minima with good renders but wrong geometries. Second, the InterHand2.6M dataset we use for training has a large number of images (130k) but they only correspond to 27 different users, compromising the generalization of the model.
To alleviate these problems, we build a shape prior using the 3DH dataset , which contains 13k 3D posed hand scans of 183 different users. The scans are used to pre-train the geometry MLPs in , which we denote , and which are responsible for predicting the signed distance :
We pre-train with two additional losses. First, assuming to be a point of a 3D scan, we enforce to predict a distance on that point:
We also supervise the gradient of the signed distance with the ground truth normal at :
where is the 3D normal direction at .
With these two losses, jointly with losses , \mathcal{L}_{\textrm{\mathbf{w}}} and the regularization , we learn a prior on which is used to initialize the full optimization of the model in Eq. 17. As we show in the experimental section, this prior allows to significantly boost the performance of LISA.
3 Inference
In the experimental section we apply the learned model to 3D reconstruction from point-clouds and to 3D reconstruction from images. Both of these applications involve an optimization scheme which we describe below.
Reconstruction from point clouds. Let be a point-cloud with 3D points. To fit our trained model to this data, we follow a very similar pipeline as the one used to learn the prior. Specifically, we minimize the following objective function:
Reconstruction from monocular or multiview images. Given an input image , we assume we have a coarse foreground mask and that the 2D locations of hand joints, denoted as , are available. These locations can be detected using, e.g., OpenPose . To fit LISA to this data, we minimize the following objective:
where the first two terms correspond to the color loss of Eq. 13 (expanded to all viewing directions intersecting the pixels of the input image), and the shape and pose regularization loss in Eq. 16. The last term is a joint-based data term that penalizes the 2D distance between the estimated 2D joints and the projected 3D joints computed from the estimated pose parameters :
where is the 3D-to-2D projection. We also use extrinsic camera parameters in case of multi-view reconstruction.
Experiments
In this section, we evaluate LISA on the tasks of hand reconstruction from point clouds and hand reconstruction from RGB images, and demonstrate that it outperforms the state of the art by a considerable margin.
Datasets. We train LISA on a non-released version of the InterHand2.6M dataset , which contains multi-view sequences showing hands of 27 users. In total, there are 5804 multi-view frames and 131k images with the resolution of . Every frame has 22 views on average, two of which were not used for training and left for validation. The dataset also provides a pseudo ground truth of the 3D joints, and we remove background in all images using hand masks obtained by a Mask R-CNN model provided by the authors of the dataset. The geometry prior is learned on the 3DH dataset which contains sequences of 3D scans of 183 users (we use the same training/test split of 150/33 users proposed by the authors). For evaluating hand reconstruction from point clouds, we use the test split of the MANO dataset , which includes 50 3D scans of a single user, and the test set of 3DH, which includes scans of 33 users. For hand reconstruction from images, we use the DeepHandMesh dataset , which is annotated with ground-truth 3D hand scans.
Evaluated hand models. As LISA is the first neural model able to simultaneously represent hand geometry and texture, there are no published methods that would be directly comparable. To define baselines, we have therefore re-implemented several recent methods based on articulated implicit representations from the related field of human body modeling. We adapt NASA and NARF to our setup by changing their geometry representation to signed distance fields, adding a positional encoding to NASA, and duplicating their geometry MLPs to predict also color. We train these methods on the InterHand2.6M dataset with supervision on the skinning weights. We did not manage to extend SNARF , as it relies on an intermediate non-differentiable optimization during the forward pass that impedes calculating the output gradient with respect to the input points, which is necessary for applying the Eikonal loss. We also compare to the original MANO model and to our implementation of VolSDF parameterized by the pose, shape and color vectors, but which does not consider a per-bone reasoning. Besides, we ablate the following versions of the proposed model: the full model when trained with images and the geometric prior (LISA-full), a version trained solely with images (LISA-im), and a version trained only with the geometric prior (LISA-geom).
2 Shape reconstruction from point clouds
Table 1 summarizes the results of hand reconstruction from point clouds from the 3DH and MANO datasets. As the evaluation metrics, we report the vertex-to-vertex (V2V) and vertex-to-surface (V2S) distances (in millimeters). We compute these metrics in both directions, i.e. from the reconstruction to the scan and the other way around. For a fair comparison, all reconstructions from all methods based on implicit representations are obtained with the same Marching Cubes resolution (). Since MANO uses a mesh with only 778 vertices, we subdivide its reconstructed surface into 100k vertices.
The results show that LISA-im consistently outperforms the other methods when only images are used for training. Adding the geometric prior (LISA-full) yields a significant boost in performance. When the model is trained solely with the geometric prior (LISA-geom), it yields even lower errors than when trained using both the geometric prior and images (LISA-full). This is because we segmented out the hand in the training images and LISA-full and LISA-im therefore learned to close the surface right after the wrist. This spurious surface increases the measured error.
Figure 3 visualizes examples of the reconstructions. Clear artifacts can be seen in most implicit models, except of LISA-full and the parametric MANO model.
3 Shape and color reconstruction from images
Table 2 evaluates the hand models on the task of 3D reconstruction from single and multiple views on the DeepHandMesh dataset . Among methods trained on images only, LISA-im is consistently superior in 3D shape reconstruction, and its performance is further boosted when the geometric prior is employed (LISA-full).
Color reconstruction from DeepHandMesh images is evaluated in Table 2 by the PSNR metric calculated on renderings of the hand model from novel views. LISA-im is slightly superior in this metric, with exception of the case when a single image is used for the reconstruction, where the performance of LISA-im is on par with NASA. Color reconstruction from InterHand2.6M images is evaluated in Table 3 by the PSNR, SSIM and LPIPS metrics. In this case, all methods are fairly comparable in terms of rendering quality, which we suspect is likely due to noisy hand masks used in training.
Qualitative results are shown in Figure 4. Additionally, we demonstrate in Figure 5 that LISA can reconstruct hands from images in the wild, even in cases where the hand is partially occluded by an object. We refer the reader to the supplementary material for additional qualitative results.
4 Inference speed
To reconstruct the LISA hand model from one or multiple views, we first optimize the pose parameters for 1k iterations, after which we jointly optimize shape, pose and color parameters for additional 5k iterations. This process takes approximately 5 minutes. After converging, we reconstruct meshes at the resolution of , which takes around 5 seconds, or render novel views in approximately one minute. These measurements were made on images with a single Nvidia Tesla P100 GPU. The inference speed is similar for the NASA and NARF models, which also perform per-bone predictions. VolSDF is 2 faster due to the fact that it only uses a single MLP.
Conclusion
We have introduced LISA, a novel neural representation of textured hands, which we learn by combining volume rendering approaches with a hand geometric prior. The resulting model is the first one to allow full and independent control of pose, shape and color domains. We show the utility of LISA in two challenging problems, hand reconstruction from point-clouds and hand reconstruction from images. In both of these applications we obtain highly accurate 3D shape reconstructions, achieving a sub-millimeter error in point-cloud fitting and surpassing the evaluated baselines by large margins. This level of accuracy is not possible to achieve with low-resolution parametric meshes such as MANO or with models representing a single person such as DeepHandMesh . Future research directions include exploring temporal consistency for tracking applications, eliminating the need of rough 2D/3D pose of the hand skeleton and foreground mask at inference, improving the run-time efficiency, or enhancing the expressiveness in terms of high-frequency textural details while maintaining the generalization capability.
Acknowledgements: This work is supported in part by the Spanish government with the project MoHuCo PID2020-120049RB-I00.