Learning joint reconstruction of hands and manipulated objects

Yana Hasson, Gül Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J. Black, Ivan Laptev, Cordelia Schmid

Introduction

Accurate estimation of human hands, as well as their interactions with the physical world, is vital to better understand human actions and interactions. In particular, recovering the 3D shape of a hand is key to many applications including virtual and augmented reality, human-computer interaction, action recognition and imitation-based learning of robotic skills.

Hand analysis in images and videos has a long history in computer vision. Early work focused on hand estimation and tracking using articulated models or statistical shape models . The advent of RGB-D sensors brought remarkable progress to hand pose estimation from depth images . While depth sensors provide strong cues, their applicability is limited by the energy consumption and environmental constrains such as distance to the target and exposure to sunlight. Recent work obtains promising results for 2D and 3D hand pose estimation from monocular RGB images using convolutional neural networks . Most of this work, however, targets sparse keypoint estimation which is not sufficient for reasoning about hand-object contact. Full 3D hand meshes are sometimes estimated from images by fitting a hand mesh to detected joints or by tracking given a good initialization . Recently, the 3D shape or surface of a hand using an end-to-end learnable model has been addressed with depth input .

Interactions impose constraints on relative configurations of hands and objects. For example, stable object grasps require contacts between hand and object surfaces, while solid objects prohibit penetration. In this work we exploit constraints imposed by object manipulations to reconstruct hands and objects as well as to model their interactions. We build on a parametric hand model, MANO , derived from 3D scans of human hands, that provides anthropomorphically valid hand meshes. We then propose a differentiable MANO network layer enabling end-to-end learning of hand shape estimation. Equipped with the differentiable shape-based hand model, we next design a network architecture for joint estimation of hand shapes, object shapes and their relative scale and translation. We also propose a novel contact loss that penalizes penetrations and encourages contact between hands and manipulated objects. An overview of our method is illustrated in Figure 2.

Real images with ground truth shape for interacting hands and objects are difficult to obtain in practice. Existing datasets with hand-object interactions are either too small for training deep neural networks or provide only partial 3D hand or object annotations . The recent dataset by Garcia-Hernando et al. provides 3D hand joints and meshes of 44 objects during hand-object interactions.

Synthetic datasets are an attractive alternative given their scale and readily-available ground truth. Datasets with synthesized hands have been recently introduced but they do not contain hand-object interactions. We generate a new large-scale synthetic dataset with objects manipulated by hands: ObMan (Object Manipulation). We achieve diversity by automatically generating hand grasp poses for 2.7K2.7K everyday object models from 88 object categories. We adapt MANO to an automatic grasp generation tool based on the GraspIt software . ObMan is sufficiently large and diverse to support training and ablation studies of our deep models, and sufficiently realistic to generalize to real images. See Figure 1 for reconstructions obtained for real images when training our model on ObMan.

In summary we make the following contributions. First, we design the first end-to-end learnable model for joint 3D reconstruction of hands and objects from RGB data. Second, we propose a novel contact loss penalizing penetrations and encouraging contact between hands and objects. Third, we create a new large-scale synthetic dataset, ObMan, with hand-object manipulations. The ObMan dataset and our pre-trained models and code are publicly availablehttp://www.di.ens.fr/willow/research/obman/.

Related work

In the following, we review methods that address hand and object reconstructions in isolation. We then present related works that jointly reconstruct hand-object interactions.

Hand pose estimation. Hand pose estimation has attracted a lot of research interest since the 90s90s . The availability of commodity RGB-D sensors led to significant progress in estimating 3D hand pose given depth or RGB-D input . Recently, the community has shifted its focus to RGB-based methods . To overcome the lack of 3D annotated data, many methods employed synthetic training images . Similar to these approaches, we make use of synthetic renderings, but we additionally integrate object interactions.

3D hand pose estimation has often been treated as predicting 3D positions of sparse joints . Unlike methods that predict only skeletons, our focus is to output a dense hand mesh to be able to infer interactions with objects. Very recently, Panteleris et al. and Malik et al. produce full hand meshes. However, achieves this as a post-processing step by fitting to 2D predictions. Our hand estimation component is most similar to . In contrast to , our method takes not depth but RGB images as input, which is more challenging and more general.

Regarding hand pose estimation in the presence of objects, Mueller et al. grasp 7 objects in a merged reality environment to render synthetic hand pose datasets. However, objects only serve the role of occluders, and the approach is difficult to scale to more object instances.

Object reconstruction. How to represent 3D objects in a CNN framework is an active research area. Voxels , point clouds , and mesh surfaces have been explored. We employ the latter since meshes allow better modeling of the interaction with the hand. AtlasNet inputs vertex coordinates concatenated with image features and outputs a deformed mesh. More recently, Pixel2Mesh explores regularizations to improve the perceptual quality of predicted meshes. Previous works mostly focus on producing accurate shape and they output the object in a normalized coordinate frame in a category-specific canonical pose. We employ a view-centered variant of to handle generic object categories, without any category-specific knowledge. Unlike existing methods that typically input simple renderings of CAD models, such as ShapeNet , we work with complex images in the presence of hand occlusions. In-hand scanning , while performed in the context of manipulation, focuses on object reconstruction and requires RGB-D video inputs.

Hand-object reconstruction. Joint reconstruction of hands and objects has been studied with multi-view RGB and RGB-D input with either optimization or classification approaches. These works use rigid objects, except for a few that use articulated or deformable objects . Focusing on contact points, most works employ proximity metrics , while directly regresses them from images, and uses contact measurements on instrumented objects. integrates physical constraints for penetration and contact, attracting fingers onto the object uni-directionally. On the contrary, symmetrically attracts the fingertips and the object surface. The last two approaches evaluate all possible configurations of contact points and select the one that provides the most stable grasp or best matches visual evidence . Most related to our work, given an RGB image, Romero et al. query a large synthetic dataset of rendered hands interacting with objects to retrieve configurations that match the visual evidence. Their method’s accuracy, however, is limited by the variety of configurations contained in the database. In parallel work to ours jointly estimates hand skeletons and 6DOF for objects.

Our work differs from previous hand-object reconstruction methods mainly by incorporating an end-to-end learnable CNN architecture that benefits from a differentiable hand model and differentiable physical constraints on penetration and contact.

Hand-object reconstruction

As illustrated in Figure 2, we design a neural network architecture that reconstructs the hand-object configuration in a single forward pass from a rough image crop of a left hand holding an object. Our network architecture is split into two branches. The first branch reconstructs the object shape in a normalized coordinate space. The second branch predicts the hand mesh as well as the information necessary to transfer the object to the hand-relative coordinate system. Each branch has a ResNet18 encoder pre-trained on ImageNet . At test time, our model can process 20fps on a Titan X GPU. In the following, we detail the three components of our method: hand mesh estimation in Section 3.1, object mesh estimation in Section 3.2, and the contact between the two meshes in Section 3.3.

Following the methods that integrate the SMPL parametric body model as a network layer , we integrate the MANO hand model as a differentiable layer. MANO is a statistical model that maps pose (θ\theta) and shape (β\beta) parameters to a mesh. While the pose parameters capture the angles between hand joints, the shape parameters control the person-specific deformations of the hand; see for more details.

Hand pose lives in a low-dimensional subspace . Instead of predicting the full 4545-dimensional pose space, we predict 3030 pose PCA components. We found that performance saturates at 3030 PCA components and keep this value for all our experiments (see Appendix A.2).

Supervision on vertex and joint positions (LVHand,LJ\mathcal{L}_{V_{\mathit{Hand}}},\mathcal{L}_{J}). The hand encoder produces an encoding ΦHand\Phi_{\mathit{Hand}} from an image. Given ΦHand\Phi_{\mathit{Hand}}, a fully connected network regresses θ\theta and β\beta. We integrate the mesh generation as a differentiable network layer that takes θ\theta and β\beta as inputs and outputs the hand vertices VHandV_{\mathit{Hand}} and 1616 hand joints. In addition to MANO joints, we select 55 vertices on the mesh as fingertips to obtain 2121 hand keypoints JJ. We define the supervision on the vertex positions (LVHand\mathcal{L}_{V_{\mathit{Hand}}}) and joint positions (LJ\mathcal{L}_{J}) to enable training on datasets where a ground truth hand surface is not available. Both losses are defined as the L2 distance to the ground truth. We use root-relative 3D positions as supervision for LVHand\mathcal{L}_{V_{\mathit{Hand}}} and LJ\mathcal{L}_{J}. Unless otherwise specified, we use the wrist defined by MANO as the root joint.

Regularization on hand shape (Lβ\mathcal{L}_{\beta}). Sparse supervision can cause extreme mesh deformations when the hand shape is unconstrained. We therefore use a regularizer, Lβ = ∥β∥2\mathcal{L}_{\beta}~{}=~{}\|\beta\|^{2}, on the hand shape to constrain it to be close to the average shape in the MANO training set, which corresponds to β=0⃗∈\mathdsR10\beta=\vec{0}\in\mathds{R}^{10}.

The resulting hand reconstruction loss LHand\mathcal{L}_{\mathit{Hand}} is the summation of all LVHand\mathcal{L}_{V_{\mathit{Hand}}}, LJ\mathcal{L}_{J} and Lβ\mathcal{L}_{\beta} terms:

Our experiments indicate benefits for all three terms (see Appendix A.1). Our hand branch also matches state-of-the-art performance on a standard benchmark for 3D hand pose estimation (see Appendix A.3).

2 Object mesh estimation

Following recent methods , we focus on genus 0 topologies. We use AtlasNet as the object prediction component of our neural network architecture. AtlasNet takes as input the concatenation of point coordinates sampled either on a set of square patches or on a sphere, and image features ΦObj\Phi_{\mathit{Obj}}. It uses a fully connected network to output new coordinates on the surface of the reconstructed object. AtlasNet explores two sampling strategies: sampling points from a sphere and sampling points from a set of squares. Preliminary experiments showed better generalization to unseen classes when input points were sampled on a sphere. In all our experiments we deform an icosphere of subdivision level 3 which has 642 vertices. AtlasNet was initially designed to reconstruct meshes in a canonical view. In our model, meshes are reconstructed in view-centered coordinates. We experimentally verified that AtlasNet can accurately reconstruct meshes in this setting (see Appendix B.1). Following AtlasNet, the supervision for object vertices is defined by the symmetric Chamfer loss between the predicted vertices and points randomly sampled on the ground truth external surface of the object.

Regularization on object shape (LE,LL\mathcal{L}_{E},\mathcal{L}_{L}). In order to reason about the inside and outside of the object, it is important to predict meshes with well-defined surfaces and good quality triangulations. However AtlasNet does not explicitly enforce constraints on mesh quality. We find that when learning to model a limited number of object shapes, the triangulation quality is preserved. However, when training on the larger variety of objects of ObMan, we find additional regularization on the object meshes beneficial. Following we employ two losses that penalize irregular meshes. We penalize edges with lengths different from the average edge length with an edge-regularization loss, LE\mathcal{L}_{E}. We further introduce a curvature-regularizing loss, LL\mathcal{L}_{L}, based on , which encourages the curvature of the predicted mesh to be similar to the curvature of a sphere (see details in Appendix B.2. We balance the weights of LE\mathcal{L}_{E} and LL\mathcal{L}_{L} by weights μE\mu_{E} and μL\mu_{L} respectively, which we empirically set to 2 and 0.1. These two losses together improve the quality of the predicted meshes, as we show in Figure A.4 of the appendix. Additionally, when training on the ObMan dataset, we first train the network to predict normalized objects, and then freeze the object encoder and the AtlasNet decoder while training the hand-relative part of the network. When training the objects in normalized coordinates, noted with nn, the total object loss is:

Hand-relative coordinate system (LS,LT\mathcal{L}_{S},\mathcal{L}_{T}). Following AtlasNet , we first predict the object in a normalized scale by offsetting and scaling the ground truth vertices so that the object is inscribed in a sphere of fixed radius. However, as we focus on hand-object interactions, we need to estimate the object position and scale relative to the hand. We therefore predict translation and scale in two branches, which output the three offset coordinates for the translation (i.e., x,y,zx,y,z) and a scalar for the object scale. We define LT=∥T−T^∥22\mathcal{L}_{T}=\|T-\hat{T}\|^{2}_{2} and LS=∥S−S^∥22\mathcal{L}_{S}=\|S-\hat{S}\|^{2}_{2}, where T^\hat{T} and S^\hat{S} are the predicted translation and scale. TT is the ground truth object centroid in hand-relative coordinates and SS is the ground truth maximum radius of the centroid-centered object.

Supervision on object vertex positions (LVObjn,LVObj\mathcal{L}^{n}_{V_{\mathit{Obj}}},\mathcal{L}_{V_{\mathit{Obj}}}). We multiply the AtlasNet decoded vertices by the predicted scale and offset them according to the predicted translation to obtain the final object reconstruction. Chamfer loss (LVObj\mathcal{L}_{V_{\mathit{Obj}}}) is applied after translation and scale are applied. When training in hand-relative coordinates the loss becomes:

3 Contact loss

So far, the prediction of hands and objects does not leverage the constraints that guide objects interacting in the physical world. Specifically, it does not account for our prior knowledge that objects can not interpenetrate each other and that, when grasping objects, contacts occur at the surface between the object and the hand. We formulate these contact constraints as a differentiable loss, LContact\mathcal{L}_{Contact}, which can be directly used in the end-to-end learning framework. We incorporate this additional loss using a weight parameter μC\mu_{C}, which we set empirically to 1010.

We rely on the following definition of distances between points. d(v,VObj)=inf⁡w∈VObj∥v−w∥2d(v,V_{\mathit{Obj}})=\inf_{w\in V_{\mathit{Obj}}}\|v-w\|_{2} denotes distances from point to set and d(C,VObj)=inf⁡v∈Cd(v,VObj)d(C,V_{\mathit{Obj}})=\inf_{v\in C}d(v,V_{\mathit{Obj}}) denotes distances from set to set. Moreover, we define a common penalization function lα(x)=αtanh⁡(xα)l_{\alpha}(x)=\alpha\tanh\left(\frac{x}{\alpha}\right), where α\alpha is a characteristic distance of action.

We compute statistics on the automatically-generated grasps described in the next section to determine which vertices on the hand are frequently involved in contacts. We compute for each MANO vertex how often across the dataset it is in the immediate vicinity of the object (defined as less than 3mm3\textrm{mm} away from the object’s surface). We find that by identifying the vertices that are close to the objects in at least 8% of the grasps, we obtain 66 regions of connected vertices {Ci}i∈[ ⁣ ⁣]\{{C_{i}}\}_{i\in[\!\!]} on the hand which match the 55 fingertips and part of the palm of the hand, as illustrated in Figure 3 (left). The attraction term LA\mathcal{L}_{A} penalizes distances from each of the regions to the object, allowing for sparse guidance towards the object’s surface:

Our final contact loss LContact\mathcal{L}_{\mathit{Contact}} is a weighted sum of the attraction LA\mathcal{L}_{A} and the repulsion LR\mathcal{L}_{R} terms:

where λR∈\lambda_{R}\in is the contact weighting coefficient, e.g., λR=1\lambda_{R}=1 means only the repulsion term is active. We show in our experiments that the balancing between attraction and repulsion is very important for physical quality.

Our network is first trained with LHand+LObject\mathcal{L}_{\mathit{Hand}}+\mathcal{L}_{\mathit{Object}}. We then continue training with LHand+LObject+μCLContact\mathcal{L}_{\mathit{Hand}}+\mathcal{L}_{\mathit{Object}}+\mu_{C}\mathcal{L}_{\mathit{Contact}} to improve the physical quality of the hand-object interaction. Appendix C.1 gives further implementation details.

ObMan dataset

To overcome the lack of adequate training data for our models, we generate a large-scale synthetic image dataset of hands grasping objects which we call the ObMan dataset. Here, we describe how we scale automatic generation of hand-object images.

Objects. In order to find a variety of high-quality meshes of frequently manipulated everyday objects, we selected models from the ShapeNet dataset. We selected 88 object categories of everyday objects (bottles, bowls, cans, jars, knifes, cellphones, cameras and remote controls). This results in a total of 27722772 meshes which are split among the training, validation and test sets.

Grasps. In order to generate plausible grasps, we use the GraspIt software following the methods used to collect the Grasp Database . In the robotics community, this dataset has remained valuable over many years and is still a reference for the fast synthesis of grasps given known object models .

We favor simplicity and robustness of the grasp generation over the accuracy of the underlying model. The software expects a rigid articulated model of the hand. We transform MANO by separating it into 16 rigid parts, 3 parts for the phalanges of each finger, and one for the hand palm. Given an object mesh, GraspIt produces different grasps from various initializations. Following , our generated grasps optimize for the grasp metric but do not necessarily reflect the statistical distribution of human grasps. We sort the obtained grasps according to a heuristic measure (see Appendix C.2) and keep the two best candidates for each object. We generate a total of 21K21K grasps.

Textures. Object textures are randomly sampled from the texture maps provided with ShapeNet models. The body textures are obtained from the full body scans used in SURREAL . Most of the scans have missing color values in the hand region. We therefore combine the body textures with 176176 high resolution textures obtained from hand scans from 2020 subjects. The hand textures are split so that textures from 1414 subjects are used for training and 33 for test and validation sets. For each body texture, the skin tone of the hand is matched to the subject’s face color. Based on the face skin color, we query in the HSV color space the 33 closest hand texture matches. We further shift the HSV channels of the hand to better match the person’s skin tone.

Rendering. Background images are sampled from both the LSUN and ImageNet datasets. We render the images using Blender . In order to ensure the hand and objects are visible we discard configurations if less than 100100 pixels of the hand or if less than 40%40\% of the object is visible.

For each hand-object configuration, we render object-only, hand-only, and hand-object images, as well as the corresponding segmentation and depth maps.

Experiments

We first define the evaluation metrics and the datasets (Sections 5.1, 5.2) for our experiments. We then analyze the effects of occlusions (Section 5.3) and the contact loss (Section 5.4). Finally, we present our transfer learning experiments from synthetic to real domain (Sections 5.5, 5.6).

Our output is structured, and a single metric does not fully capture performance. We therefore rely on multiple evaluation metrics.

Hand error. For hand reconstruction, we compute the mean end-point error (mm) over 2121 joints following .

Object error. Following AtlasNet , we measure the accuracy of object reconstruction by computing the symmetric Chamfer distance (mm) between points sampled on the ground truth mesh and vertices of the predicted mesh.

Contact. To measure the physical quality of our joint reconstruction, we use the following metrics.

Penetration depth (mm), Intersection volume (cm3\textrm{cm}^{3}): Hands and objects should not share the same physical space. To measure whether this rule is violated, we report the intersection volume between the object and the hand as well as the penetration depth. To measure the intersection volume of the hand and object we voxelize the hand and object using a voxel size of 0.5cm0.5\textrm{cm}. If the hand and the object collide, the penetration depth is the maximum of the distances from hand mesh vertices to the object’s surface. In the absence of collision, the penetration depth is .

Simulation displacement (mm): Following , we use physics simulation to evaluate the quality of the produced grasps. This metric measures the average displacement of the object’s center of mass in a simulated environment assuming the hand is fixed and the object is subjected to gravity. Details on the setup and the parameters used for the simulation can be found in . Good grasps should be stable in simulation. However, stable simulated grasps can also occur if the forces resulting from the collisions balance each other. For estimating grasp quality, simulated displacement must be analyzed in conjunction with a measure of collision. If both displacement in simulation and penetration depth are decreasing, there is strong evidence that the physical quality of the grasp is improving (see Section 5.4 for an analysis). The reported metrics are averaged across the dataset.

2 Datasets

We present the datasets we use to evaluate our models. Statistics for each dataset are summarized in Table 1.

3 Effect of occlusions

For each sample in our synthetic dataset, in addition to the hand-object image (HO-img) we render two images of the corresponding isolated and unoccluded hand (H-img) or object (O-img). With this setup, we can systematically study the effect of occlusions on ObMan, which would be impractical outside of a synthetic setup.

We study the effect of objects occluding hands by training two networks, one trained on hand-only images and one on hand-object images. We report performance on both unoccluded and occluded images. A symmetric setup is applied to study the effect of hand occlusions on objects. Since the hand-relative coordinates are not applicable to experiments with object-only images, we study the normalized shape reconstruction, centered on the object centroid, and scaled to be inscribed in a sphere of radius 1.

Unsurprisingly, the best performance is obtained when both training and testing on unoccluded images as shown in Table 2. When both training and testing on occluded images, reconstruction errors for hands and objects drop significantly, by 12%12\% and 25%25\% respectively. This validates the intuition that estimating hand pose and object shape in the presence of occlusions is a harder task.

We observe that for both hands and objects, the most challenging setting is training on unoccluded images while testing on images with occlusions. This shows that training with occlusions is crucial for accurate reconstruction of hands-object configurations.

4 Effect of contact loss

In Figure 6, we study the effect of introducing our contact loss as a fine-tuning step. We linearly interpolate λR\lambda_{R} in [​​] to explore various relative weightings of the attraction and repulsion terms.

5 Synthetic to real transfer

Large-scale synthetic data can be used to pre-train models in the absence of suitable real datasets. We investigate the advantages of pre-training on ObMan when targeting FHB and HIC. We investigate the effect of scarcity of real data on FHB by comparing pairs of networks trained using subsets of the real dataset. One is pre-trained on ObMan while the other is initialized randomly, with the exception of the encoders, which are pre-trained on ImageNet . For these experiments, we do not add the contact loss and report means and standard deviations for 5 distinct random seeds. We find that pre-training on ObMan is beneficial in low data regimes, especially when less than 10001000 images from the real dataset are used for fine-tuning, see Figure 8.

The HIC training set consists of only 250250 images. We experiment with pre-training on variants of our synthetic dataset. In addition to ObMan, to which we refer as (a) in Figure 9, we render 20K20K images for two additional synthetic datasets, (b) and (c), which leverage information from the training split of HIC (d). We create (b) using our grasping tool to generate automatic grasps for each of the object models of HIC and (c) using the object and pose distributions from the training split of HIC. This allows to study the importance of sampling hand-object poses from the target distribution of the real data. We explore training on (a), (b), (c) with and without fine-tuning on HIC. We find that pre-training on all three datasets is beneficial for hand and object reconstructions. The best performance is obtained when pre-training on (c). In that setup, object performance outperforms training only on real images even before fine-tuning, and significantly improves upon the baseline after. Hand pose error saturates after the pre-training step, leaving no room for improvement using the real data. These results show that when training on synthetic data, similarity to the target real hand and pose distribution is critical.

6 Qualitative results on CORe50

FHB is a dataset with limited backgrounds, visible magnetic sensors and a very limited number of subjects and objects. In this section, we verify the ability of our model trained on ObMan to generalize to real data without fine-tuning. CORe50 is a dataset which contains hand-object interactions with an emphasis on the variability of objects and backgrounds. However no 3D hand or object annotation is available. We therefore present qualitative results on this dataset. Figure 7 shows that our model generalizes across different object categories, including light-bulb, which does not belong to the categories our model was trained on. The global outline is well recovered in the camera view while larger mistakes occur in the perpendicular direction. More results can be found in Appendix D.

Conclusions

We presented an end-to-end approach for joint reconstruction of hands and objects given a single RGB image as input. We proposed a novel contact loss that enforces physical constraints on the interaction between the two meshes. Our results and the ObMan dataset open up new possibilities for research on modeling object manipulations. Future directions include learning grasping affordances from large-scale visual data, and recognizing complex and dynamic hand actions.

This work was supported in part by ERC grants ACTIVIA and ALLEGRO, the MSR-Inria joint lab, the Louis Vuitton ENS Chair on AI and the DGA project DRAAF. We thank Tsvetelina Alexiadis, Jorge Marquez and Senya Polikovsky from MPI for help with scan acquisition, Joachim Tesch for the hand-object rendering, Mathieu Aubry and Thibault Groueix for advices on AtlasNet, David Fouhey for feedback. MJB has received research gift funds from Intel, Nvidia, Adobe, Facebook, and Amazon. While MJB is a part-time employee of Amazon, his research was performed solely at, and funded solely by, MPI. MJB has financial interests in Amazon and Meshcapade GmbH.

References

APPENDIX

Our main paper proposed a method for joint reconstruction of hands and objects. Below we present complementary analysis for hand-only reconstruction in Section A and object-only reconstruction in Section B. Section C presents implementation details.

Appendix A Hand pose estimation

We first present an ablation study for the different losses we defined on the MANO hand model (Section A.1). Then, we study the latent hand representation (Section A.2). Finally, we validate our hand pose estimation branch and demonstrate its competitive performance compared to the state-of-the-art methods on a benchmark dataset (Section A.3).

As explained in Section 3.1 of the main paper, we define three losses for the differentiable hand model while training our network: (i) vertex positions LVHand\mathcal{L}_{V_{\mathit{Hand}}}, (ii) joint positions LJ\mathcal{L}_{J}, and (iii) shape regularization Lβ\mathcal{L}_{\beta}. The shape is only predicted in the presence of Lβ\mathcal{L}_{\beta}. In the absence of shape regularization, when only sparse keypoint supervision is provided, predicting β\beta without regularizing it produces extreme deformations of the hand mesh, and we therefore fix β\beta to the average hand shape.

Table A.1 summarizes the contribution of each of these losses. Note that the dense vertex supervision is available on our synthetic dataset ObMan, and not available on the real datasets FHB and StereoHands .

We find that predicting β\beta while regularizing it with Lβ\mathcal{L}_{\beta} significantly improves the mean end-point-error on keypoints. On the synthetic dataset ObMan, we find that adding LV\mathcal{L}_{V} yields a small additional improvement. We therefore use all three losses whenever dense vertex supervision is available, and LJ\mathcal{L}_{J} in conjunction with Lβ\mathcal{L}_{\beta} when only keypoint supervision is provided.

A.2 MANO pose representation

As described in Section 3.1 of the main paper, our hand branch outputs a 30-dimensional vector to represent the hand. These are the 30 first PCA components from the 45-dimensional full pose space. We experiment with different dimensionality for the latent hand representation and summarize our findings in Table A.2. While low-dimensionality fails to capture some poses present in the datasets, we do not observe improvements after increasing the dimensionality more than 30. Therefore, we use this value for all experiments in the main paper.

A.3 Comparison with the state of the art

Using the MANO branch of the network, we can also estimate the hand pose for images in which the hands are not interacting with objects, and compare our results with previous methods. We train and test on the StereoHands dataset , and follow the evaluation protocol of by training on 10 sequences from StereoHands and testing on the 2 remaining ones. For fair comparison, we add a palm joint to the MANO model by averaging the positions of two vertices on the front and back of the hand model at the level of the palm. Although the hand shape parameter β\beta allows to capture the variability of hand shapes which occurs naturally in human populations, it does not account for the discrepancy between different joint conventions. To account for skeleton mismatch, we add a linear layer initialized to identity which maps from the MANO joints to the final joint annotations.

We report the area under the curve (auc) on the percentage of correct keypoints (PCK). Figure A.2 shows that our differentiable hand model is on par with the state of the art. Note that the StereoHands benchmark is close to saturation. In contrast to other methods that only predicts sparse skeleton keypoints, our model produces a dense hand mesh. Figure A.1 presents some qualitative results from this dataset.

Appendix B Object reconstruction

In the following, we validate our design choices for the object reconstruction branch. We experiment with object reconstruction (i) in the camera viewpoint (Section B.1) and (ii) with regularization losses (Section B.2).

As explained in Section 3.2 of the main paper, we perform object reconstructions in the camera coordinate frame. To validate that AtlasNet can successfully predict objects in camera view as well as in canonical view, we reproduce the training setting of the original paper . We use the setting where 2500 points are sampled on a sphere and train on the rendered images from ShapeNet . To obtain the rotated reference for the object, we apply the ground truth azimuth and elevation provided with the renderings so that the 3D ground truth matches the camera view. We use the original hyperparameters (Adam with a learning rate of 0.001) and train both networks for 25 epochs. Both for supervision and evaluation metrics, we report the Chamfer distance LVObj=12(∑pminq∥p−q∥22+∑qminp∥q−p∥22)\mathcal{L}_{V_{Obj}}=\frac{1}{2}(\sum_{p}min_{q}\|p-q\|^{2}_{2}+\sum_{q}min_{p}\|q-p\|^{2}_{2}) where qq spans the predicted vertices and pp spans points uniformly sampled on the surface of the ground truth object. We always sample the same number of points on the surface as there are vertices in the predicted mesh. We find that both numerically and qualitatively the performance is comparable for the two settings. Some reconstructed meshes in camera view are shown in Figure A.3. For better readability they also multiply the Chamfer loss by 10001000. In order to provide results directly comparable with the original paper , we also report numbers with the same scaling in Table A.3. Table A.3 reports the Chamfer distances for their released model, our reimplementation in canonical view, and our implementation in non-canonical view. We find that our implementation allows us to train a model with similar performances to the released model. We observe no numerical or qualitative loss in performance when predicting the camera view instead of the canonical one.

B.2 Object mesh regularization

We find that in the absence of explicit regularization on their quality, the predicted meshes can be very irregular. Sharp discontinuities in curvature occur in regions where the ground truth mesh is smooth, and the mesh triangles can be of very different dimensions. These shortcomings can be observed on all three reconstructions in Figure A.3. Following recent work on mesh estimation from image inputs , we introduce regularization terms on the object mesh.

Laplacian smoothness regularization (LL\mathcal{L}_{L}). In order to avoid unwanted discontinuities in the curvature of the mesh, we enforce a local prior of smoothness. We use the discrete Laplace-Beltrami operator to estimate the curvature at each mesh vertex position, as we have no prior on the final shape of the geometry, we compute the graph laplacian LL on our mesh, which only takes into account adjacency between mesh vertices. Multiplying the laplacian LL by the positions of the object vertices VObj\mathcal{V}_{Obj} produces vectors which have the same direction as the vertex normals and their norm proportional to the curvature. Minimizing the norm of these vector therefore minimizes the curvature. We minimize the mean curvature over all vertices in order to encourage smoothness on the mesh.

Laplacian edge length regularization (LE\mathcal{L}_{E}). LE\mathcal{L}_{E} penalizes configurations in which the edges of the mesh have different lengths. The edge regularization is defined as:

where EL\mathcal{E}_{L} is the set of edge lengths, defined as the L2 norms of the edges, and μ(EL2)\mu({\mathcal{E}_{L}^{2})} is the average of the square of edge lengths.

To evaluate the effect of the two regularization terms we train four different models. We train a model without any regularization, two models for which only one of the two regularization terms are active, and finally a model for which the two regularization terms are applied simultaneously. Each of these models is trained for 200200 epochs.

Figure A.4 shows the qualitative benefits of each term. While edge regularization LE\mathcal{L}_{E} alone already significantly improves the quality of the predicted mesh, note that unwanted bendings of the mesh still occur, for instance in the last row for the cellphone reconstruction. Adding the laplacian smoothness LL\mathcal{L}_{L} resolves these irregularities. However, adding each regularization term negatively affects the final reconstruction score. Particularly we observe that introducing edge regularization increases the Chamfer loss by 22% while significantly improving the perceptual quality of the predicted mesh. Introducing the regularization terms contributes to the coarseness of the object reconstructions, as can be observed on the third row, where sharp curvatures of the object in the input image are not captured in the reconstruction.

Appendix C Implementation details

We give implementation details on our training procedure (Section C.1) and our automatic grasp generation (Section C.2).

For all our experiments, we use the Adam optimizer . As we observe instabilities in validation curves when training on synthetic datasets, we freeze the batch normalization layers. This fixes their weights to the original values from the ImageNet pre-trained ResNet18 .

For the final model trained on ObMan, we first train the (normalized) object branch using LObjectn\mathcal{L}^{n}_{\mathit{Object}} for 250 epochs, we start with a learning rate of 10−410^{-4} and decrease it to 10−510^{-5} at epoch 200. We then freeze the object encoder and the AtlasNet decoder, as explained in Section 3.2 of the main paper. We further train the full network with LHand+LObject\mathcal{L}_{\mathit{Hand}}+\mathcal{L}_{\mathit{Object}} for 350350 additional epochs, decreasing the learning rate from 10−410^{-4} to 10−510^{-5} after the first 200200 epochs.

When fine-tuning from our main model trained on synthetic data to smaller real datasets, we unfreeze the object reconstruction branch.

For the FHBc dataset, we train all the parts of the network simultaneously with the supervision LHand+LObject\mathcal{L}_{\mathit{Hand}}+\mathcal{L}_{\mathit{Object}} for 400400 epochs, decreasing the learning rate from 10−410^{-4} to 10−510^{-5} at epoch 300300.

When fine-tuning our models with the additional contact loss, LHand+LObject+μCLContact\mathcal{L}_{\mathit{Hand}}+\mathcal{L}_{\mathit{Object}}+\mu_{C}\mathcal{L}_{\mathit{Contact}}, we use a learning rate of 10−510^{-5}. We additionally set the momentum of the Adam optimizer to zero, as we find that momentum affects negatively the training stability when we include the contact loss.

In all experiments, we keep the relative weights between different losses as provided in the main paper and normalize them so that the sum of all the weights equals 1.

C.2 Heuristic metric for sorting GraspIt grasps

We use GraspIt to generate grasps for the ShapeNet object models. GraspIt generates a large variety of grasps by exploring different initial hand poses. However, some initializations do not produce good grasps. Similarly to we filter the grasps in a post-processing step in order to retain grasps of good quality according to a heuristic metric we engineer for this purpose.

For each grasp, GraspIt provides two grasp quality metrics ε\varepsilon and vv . Each grasp produced by GraspIt defines contact points between the hand and the object. Assuming rigid contacts with friction, we can compute the space of wrenches which can be resisted by the grasp: the grasp wrench space (GWS). This space is normalized with relation to the scale of the object, defined as the maximum radius of the object, centered at its center of mass. The grasp is suitable for any task that involves external wrenches that lie within the GWS. vv is the volume of the 6-dimensional GWS, which quantifies the range of wrenches the grasp can resist. The GWS can further be characterized by the radius ε\varepsilon of the largest ball which is centered at the origin and inscribed in the grasp wrench space. ε\varepsilon is the maximal wrench norm that can be balanced by the contacts for external wrenches applied coming from arbitrary directions. ε\varepsilon belongs to $$ in the scale-normalized GWS, and higher values are associated with a higher robustness to external wrenches.

We require a single value to reflect the quality of the grasp in order to sort different grasps. We use the norm of the [ε,v][\varepsilon,v] vector in our heuristic measure of grasp quality. We find that in the grasps produced by GraspIt, power grasps, as defined by in which larger surfaces of the hand and the object are in contact, are rarely produced. To allow for a larger proportion of power grasps, we use a multiplier γpalm\gamma_{palm} which we empirically set to 11 if the palm is not in contact and 33 otherwise. We further favor grasps in which a large number of phalanges are in contact with the object by weighting the final grasp score using NpN_{p}, the number of phalanges in contact with the object, which is computed by the software.

The final grasp quality score GG is defined as:

We find that keeping the two best grasps for each object produces both diverse grasps and grasps of good quality.

Appendix D Qualitative results on CORe50 dataset

We present additional qualitative results on the CORe50 dataset. We present a variety of diverse input images from CORe50 in Figure A.5 alongside the predictions of our final model trained solely on ObMan.

The first row presents results on various shapes of light bulbs. Note that this category is not included in the synthetic object models of ObMan. Our model can therefore generalize across object categories. The last column shows some reconstructions of mugs, showcasing the topological limitations of the sphere baseline of AtlasNet which cannot, by construction, capture handles.

However, we observe that the object shapes are often coarse, and that fine details such as phone antennas are not reconstructed. We also observe errors in the relative position between the object and the hand, which is biased towards predicting the object’s centroid in the palmar region of the hand, see Figure A.5, fourth column. As hard constraints on collision are not imposed, hand-object interpenetration occurs in some configurations, for instance in the top-right example. In the bottom-left example we present a failure case where the hand pose violates anatomical constraints. Note that while our model predicts hand pose in a low-dimensional space, which implicitly regularizes hand poses, anatomical validity is not guaranteed.