CPF: Learning a Contact Potential Field to Model the Hand-Object Interaction

Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, Cewu Lu

Introduction

It is essential to model hand-object interaction from a single image for understanding the human activities, in which simulating a physically plausible grasp is also crucial for VR//AR, teleoperation, and grasping applications. Given an image as input, the problem aims to not only estimate proper hand-object pose but also to recover a natural grasp configuration. While estimating hand or object alone has made a considerable success over the past decades, simultaneously estimating hand-object pose with interaction has only emerged in the past few years.

Previous works on joint hand-object estimation usually treat the contact as a result of the correct pose estimation . Apparently, if the hand and object can be perfectly recovered, the contact between them will also be satisfied. Yet, such perfection cannot be achieved in practice. Since contact can provide rich cues to guide accurate pose and natural grasp, more attention has been recently drawn to the contact modeling and contact representation . And several contact datasets have been released to the community. However, a solution of properly integrating contact modeling into the current hand-object pose estimation pipeline has remained an open research question. The existing methods either exploit distance-based attraction and repulsion to mitigate disjointedness and interpenetration, or refine the predicted pose in virtue of physics simulators . While the both solutions are considered to be irrelevant to contact semantics, which we will explain later, the latter solutions also lack flexibility on hand pose and shape.

To model the contact, we propose an explicit representation named Contact Potential Field (CPF, §4). It is built upon the idea that the contact between a hand and an object mesh under grasp configuration is multi-point contact, which involves multiple hand-object vertex pair affinities. These affinities are regarded as the contact semantics, which depict the pairing of the hand-object vertices that come into contact with each other during the interaction. When noisy predicted hand and object are disjointed from each other, we shall apply an attraction to pull these vertex pairs close; While the hand and object are intersected, we shall have a repulsion to push them away. Contacts of those affinitive vertex pairs are the result of equilibrium between the attraction and repulsion. In this paper, we treat each contacting HO vertex pair as a spring-mass system. First, the two end-points of spring is a counterpart of the two HO vertices in affinity. Second, the spring’s elastic property is another counterpart of the intensity of the vertex pair affinity. In this way, we can model the HO interaction with a potential field, as we call it CPF, which is determined by minimal elastic energy at the grasp position. Therefore, estimating the HO pose under contact is equivalent to minimizing the elastic energy inside CPF. Representing contact as CPF has two main advantages. First, compared with contact heuristic with proximity metrics or distance field , CPF is able to assign per-vertex contact semantics (contact points on different hand part) to object mesh. Second, by minimizing the elastic energy, CPF can uniformly avoid interpenetration and control the disjointedness. Based on CPF, we also propose a novel learning-fitting hybrid framework namely for Modeling the Interaction of Hand and Object, as we call it MIHO (§5).

Another problem with the existing methods is the representation of the hand model. Most researches adopted a skinning model, MANO , to represent hand. MANO is considered to be flexible and deformable with its pose and shape parameters. However, fitting on these high DoFs parameters is prone to anatomical abnormality. Researches in the robotics community adopted a dexterous hand in the off-the-shelf grasping software , which can almost guarantee a valid pose. But the rigidity of those rod-like hand is less suitable for applications in CV//CG. To make the best of both worlds, we propose a novel anatomically constrained hand model namely A-MANO (§3). It inherits the formulation of the skinning model and constrains the hand joints’ rotation within a proposed twist-splay-bend frame (Fig. 2).

For evaluation, we report our scores on FHB and HO3D dataset in terms of reconstruction and physical quality metrics. Note that, the ground truth of FHB is noisy and suffers from severe interpenetration . Since our method can avoid the penetration in the first place, our results are more visually and physically plausible. Therefore, we argue that, in this dataset, a higher reconstruction score does not necessarily benchmark the performance of the method. While on HO3D, we achieve state-of-the-art performance on both reconstruction and physical metrics. The contributions of this paper are as follows.

We highlight contact in the hand-object interaction modeling task by proposing an explicit representation named CPF.

We introduce A-MANO, a novel anatomical-constrained hand model that helps to mitigate pose’s abnormality during optimization.

We present a novel framework, MIHO, for modeling hand-object interaction. It can achieve state-of-the-art performance on several benchmarks.

Related Work

3D Hand Reconstruction. Most of the existing 3D hand reconstruction methods adopted a parametric skinning hand, e.g. MANO as a template. To drive MANO, it is crucial to obtain joint rotation along hand kinematic tree. Boukhayma et al. firstly proposed to regress the PCA components of the rotations. Later, directly regressing the full rotations from 3D positions has shown better performance. However, those high DoF regression is prone to pose abnormality. Thus, Spurr et al. exploited biomechanical constraints over hand joints in training scheme. Different from , we apply rotation constraints over the axes and angles in the proposed twist-splay-bend coordinate frame.

Hand-object Pose Estimation. In a wide range of topics in modeling hand-object interaction, the most commonly referred one is HO pose estimation . In this regard, the earlier methods focused on either hand or object pose alone, or estimated hand in grasping pose with knowing object shape prior . Jointly estimating hand and object pose was firstly presented by Romero et al. via searching for nearest neighbors in a large database. Recently, learning-based frameworks have emerged in this area. Hasson et al. proposed two learning frameworks to recover hand-object meshes, one by synthesizing HO data under manipulation and the other by exploiting photometric consistency over video sequence . Doosti et al. employed the graph neural networks to lift the 2D HO keypoints into 3D space. Tekin et al. adopted 3D YOLO to predict HO pose in one stage. Korrawe et al. recovered HO model in a form of Signed Distance Function .

Contact Heuristic. Exploiting contact heuristic in hand-object interaction can be traced back to several decades before . Early works utilized some shape-specified contact physics (e.g. cones and blocks ) or predefined grasp as prior. Studies on capturing or imitating HO interaction also leveraged contact to satisfy the reality. Later, the studies on grasping synthesis and tracking turned to physical simulators to circumvent model’s intersection. Multi-point contact formulation was proposed in , which we found useful when applying physical constraints, e.g. used contact points to resolve penetration. For unified attraction and repulsion, most works employed heuristic such as proximity metric , signed distance function , predefined contact pattern , or turned to simulator for simplicity. Recently, Antotsiou et al. refined the grasp by attracting fingers to its nearest point on object surface w.r.t distance-based energy. Hasson et al. applied well-designed interaction losses which are also based on proximity metric. Although our method differs from all of the previous methods in terms of contact heuristic, we consider that both and are still strong baselines. Thus we will compare our contact heuristic with theirs.

Anatomically Constrained A-MANO

Fitting on 15 joint rotations of MANO requires high DoFs regression which may cause abnormal hand posture as shown in Fig. 7. Since the human hand can be modeled in a kinematic tree, and the majority of the joints only have one DoF about the bend axis, we can impose constraints over the rotation about the unwanted axes. Therefore the proposed twist-splay-bend Cartesian coordinate frame can be assigned to each joint along the kinematic tree. The frame’ s x,y,zx,y,z axes are coaxial to the 3 revolute directions: twist, splay, and bend direction on the basis of hand anatomy (Fig. 2 right). Then we can impose axial constraints in the twist and splay axes, and impose angular constraints w.r.t the bend angle. Details of the twist-splay-bend frame are elaborated in Supp A.1.

Anchors.

Since the hand mesh of different subjects are almost identical in the subdivision of hand region (e.g. phalanges), we can interpolate several representative points (later we call it anchors) on hand mesh to largely reduce the number of HO vertex pairs. Instead of attaching springs from object mesh to all the affinitive vertices on hand mesh, we only attach them on the several hand subregion centers, as we call it anchors (Fig. 2 left). According to the statistics on the contact frequency of different hand parts, we first divide the full hand palm into 17 subregions: 3 for each phalange of 5 fingers, 1 for metacarpals, and another for carpals. Then, we interpolate up to 4 anchors for each subregion. We ignore all the vertices on the back side of the hand. Details of subregion division and anchors interpolation are described in Supp A.2, A.3.

Contact Potential Field

A single contact is modeled as a spring-mass system which consists of a spring and two mass points on each side (hand and object). When the spring is at its rest position, it does not store energy, whilst it is stretched or compressed, according to Hooke’s Lawhttps://en.wikipedia.org/wiki/Hookes_law, it will store the elastic potential energy with the form: 12k∣Δl∣2\frac{1}{2}k|\bm{\Delta l}|^{2}, where kk is the spring elasticity, and ∣Δl∣|\bm{\Delta l}| is a certain “distance” metric w.r.t. the spring’s rest position.

In CPF, we define two types of spring: attractive spring and repulsive spring. The goal of attractive spring is to pull the hand vertex vh\bm{v}^{h} toward the object vertex vo\bm{v}^{o} based on a given HO vertex pair affinity. And the goal of repulsive spring is to push the vh\bm{v}^{h} away from vo\bm{v}^{o} along the vo\bm{v}^{o} ’s normal if the vh\bm{v}^{h} is in the vicinity of vo\bm{v}^{o}. Apart from these definitions, we should also point out that the attractive spring is bound with a certain pair of HO vertex affinity, while the repulsive spring only takes effect in the neighborhood of HO vertex pairs at some point.

- Attractive Spring. We define the rest length of attractive spring as in which the hand vertex and object vertex are in perfect contact, and the distance metric ∣Δl∣|\bm{\Delta l}| as Euclidean distance. Given a HO affinity that includes a vertex pair: vih\bm{v}^{h}_{i} and vjo\bm{v}^{o}_{j}, the ∣Δlijatr∣|\bm{\Delta l}^{\rm{atr}}_{ij}| is equal to ∥vih−vjo∥2\left\|\bm{v}^{h}_{i}-\bm{v}^{o}_{j}\right\|_{2}. The potential energy of the current attractive spring is given by:

- Repulsive Spring. We hope that the repulsion energy is high when vih\bm{v}^{h}_{i} is penetrating or in the vicinity of vjo\bm{v}^{o}_{j}, but gradually decays as the vih\bm{v}^{h}_{i} moves away from the object, and finally becomes negligible at certain distance. Given a proximate HO vertex pair: vih\bm{v}^{h}_{i} and vjo\bm{v}^{o}_{j}, We define a repulsive spring to model this behavior. Supposing that the repulsive spring has the rest position at +∞+\infty away along the object normal njo\mathbf{n}^{o}_{j}. We adopt a heuristic distance metric ∣Δl∣=e−∣Δlijrpl∣−e−∞=e−∣Δlijrpl∣|\bm{\Delta l}|=e^{-|\bm{\Delta l}^{\rm{rpl}}_{ij}|}-e^{-\infty}=e^{-|\bm{\Delta l}^{\rm{rpl}}_{ij}|}, where ∣Δlijrpl∣=(vih−vjo)⋅njo|\bm{\Delta l}^{\rm{rpl}}_{ij}|=(\bm{v}_{i}^{h}-\bm{v}_{j}^{o})\cdot\mathbf{n}_{j}^{o} is the projection of the (vih−vjo)(\bm{v}_{i}^{h}-\bm{v}_{j}^{o}) on the object normal njo\mathbf{n}_{j}^{o}. Thus, the potential energy of the current repulsive spring is

In literature, adopting repulsive effect along surface normal can be found in . (Eq. 10) also discussed that e−(⋅)e^{-(\cdot)} is an efficient heuristic concerning sub-sampled set of vertices.

Grasping inside Contact Potential Field.

By collecting all the attractive and repulsive springs, to form a natural grasp is equivalent to minimize the elastic energy:

As discussed in §3, the hand vertices can be simplified to subregion anchorsanchors, which will largely relax the difficulty of learning and fitting inside the CPF. Thus, for attractive spring, we replace the Δlij\bm{\Delta l}_{ij} in Eq.1 to Δlij′=ai−vjo\bm{\Delta l}^{\prime}_{ij}=\bm{a}_{i}-\bm{v}^{o}_{j}, where ai\bm{a}_{i} is the closest anchor to vih\bm{v}^{h}_{i}. Besides, we would like to have the repulsion force be only applied to those HO affinity pairs that are of vertices in vicinity. Thus we set zero the repulsion energy when the vertex distance ∥vjo−vih∥2\|\bm{v}_{j}^{o}-\bm{v}_{i}^{h}\|_{2} is greater than a threshold trpl=20 mmt_{\rm{rpl}}=20\ mm.

While the attraction energy is bound with certain HO affinities, the repulsion energy is rather ambient and affinity-agnostic. To integrate the CPF into learning framework, we only consider the kijatrk_{ij}^{\rm{atr}} as the prediction of neural network. To enable this, network shall have the abilities of 1) pairing the hand anchors and object vertices into HO affinity pair, e.g. (ai,vjo)(\bm{a}_{i},\bm{v}^{o}_{j}); and 2) regressing the intensity of those affinity pairs, e.g. kijatrk_{ij}^{\rm{atr}}. These require annotation of the attractive springs katrk^{\rm{atr}}.

Given the ground-truth (gt.) HO pose and their mesh model, we automatically annotate each kijatrk_{ij}^{\rm{atr}} based on a heuristic of the (ai,vjo)(\bm{a}_{i},\bm{v}^{o}_{j}) pair distance. Since each ai\bm{a}_{i} may be included in several affinity pairs, we hope the attraction energy stored in each spring at gt. HO pose is well balanced. Thus we assign the gt. k^ijatr\hat{k}^{\rm{atr}}_{ij} a value that is inverse-proportional to the gt. ∣Δl^ijatr∣|\bm{\Delta}\hat{\bm{l}}^{\rm{atr}}_{ij}|. In order to train the network, we also bound the magnitude of k^ijatr\hat{k}^{\rm{atr}}_{ij} by and 11. Here we only provide a glimpse of the annotation heuristic of k^ijatr\hat{k}^{\rm{atr}}_{ij}:

Empirically, we set the scale factor s=20 mms=20\ mm and reject those HO affinities with gt. ∣Δl^ijatr∣≥20 mm|\bm{\Delta}\hat{\bm{l}}^{\rm{atr}}_{ij}|\geq 20\ mm. As for the elasticity of repulsive spring, we empirically set all kijrpl{k}_{ij}^{\rm{rpl}} to 1×10−31\times 10^{-3}. Detailed analysis of the gt. k^atr\hat{k}^{\rm{atr}} and the attraction-repulsion equilibrium are provided in Supp B.1, B.2,

Hybrid Framework – MIHO

With respect to the proposed CPF (§4), our approach MIHO models the hand-object interaction in three stages, namely HoNet (§5.1), PiCR (§5.2), and GeO (§5.3).

2 Pixel-wise Contact Recovery Module, PiCR

With the coarse meshes of hand and object in HoNet, PiCR learns to recover the CPF by firstly paring the hand anchors and object vertices into HO affinity pairs and then regressing the spring elasticities that describe the affinities. To achieve this, PiCR yields three cascaded outcomes: 1) Vertex Contact (VC) decides which vertices on object are in contact with hand; 2) Contact Region (CR) decides the subregion that is most likely to contact with those vertices in VC; 3) Anchor Elasticity (AE) represents the elasticities of the attractive springs. With VC, CR, and AE, we can then recover the CPF as illustrated in Fig. 4.

where fj=pjf_{j}=p_{j} if the gt. v^jo\hat{\bm{v}}^{o}_{j} belongs to any HO affinity, otherwise fj=(1−pj)f_{j}=(1-p_{j}), and the pjp_{j} is the predicted probability at VC[j][j]. \mathds1jimg\mathds{1}^{img}_{j} denotes whether the vertex vjo\bm{v}^{o}_{j} is projected inside the image. αj\alpha_{j} is inverse class frequency and γ\gamma is empirically set to 22.

where the k^ijatr\hat{k}_{ij}^{\rm{atr}} is the gt. elasticity described in §4.

With the predicted VC, CR and AE, as well as the coarse meshes Vo,Vh\mathcal{V}^{o},\mathcal{V}^{h} in HoNet, PiCR finally recovers the CPF and collects the elastic energy EelastE_{\rm{elast}} as described in Algm.1. We empirically set the probability threshold of VC: tvc=0.8t_{\rm{vc}}=0.8 and the distance threshold: trpl=20 mmt_{\rm{rpl}}=20\ mm.

PiCR’s Framework. The proposed PiCR consists of a backbone bb that extracts features from image, an encoder pp that converts image features to object vertex features, and 3 heads hvch_{\rm{vc}}, hcrh_{\rm{cr}} and haeh_{\rm{ae}} which sequentially convert those features into VC, CR, and AE. As illustrated in Fig. 3, the process of feature extraction in PiCR can be expressed as:

where b(⋅)b(\cdot) is the hourglass networks , π(⋅)\pi(\cdot) is the perspective camera projection, and f(⋅)f(\cdot) stands for aligning Vo\mathcal{V}^{o} ’s 2D projection π(Vo)\pi(\mathcal{V}^{o}) with the image features b(I)b(\mathcal{I}) through bilinear sampling. Inspired from Eq.(1) in , we also append the object’s root-relative zz value z(Vo)z(\mathcal{V}^{o}) at the end of f(⋅)f(\cdot) to form the pixel-wise features F′\mathcal{F}^{\prime}. Next, a PointNet encoder p(⋅)p(\cdot) is adopted to convert F′\mathcal{F}^{\prime} to its point-wise features F\mathcal{F}.

The process of three PiCR’s heads can be expressed as:

where all the heads are presented as multi-layer perceptrons. We provide implementation details in Supp D.1.

3 Grasping Energy Optimizer, GeO

The fitting part: Grasping Energy Optimizer (GeO) aims to refine the HO pose w.r.t. the recovered CPF. For the object part, we adjust its 6D pose Po∈se(3)\bm{P}_{o}\in\mathfrak{se}(3). For the hand part, we jointly adjust the A-MANO’ s 15 joint rotations {Rj∈so(3) ∣ j≤15}\{\bm{R}_{j}\in\mathfrak{so}(3)\ |\ j\leq 15\} and a wrist pose Pw∈se(3)\bm{P}_{w}\in\mathfrak{se}(3).

In order to mitigate the abnormal hand posture during optimization, we also define an anatomical cost function Lanat\mathcal{L}_{\rm{anat}} that penalizes the unwanted axial components and angular values of the 1515 rotations in the proposed twist-splay-bend coordinate frame. First, for the joints along hand kinematic tree, we penalize the component of rotation axis arot\mathbf{a}^{rot} on twist direction: ntwist\mathbf{n}^{twist}, since any component that causes the finger twisting along its pointing direction is prohibited. Second, for the joints that do not belongs to 5 knuckles, we also penalize the component of arot\mathbf{a}^{rot} on splay direction: nsplay\mathbf{n}^{splay}. Last, we penalize the rotation angle ϕbend\phi^{bend} that revolves about the bend axis if it is greater than π/2\pi/2. The total anatomical cost can be written as:

We also penalize the offset of the refined hand-object vertices ∗Vo{}^{*}\mathcal{V}^{o}, ∗Vh{}^{*}\mathcal{V}^{h} from their initial estimation Vh{\mathcal{V}}^{h}, Vo{\mathcal{V}}^{o} in form of l2 distance: Loffset\mathcal{L}_{\rm{offset}}. We implement GeO in PyTorch with Adam solver. The whole optimization process can be expressed as:

Experiments and Results

We would like to train and evaluate MIHO w.r.t. the real-world dataset that involves human hand interacting with textured object. In the community, there exist mainly four datasets that contain images and ground-truth 3D HO annotation, namely ObMan , FHB and HO3D and ContactPose . However, only FHB and HO3D satisfy our requirements in this study.

FHB is a first-person RGBD video dataset of hand in manipulation with objects. The ground-truth of hand poses was captured via magnetic sensors. In our experiments, we use a subset of FHB that contains 4 objects with a scanned model and pose annotation. We adopt the action split following the protocol given by , and filter out the samples with a minimum HO distance greater than 5 mmmm, which yields us 7223 samples for training and 7373 for testing.

HO3D.

HO3D is another dataset that contains precise hand-object pose during the interaction. Due to historical reasons, there is two versions of HO3D, namely v1 and v2 . In our experiments, we mainly compare our methods with the baseline on HO3Dv1, but also conduct several comparisons with the recently released pre-trained model of on HO3Dv2. Similar to FHB, we filter out samples with distance threshold 5mmmm. It’s also worth mentioning that, since our method requires a known object model, as well as a stable grasping configuration, nearly 5448 samples in HO3Dv2 test set are not suitable for our methods to report. Therefore, we manually select 6076 samples in HO3dv2 test set to compare MIHO with . We call this split by HO3Dv2-. Besides, training HO3Dv1 in previous methods requires an extra synthetic dataset that is not publicly available. Thus we manually augment the HO3Dv1 training set (referred as HO3Dv1+) and reproduce the results (referred as +) comparable with those in . Details of HO3Dv2- selection and the augmentation procedures are provided in Supp C.1, C.2.

2 Metrics

Modeling the HO interaction requires not only a proper pose of both hand and object but also a natural grasp configuration. Here, we report 5 metrics in total that cover both reconstruction and grasp quality. Note that, since considering either of those metrics alone may yield misleading comparison, we consider them together for evaluation.

MPVPE. We compute the mean per vertex position error for both hand and object in camera space to assess the quality of pose estimation.

Penetration Depth (PD). To measure how deep that the hand is penetrating the object’s surface, we calculate the penetration depth that is the maximum distance of all the penetrated hand vertices to their closest object surface.

Solid Intersection Volume (SIV). To measure how much space intersection that occurs during estimation, we voxelize the object mesh into 80380^{3} voxels, and calculate the sum of the voxel volume inside the hand surface.

Disjointedness Distance (DD). We also encourage stable HO contact, which can be depicted as attracting fingertips onto the object surface. Therefore, we define the disjointedness metrics as the average distance of hand vertices in 5 fingertips region to their closet object surface.

Simulation Displacement (SD). We further evaluate the grasp stability in a modern physics simulator . We measure the average displacement of object’s center over a fixed time period by holding the hand steadily and applying gravity to the object .

3 Comparison with State-of-the-Arts

For the FHB dataset, we compare our methods with the previous SOTA of hand-object reconstruction. For , we select the results under the setting of full data supervision. Since didn’t exploit any repulsion and attraction loss during training, direct comparisons on intersection and disjointedness may not be convincing enough. While the contact losses were considered in another work named ObMan , it only represented the genus 0 object mesh as a deformable icosphere, which is also not directly comparable with ours (known object model). To ensure rational comparisons, we migrate the repulsion loss and attraction loss from ObMan to the MeshRegNet in , and reproduce the results on par with it. We call this adaptation: ObMan∗. For the HO3Dv1 dataset, we compare our results with the reproduced +.

We report our results under two experimental settings: 1) hand-alone that fixes the object at the initial prediction in HoNet, and only optimizes the hand pose in GeO; 2) hand-object that jointly optimizes the hand and object poses in GeO. In Tab.1 we show our comparisons with the previous SOTA in all 5 metrics. For FHB dataset, as analyzed in , its ground-truth suffers from frequent interpenetration. We find that lower vertex error does not necessarily benchmark a higher reconstruction quality. Indeed, as shown in Tab.1 (col. 4, 5), either ground-truth or reveals substantial solid intersection volume, penetration depth and disjointedness. We find that MIHO outperforms by a margin of 3.71 mmmm in penetration depth, 9.34 cm3cm^{3} in solid intersection volume, and 14.99 mmmm in disjointedness distance, while only suffers from minor performance cost in hand MPVPE of 2.03 mmmm and object MPVPE of 0.51 mmmm. In the mean time, our simulation displacement also demonstrates the stability of our predicted grasp. These are consistent with our expectation that the CPF can by nature repulse intersection away and attract disjointedness to touch. As for HO3Dv1 testing set, our method also outperformed the previous SOTA over the most metrics. In terms fo simulation displacement, we found + slightly outperforms us by 1.98 mmmm. Based on our inspection in the Bullet simulator, their stability are mainly attributed to the forces resulting from the intersection that balance each other. Visual comparisons are shown in Fig. 5. As for HO3Dv2, since we only test MIHO on the subset: HO3Dv2-, our results are not directly suitable for submitting to its online evaluation server. Thus, we only report the object 3D vertex errors on HO3Dv2- based on the given annotation. We firstly align the predicted object vertex to the predicted hand wrist joint, then compute the wrist-relative object vertex error with those in ground-truth. Detailed comparisons are in Tab. 1 (col. 11, 12).

4 Ablation Study

In this experiment, we further evaluate the effectiveness of the proposed CPF and A-MANO. In the main text, we include three of the most representative studies. The ablation studies are mainly conducted on the FHB test set with action split. For more studies on 1) impact from the magnitude of krplk^{\rm{rpl}}; 2) A-MANO with PCA pose; 3) unwanted twist correction; please visit Supp D.2.

Comparison with simple Distance-based Contact Heuristics. To show the superiority of the CPF over the distance-based contact heuristics, we compare the fitting stage of MIHO with two simple yet strong baselines: (a) Vanilla Contact that removes the EelastE_{\rm{elast}} term in Eq. 11 and purely attracts the anchors on fingertips to its nearest object vertex (similar to ) in a given threshold which we set as 20 mmmm; (b) ObMan Contact that replaces EelastE_{\rm{elast}} in Eq. 11 by the well-designed interaction losses in ObMan . All the three experiments start from the same HO pose predicted by HoNet (§5.1). We show in Tab. 2 that by exploiting CPF, MIHO can surpass the simple baselines on most of the metrics. Note that, since both (a) and (b) directly optimize the disjointedness term, their results show better resistance on it. The last column in Tab. 2 shows that our methods can save average time per iteration by 46% compared with ObMan Contact. We also conduct two qualitative comparisons in Fig. 6. The first one shows that CPF can learn the contact semantics to guide the optimization that better matches visual cues, whereas the Vanilla Contact fails to form a valid grasp. The second shows that CPF can maintain subtle interaction, as no attraction will be applied on those non-affinitive vertex pairs (see ring and pinky fingers when unscrewing the juice cap).

Effectiveness of Repulsive Springs. To measure the efficacy of repulsive springs in CPF, we remove all the repulsion energy ErplE^{\rm{rpl}} induced by them, leaving the attraction as the unique type of energy applied on hand and object. As we expected, the result in Tab. 3 witnesses the accumulation of PD and SIV. To note, even without the repulsive springs, we still witness a remarkable improvement of PD and SIV over the FHB ground-truth. This is attributed to the repulsive behavior of the attractive springs: when hand is inside the object surface, the energy stored in the attractive springs will act as repulsion that pushes out the hand.

Effectiveness of the Anatomical Constraints. We further highlight the efficacy of adopting the anatomical constraints. We conduct a contrastive experiment whose only difference is the absence of Lanat\mathcal{L}_{\rm{anat}}. Both experiments start from a zero (flat) hand and minimize the EelastE_{\rm{elast}} based on the same predicted CPF. We show in Fig. 7 that the anatomical constraints are able to effectively prevent abnormality during the optimization.

Conclusion

In this work, we propose a novel contact representation named CPF and a learning-fitting hybrid framework MIHO to help modeling hand and object interaction. Comprehensive evaluations show that our methods, while being able to recover precise hand-object pose, can also effectively 1) avoid interpenetration and control disjointedness, and 2) prevent abnormality in hand pose. We hope CPF can serve as an effective contact representation for future works on hand-object interaction. Later, we also plan to develop for an object-agnostic representation of CPF, for the interaction in general cases.

References

Appendix

In the supplemental document, we provide:

Detailed Analysis of the Spring’s Elasticity.

Appendix A Anatomically Constrained A-MANO

In this section, we introduce the proposed twist-splay-bend frame of A-MANO. Both the original MANO and our A-MANO hand model are driven by the relative rotation at each articulation. To mitigate the pose abnormality, we apply constraints on the rotation axis-angle Rotation cay be represented as rotating along an axis by an angle.. We intend to decompose the rotation axis into three components to the three axes of a Euclidean coordinate frame, in which each component depicts the proportion of rotation along that axis. Obviously, there have infinity choices of the three orthogonal axes. MANO adopts 16 identical coordinate frames whose 3 orthogonal axes are not coaxial to the direction of the hand kinematic tree (Fig. 8 left). Different from MANO, we follow the Universal Robot Description Format (URDF) that describe each articulation along the hand kinematic tree as a revolute jointhttps://en.wikipedia.org/wiki/Revolute_joint. Nevertheless, a revolute joint only has one degree of freedom, which is not enough to drive the motion of a real hand. Thus, we assign each articulation with three revolute joints, named as twist, splay and bend (Fig. 8 right),

Here, we elaborate the conversion from the MANO’s all identical coordinate system of to our twist-splay-bend frame in three steps. For each articulation, we first compute the twist axis as the vector from the child of the current joint to itself. Then we employ MANO’s yy (up) axis and derive the bend axis that is calculated from cross product on the twist and yy axes. Finally, we obtain the splay axis by applying cross product on the bend and twist axes. We illustrate the above procedures in Fig. 9.

A.2 Hand Subregion Assignment

As introduced in main text §3 (Anchors.), we divide the hand palm into 17 subregions, and interpolate the vertices in each subregion into representative anchor / anchors. In this part, we will firstly discuss how we assign hand vertices to several subregion.

According to hand anatomy, the linkage bones consists of carpal bones, metacarpal bones, and phalanges, where phalanges can be further divided into three kinds: proximal phalanges, intermediate phalanges, and distal phalanges. Here we assume the link between MANO joints are a counterpart of linkage bones on hand. We now assign the vertices of MANO into 17 subregions based on the linkage bones. The subregions’ names and abbreviations are defined in Fig. 10. For clarity, we number the MANO links from 1 to 20 as illustrated in Fig. 11 (left).

To assign the MANO vertices to its corresponding region, we need firstly assign the vertices to the link that lies inside the region. This is achieved by control points. For link 0-3, 5-7, 9-11, 13-15, 17-20, we set one control point at the midpoint of the link’s ends, while for link 4, 8, 12, and 16, we set two control points at the upper and lower third of the link’s ends. For clarity, we also number the control points from 0 to 23 as shown in Fig.11 (middle). After a list of control points are obtained, we label each hand vertex to one of these control points by querying which control point it has the least distance from. Finally, we merge the vertices that belong to control points 0, 5, 10, 15, and 20 to derive subregion of Palm Metacarpal, and merge those vertices that belong to control points 4, 9, 14, 19 to derive subregion of Carpal .

A.3 Hand Anchor Selection

Here we elaborate on how we select the anchors based on the subregions and their control points. To ensure these anchors can be used in a common optimization framework and keep their representative power during the process of optimization, we propose the following three protocols: a) Anchors should be located on the surface of the hand mesh. b) Anchors should distribute uniformly on the surface of the region it represents. c) Anchors can be derived from hand vertices in a differentiable way.

We utilize control points introduced in §A.2 to derive anchors. Since the anchor selection is independent of hand’s configuration, we adopt a flat hand in the canonical coordinate system. As illustrated in Fig.11 (middle, right), the control points are roughly uniformly distributed in each subregion. Each control point will correspond to an anchor of that subregion. The Carpal is an only exception: we select only 3 over 5 (ID: 5, 10, 20) of the control points in the subregion of Carpal for anchor derivation.

To derive an anchor from a control point, we need to get one face (consist of 3 integers) and two weights. 1) Non-tip regions. For non-tip regions, we cast a ray that is originated from each control point in a certain subregion, and pointing to the palm surface. We retrieve the first intersection of the ray with hand mesh. This intersection will be the anchor that correspond to the control point, also the subregion. 2) Tip regions. For tip regions, we would select three anchors of each control point to increase the density of anchors in that subregion, as tip involves more contact information during manipulation. For the control point in tip subregions, we first cast a ray originated from the control point and get the intersection point on the hand mesh. Then a cone is created with the control point as apex, the intersection point as the base center, and a base radius. The base radius is estimated by the maximum distance of vertices in the subregion to their control point. Three generatrices equally distributed on cone surface are selected as new ray casting directions. We cast three rays from the control point in the direction of the three generatrices and retrieve the intersection points with hand mesh. These intersection points will be selected as anchors to that control point in the fingertip regions.

Appendix B Spring’s Elasticity

Here we illustrate elastic energy between a pair of points vih\bm{v}_{i}^{h} and vjo\bm{v}_{j}^{o}, denoting one vertex on hand surface and another vertex on object surface respectively. The vertex on object surface binds with a vector njo\mathbf{n}_{j}^{o} representing the normal direction at this vertex (also the direction of repulsion). Then we compute the offset vector Δlijatr=vih−vjo\bm{\Delta l}_{ij}^{\rm{atr}}=\bm{v}_{i}^{h}-\bm{v}_{j}^{o}, and the projection of the offset vector on object normal njo\mathbf{n}_{j}^{o}: ∣Δlijrpl∣=(vih−vjo)⋅njo|\bm{\Delta l}_{ij}^{\rm{rpl}}|=(\bm{v}_{i}^{h}-\bm{v}_{j}^{o})\cdot\mathbf{n}_{j}^{o}. ∣Δlijrpl∣|\bm{\Delta l}_{ij}^{\rm{rpl}}| is positive if vih\bm{v}_{i}^{h} falls outside the object, and negative if vih\bm{v}_{i}^{h} falls inside the object. We use an exponential function here to provide magnitude and gradient heuristic for optimizer: a) the less ∣Δlijrpl∣|\bm{\Delta l}_{ij}^{\rm{rpl}}| is, the more vih\bm{v}_{i}^{h} penetrates into the object. The gradient of repulsive energy will be an exponential increasing function of Δlijrpl\bm{\Delta l}_{ij}^{\rm{rpl}}. b) when vih\bm{v}_{i}^{h} intersects into the object, both the repulsion and the attraction will push vih\bm{v}_{i}^{h} towards the surface; when vih\bm{v}_{i}^{h} is outside the object, the attraction and repulsion will point to opposite directions, leading to a balance point outside but in the vicinity to the object’s surface. We provide an intuitive illustration in Fig. 12.

B.2 Anchor Elasticity Assignment

As discussed in main text §4 (Annotation of the Attractive Springs), we treat the elasticity of the attractive spring as the network prediction. Here, we shall provide the annotation heuristics of the attractive spring k^atr\hat{k}^{\rm{atr}} First, we set the anchor ai\bm{a}_{i} - vertex vjo\bm{v}^{o}_{j} pair with ground-truth distance ∣Δl^ijatr∣>20mm|\bm{\Delta}\hat{\bm{l}}^{\rm{atr}}_{ij}|>20mm as invalid contact and has k^ijatr=0\hat{k}^{\rm{atr}}_{ij}=0. Second, for those anchor-vertex pairs within the distance threshold 20 mmmm, an inverse-proportional k^ijatr\hat{k}^{\rm{atr}}_{ij} is assigned according to the ∣Δl^ijatr∣|\bm{\Delta}\hat{\bm{l}}^{\rm{atr}}_{ij}|:

To note, we do not have a strict requirement on the function of k^ijatr\hat{k}^{\rm{atr}}_{ij}. Any other functions should also work when satisfying: a) k=1k=1when ∣Δl∣=0|\bm{\Delta l}|=0; b) kkis inverse proportional to ∣Δl∣|\bm{\Delta l}| in the range of 0 to 20 mmmm; c) kkis bounded by 0 and 1. The choice of cosine function is simply due to its smoothness.

Appendix C HO3D Dataset

As we mentioned in the main text §6.1, several samples in the HO3D testing set do not suit for evaluating MIHO. Firstly, since GeO requires the predicted 6D pose of the known objects, all the grasps of the pitcher have to be removed. Secondly, many interactions of hand and objects in the testing set are not stable. For example, sliding the palm over the surface of a bleach cleanser bottle, may cause a strange contact and mislead the optimization in GeO. Therefore, we only select the grasps that can pick up the objects firmly. We show several unsuitable samples in Fig. 13. Table.4 shows our final selection on HO3Dv2 test set, as we called HO3Dv2-.

C.2 Data Augmentation

We augment the training sample in HO3Dv1 in terms of poses and grasps. a) To generate more poses, we firstly randomize a disturbance transformation to the hand and object poses in the object canonical coordinate system. Then, we apply the disturbance on the hand and object meshes and render these meshes to image by a given camera intrinsic. b) To generate more grasps, we fit more stable grasps around the object. Specifically, as we show in Fig. 14, the generation procedure is achieved by 2 steps: 1) Manually move the hand around the tightest bounding cuboid of the object. 2) Refine the hand pose in the proposed GeO. Since the attractive springs in CPF are unavailable here, we replace the attraction energy in main text Eq. 3 with the LA\mathcal{L}_{A} in Eq. 4, and retain the repulsion energy and the anatomical cost. The optimization process of grasping generation can be expressed as:

Appendix D Experiments and Results

In this section we provide more implementation details about the HoNet, PiCR, and GeO module.

The HoNet module employs ResNet-18 backbone initialized with ImageNet pretrained weights. For FHB and HO3Dv2 dataset, we use the pretrained weights released from . For the HO3Dv1 dataset, we train the HoNet with Adam solver and a constant learning rate of 5×10−45\times 10^{-4} in total 200 epochs.

PiCR.

The PiCR module employs a Stacked Hourglass Networks (with 2 stacks) as backbone, a PointNet as the point encoder, and three multi-layer perceptrons as heads. The image features yield from the two hourglass stacks are gathered together and sequentially fed into the PointNet encoder and three heads. While the loss is computed over the sum of two rounds prediction, both PointNet encoder and the three heads have only one instance throughout PiCR module. At the evaluation stage, we only use the image features from the last hourglass stack to get the prediction from three heads.

We train the PiCR module with two stages. 1) Pretraining.We pretrain the PiCR module with the input image and the ground-truth object mesh in camera space. The ground-truth object mesh are disturbed by a minor rotation and translation shift. We employ Adam solver with an initial learning rate of 1×10−31\times 10^{-3}, decaying 50% every 100 epochs. The total epochs during pretraining stage is 200. 2) Fine-tuning.At the fine-tuning stage, we feed PiCR module with the object vertices predicted from HoNet. The HoNet’s weights is freezed during PiCR fine-tuning. We employ Adam solver and set the initial learning rate in fine-tuning stage as 5×10−45\times 10^{-4}, decayed to 50% every 100 epochs, and finished at 200 epochs. In both stages, we set the training mini-batch size to 8 per GPU, and a total of 4 GPUs are used.

GeO.

The GeO is a fitting module based on the non-linear optimization. For each sample, we minimize the cost function in 400 iterations, with a initial learning rate of 1×10−21\times 10^{-2}, reduced on plateau that the cost function has stopped decaying in 20 consecutive iterations. We implement GeO in PyTorch thanks for its auto derivative, and an Adam solver is employed when updating the arguments. To note, GeO can also support any other optimization toolbox.

D.2 Ablation Study

As referred in main text §6.4 (Ablation Study), this section contains another three ablation studies. all the following experiments are under the hand-object setting.

While the elasticity katrk^{\rm{atr}} of the attractive springs are predicted in PiCR, the elasticity krplk^{\rm{rpl}} of those repulsive strings are empirically set to 1×10−31\times 10^{-3}. In order to measure the impact of the magnitude of krplk^{\rm{rpl}} on repulsion, we test our GeO with seven experiment settings in which the krplk^{\rm{rpl}} is set to {0.2, 0.6, 1.0, 1.4, 2.0, 4.0, 8.0}×10−3\{0.2,\ 0.6,\ 1.0,\ 1.4,\ 2.0,\ 4.0,\ 8.0\}\times 10^{-3}, respectively. The experiment with krpl=1×10−3k^{\rm{rpl}}=1\times 10^{-3} is in accord with the default experiment in main text. As shown in Tab. 5, while the large krplk^{\rm{rpl}} can reduce the solid interpenetration volume, it may also push the attraction apart thus is not preferable in the reconstruction metrics: hand MPVPE and object MPVPE.

A-MANO with PCA Pose.

Since the MANO can also be driven by the PCA components of joint rotation, we further conduct experiments to demonstrate the superiority of our full MANO ( MANO with 15 relative joint rotations) over the PCA MANO (MANO with 15 PCA components of rotations). Tab. 6 shows that our full MANO can achieve a notable decrease in the hand MPVPE. We attribute this to the fact that the PCA MANO tends to recovery a hand that is inclined to the mean flat pose, while our full version imposes higher flexibility on the hand pose.

However, fitting on the 15 rotations in forms of so(3)\mathfrak{so}(3) brings 15×3=4515\times 3=45 degree of freedoms, which is less stable against pose abnormality. Hence in order to fully exploit the advantages when fitting on the rotations of 15 joints, we have to combine the anatomical constrains with it.

Unwanted Twist Correction.

In this part, we show the effectiveness when fitting the 15 rotations with anatomical constrains. We observe an unwanted twist of thumb in the ground-truth pose of HO3Dv1 testing set. As shown in Fig. 15, since A-MANO imposes constraints on the twist component of the rotation axis, it can achieve a more visually pleasing result in such case.

Appendix E More Qualitative Results

We demonstrate the qualitative results of MIHO in Fig. 16 on both the FHB and HO3D dataset . Note that the ground truth of the test set in HO3Dv2- is not available.