TOCH: Spatio-Temporal Object-to-Hand Correspondence for Motion Refinement

Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, Gerard Pons-Moll

Introduction

Tracking hands that are in interaction with objects is an important part of many applications in Virtual and Augmented Reality, such as modeling digital humans capable of manipulation tasks . Although there exists a vast amount of literature about tracking hands in isolation, much less work has focused on joint tracking of objects and hands. The high degrees of freedom in possible hand configurations, frequent occlusions, noisy or incomplete observations (e.g lack of depth channel in RGB images) make the problem heavily ill-posed. We argue that tracking interacting hands requires a powerful prior learned from a set of clean interaction sequences, which is the core principle of our method.

Beyond the aforementioned challenges, subtle errors in hand estimation have a huge impact on perceived realism. For example, if the 3D object is floating in the air, is grasped in a non-physically plausible way, or hand and object intersect, the perceived quality will be poor. Unfortunately, such artifacts are common in pure hand-tracking methods. Researchers have used different heuristics to improve plausibility, such as inter-penetration constraints and smoothness priors . A recent line of work predicts likely static hand poses and grasps for a given object but those methods can not directly be used as a prior to fix common capturing and tracking errors. Although there exists work to refine hand-object interactions , it is only concerned with static grasps.

In this work, we propose TOCH, a data-driven method for refining noisy 3D hand-object interaction sequences. In contrast to previous work in interaction modeling, TOCH not only considers static interactions but can also be applied to sequences without introducing snapping artifacts. The whole approach is outlined in Figure 1. Our key insight is that estimating point-wise, spatio-temporal object-hand correspondences are crucial for realism, and sufficient to constrain the high-dimensional hand poses. Thus, the point-wise corresponcendes between object and hand are encoded in a novel spatio-temporal representation called TOCH field, which takes the object geometry and the configuration of the hand with respect to the object into account. We then learn the manifold of plausible TOCH fields from the recently released MoCap dataset of hand-object interactions using an auto-encoder and apply it to correcting noisy observations. In contrast to conventional binary contacts , TOCH fields also encode the position of hand parts that are not directly in contact with the object, making TOCH applicable to whole interaction sequences, see Figure 2. TOCH has further useful properties for practical application:

TOCH can effectively project implausible hand motions to the learned object-centric hand motion manifold and produces visually correct interaction sequences that outperform previous static approaches.

TOCH does not depend on specific sensor data (RGB image, depth map etc.) and can be used with any tracker.

TOCH can be used to transfer grasp sequences across objects sharing similar geometry, even though it is not designed for this task.

Related Work

Reconstructing 3D hand surfaces from RGB or depth observations is a well-studied problem . Existing work can generally be classified into two paradigms: discriminative approaches directly estimate hand shape and pose parameters from the observation, while generative approaches iteratively optimize a parametric hand model so that its projection matches the observation. Recently, more challenging settings such as reconstructing two interacting hands are also explored. These works ignore the presence of objects and are hence less reliable in interaction-intensive scenarios.

Joint Hand and Object Reconstruction. Jointly reconstructing hand and object in interaction has received much attention. Owing to the increasing amount of hand-object interaction datasets with annotations , deep neural networks are often used to estimate an initial hypothesis for hand and object poses, which are then jointly optimized to meet certain interaction constraints . Most works in this direction improve contact realism by encouraging a small hand-to-object distance and penalizing inter-penetrating vertices . However, these simple approaches often yield implausible interaction and do not take the whole motion sequence into account. In contrast, our method alleviates both shortcomings through a object-centric, temporal representation that also considers frames in which hand and object are not in direct contact.

2 Hand Contact and Grasp

Grasp Synthesis. Synthesizing novel hand grasp given an object has been widely studied in robotics . Traditional approaches either optimize for force-closure condition or sample and rank grasp candidates based on learned features . There are also hybrid approaches that combine the merits of both . Recently, a number of neural network-based models have been proposed for this task . In particular, represent the hand-object proximity as an implicit function. We took a similar approach and represent the hand relative to the object by signed distance values distributed on the object.

Object Manipulation Synthesis. In comparison with static grasp synthesis, generating dexterous manipulation of objects is a more difficult problem since it additionally requires dynamic hand and object interaction to be modeled. This task is usually approached by optimizing hand poses to satisfy a range of contact force constraints . Hand motions generated by these works are physically plausible but lack natural variations. Zhang et al. utilized various hand-object spatial representations to learn object manipulation from data. An IK solver is used to avoid inter-penetration. We took a different approach and solely use an object-centric spatio-temporal representation, which is shown to be less prone to interaction artifacts.

Contact Refinement. Recently, some works focus on refining hand and object contact . Both and propose to first estimate the potential contact region on the object and then fit the hand to match the predicted contact. However, limited by the proposed contact representation, they can only model hand and object in stable grasp. While we share a similar goal, our work can also deal with the case where the hand is close to but not in contact with the object, as a result of our novel hand-object correspondence representation. Hence our method can be used to refine a tracking sequence.

3 Pose and Motion Prior

It has been observed that most human activities lie on low-dimensional manifolds . Therefore natural motion patterns can be found by applying learned data priors. A pose or motion prior can facilitate a range of tasks including pose estimation from images or videos , motion interpolation , motion capture , and motion synthesis . Early attempts in capturing pose and motion priors mostly use simple statistical models such as PCA , Gaussian Mixture Models or Gaussian Process Dynamical Models . With the advent of deep generative models , recent works rely on auto-encoders and adversarial discriminators to more faithfully capture the motion distribution.

Compared to body motion prior, there is less work devoted to hand motion priors. Ng et al. learned a prior of conversational hand gestures conditioned on body motion. Our work bears the most similarity to , where an object-dependent hand pose prior was learned to foster tracking. Hamer et al. proposed to map hand parts into local object coordinates and learn the object-dependent distribution with a Parzen density estimator. The prior is learned on a few objects and subsequently transferred to objects from the same class by geometric warping. Hence it cannot truly capture the complex correlation between hand gesture and object geometry.

Method

where the parameters {βi,θi,ti}\{\boldsymbol{\beta}^{i},\boldsymbol{\theta}^{i},\boldsymbol{t}^{i}\} denote shape, pose and translation w.r.t. template hand mesh Y\boldsymbol{Y} respectively.

Naively training an auto-encoder on hand meshes is problematic, because the model could ignore the conditioning object and learn a plain hand motion prior (Sec. 4.5). Moreover, if we include the object into the formulation, the model would need to learn manifolds for all joint rigid transformation of hand and object, which leads to high problem complexity . Thus, we represent the hand as a TOCH field F\boldsymbol{F}, which is a spatio-temporal object-centric representation that makes our approach invariant to joint hand and object rotation and translation.

Representation properties. The described TOCH field representation has the following advantages. 1) It is naturally invariant to joint rotation and translation of object and hand, which reduces required model complexity. 2) By specifying the distance between corresponding points, TOCH fields enable a subsequent auto-encoder to reason about point-wise proximity of hand and object. This helps to correct various artifacts, e.g. inter-penetration can be simply detected by finding object vertices with a negative correspondence distance. 3) From surface normal directions of object points and the corresponding distances, a TOCH field can be seen as an encoding of the partial hand point cloud from the perspective of the object surface. We can explicitly derive that point cloud from the TOCH field and use it to infer hand pose and shape by fitting the hand model to the point cloud (c.f. Section 3.3).

2 Temporal Denoising Auto-encoder

where F\boldsymbol{F} denotes the groundtruth TOCH fields and BCE(c^ji,cji)\text{BCE}(\hat{c}_{j}^{i},c_{j}^{i}) is the binary cross entropy between output and target correspondence indicators. Note that we only compute the first two parts of the loss on TOCH field elements with cji=1c_{j}^{i}=1, i.e. object points that have a corresponding hand point. We use a weighted loss on the distances d^ji\hat{d}_{j}^{i}. The weights are defined as

where Ni=∑j=1NcjiN_{i}=\sum_{j=1}^{N}c_{j}^{i}. This weighting scheme encourages the network to focus on regions of close interaction, where a slight error could have huge impact on contact realism. Multiplying by the sum of correspondence ensures equal influence of all points in the sequence instead of equal influence of all frames.

3 Hand Motion Reconstruction

After projecting the noisy TOCH fields of input tracking sequence to the manifold learned by the auto-encoder, we need to recover the hand motion from the processed TOCH fields. The TOCH field is not fully differentiable w.r.t. the hand parameters, as changing correspondences would involve discontinuous function steps. Thus, we cannot directly optimize the hand pose parameters to produce the target TOCH field. Instead, we decompose the optimization into two steps. We first use the denoised TOCH fields to locate hand points corresponding to the object points. We then optimize the MANO model to find hands that best fit these points, which is a differentiable formulation.

Formally, given denoised TOCH fields Fi(H,O)={(cji,dji,yji)}j=1N\boldsymbol{F}^{i}(\boldsymbol{H},\boldsymbol{O})=\{(c_{j}^{i},d_{j}^{i},\boldsymbol{y}_{j}^{i})\}_{j=1}^{N} for frames i∈{1,...,T}i\in\{1,...,T\} on object points {oj}j=1N\{\boldsymbol{o}_{j}\}_{j=1}^{N}, we first produce the partial point clouds Y^i\hat{\boldsymbol{Y}}^{i} of the hand as seen from the object’s perspective:

Then, we fit MANO to those partial point clouds by minimizing:

The first term of Equation 6 is the hand-object correspondence loss

where LBS is the linear blend skinning function in Equation 1 and ProjY(⋅)\textnormal{Proj}_{\boldsymbol{Y}}(\cdot) projects a point to the nearest point on the template hand surface. This loss term ensures that the deformed template hand point corresponding to oi\boldsymbol{o}_{i} is at a predetermined position derived from the TOCH field.

The last term of (6) regularizes shape and pose parameters of MANO,

where p¨ki\ddot{\boldsymbol{p}}_{k}^{i} is the acceleration of hand joint kk in frame ii. Besides regularizing the norm of MANO parameters, we additionally enforce temporal smoothness of hand poses. This is necessary because (7) only constrains those parts of a hand with object correspondences. Per-frame fitting of TOCH fields leads to multiple plausible solutions, which can only be disambiguated by considering neighbouring frames. Since (6) is highly nonconvex, we optimize it in two stages. In the first stage, we freeze shape and pose, and only optimize hand orientation and translation. We then jointly optimize all the variables in the second stage.

Experiments

In this section, we evaluate the presented method on synthetic and real datasets of hand/object interaction. Our goal is to verify that TOCH produces realistic interaction sequences (Section 4.3), outperforms previous static approaches in several metrics (Section 4.4), and derives a meaningful representation for hand object interaction (Section 4.5). Before presenting the results, we introduce the used datasets in Section 4.1 and the evaluated metrics in Section 4.2.

GRAB. We train TOCH on GRAB , a MoCap dataset for whole-body grasping of objects. GRAB contains interaction sequences with 5151 objects from . We pre-select 1010 objects for validation and testing, and train with the rest sequences. Since we are only interested in frames where interaction is about to take place, we filter out frames where the hand wrist is more than 1515 cm away from the object. Due to symmetry of the two hands, we anchor correspondences to the right hand and flip left hands to increase the amount of training data.

HO-3D. HO-3D is a dataset of hand-object video sequences captured by RGB-D cameras. It provides frame-wise annotations for 3D hand poses and 66D object poses, which are obtained from a novel joint optimization procedure. To ensure fair comparison with baselines which are not designed for sequences without contact, we compare on a selected subset of static frames with hand-object contact.

2 Metrics

Mean Per-Joint Position Error (MPJPE). We report the average Euclidean distance between refined and groundtruth 3D hand joints. Since pose annotation quality varies across datasets, this metric should be jointly assessed with other perceptual metrics.

Mean Per-Vertex Position Error (MPVPE). This metric represents the average Euclidean distance between refined and groundtruth 3D meshes. It assesses the reconstruction accuracy of both hand shape and pose.

Solid Intersection Volume (IV). We measure hand-object inter-penetration by voxelizing hand and object meshes and reporting the volume of voxels occupied by both. Solely considering this metric can be misleading since it does not account for the case where the object is not in effective contact with the hand.

Contact IoU (C-IoU). This metric assesses the Intersection-over-Union between the groundtruth contact map and the predicted contact map. The contact map is obtained from the binary hand-object correspondence by thresholding the correspondence distance within ±2\pm 2 mm. We only report this metric on GRAB since the groundtruth annotations in HO-3D are not accurate enough .

3 Refining Synthetic Tracking Error

In order to use TOCH in real settings, it would be ideal to train the model on the predictions of existing hand trackers. However, this requires large amount of images/depth sequences paired with accurate hand and object annotations, which is currently not available. Moreover, targeting a specific tracker might lead to overfitting to tracker-specific errors, which is undesirable for generalization.

We observe that hand errors can be decomposed into inaccurate global translation and inaccurate joint rotations, and the inaccuracies produced by most state-of-the-art trackers are small. Therefore, we propose to synthesize tracking errors by manually perturbing the groundtruth hand poses of the GRAB dataset. To this end, we apply three different types of perturbation to GRAB: translation-dominant perturbation (abbreviated GRAB-T in the table) applies an additive noise to hand translation tH\boldsymbol{t_{H}} only, pose-dominant perturbation (abbreviated GRAB-R) applies an additive noise to hand pose θ\boldsymbol{\theta} only, and balanced perturbation (abbreviated GRAB-B) uses a combination of both. We only train on the last type while evaluate on all three. The quantitative results are shown in Table 1 and qualitative results are presented in Figure 3.

We can make the following observations. First, TOCH is most effective for correcting translation-dominant perturbations of the hand. For pose-dominant perturbations where the vertex and joint errors are already very small, the resulting hands after TOCH refinement exhibit larger errors. This is because TOCH aims to improve interaction quality of a tracking sequence, which can hardly be reflected by distance based metrics such as MPJPE and MPVPE. We argue that more important metrics for interaction are intersection volume and contact IoU. As an example, the perturbation of GRAB-R (0.30.3) only induces a tiny joint position error of 4.64.6 mm, while it results in a significant 88.6%88.6\% drop in contact IoU. This validates our observation that any slight change in pose has a notable impact on physical plausibility of interaction. TOCH effectively reduces hand-object intersection as well as boosts the contact IoU even when the noise of testing data is higher than that of training data.

4 Refining RGB(D)-based Hand Estimators

To evaluate how well TOCH generalizes to real tracking errors, we test TOCH on state-of-the-art models for joint hand-object estimation from image or depth sequences. We first report comparisons with the RGB-based hand pose estimator , and two grasp refinement methods RefineNet and ContactOpt in Table 2. Hasson et al. predict hand meshes from images, while RefineNet and ContactOpt have no knowledge about visual observations and directly refine hands based on 3D inputs. Groundtruth object meshes are assumed to be given for all the methods. TOCH achieves the best score for all three metrics on HO-3D. In particular, it reduces the mesh intersection volume, indicating an improved interaction quality. We additionally evaluate TOCH on HOnnotate , a state-of-the-art RGB-D tracker which annotates the groundtruth for HO-3D. Figure 4 shows some of its failure cases and how they are corrected by TOCH.

5 Analysis and Ablation Studies

Grasp transfer. In order to demonstrate the wide-applicability of our learned features, we utilize the pre-trained TOCH auto-encoder for grasp transfer although it was not trained for this task. The goal is to transfer grasping sequences from one object to another object while maintaining plausible contacts. Specifically, given a source hand motion sequence and a source object, we extract the TOCH fields and encode them with our learned encoder network. We then simply decode using the target object – we perform a point-wise concatenation of the latent vectors with point clouds of the target object, and reconstruct TOCH fields with the decoder. This way we can transfer the TOCH fields from the source object to the target object. Qualitative examples are shown in Figure 5.

Semantic correspondence. We argue that explicitly reasoning about dense correspondence plays a key role in modeling hand-object interaction. To show this, we train another baseline model in the same manner as in Section 3, except that we adopt a simpler representation F(H,O)={(ci,di)}i=1N\boldsymbol{F}(\boldsymbol{H},\boldsymbol{O})=\{(c_{i},d_{i})\}_{i=1}^{N}, where we keep the binary indicator and signed distance without specifying which hand point is in correspondence. The loss term (7) accordingly changes from mean squared error to Chamfer distance. We can see from Table 3 that this baseline model gives the worst results in all three metrics.

Train and test on the same objects. We test the scenario where objects in test sequences are also seen at training time. We split the dataset based on the action intent label instead of by objects. Specifically, we train on sequences labelled as ’use’, ’pass’ and ’lift‘, and evaluate on the remaining. Results from Table. 3 show that generalization to different objects works slightly better than generalizing to different actions. Note that the worse results are also partly attributed to the smaller training set under this new split.

Temporal modeling. We verify the effect of temporal modeling by replacing the GRU layer with global feature aggregation. We concatenate the global average latent code with per-frame latent codes and feed the concatenated feature of each frame to a fully connected layer. As seen in Table. 3, temporal modeling with GRU largely improves interaction quality in terms of recovered contact.

Complexity and running time. The main overhead incurred by TOCH field is in computing ray-triangle intersections, the complexity of which depends on the object geometry and the specific hand-object configuration. As an illustration, it takes around 2s per frame to compute the TOCH field on 2000 sampled object points for an object mesh with 48k vertices and 96k triangles on Intel Xeon CPU. In hand-fitting stage, TOCH is significantly faster than ContactOpt since the hand-object distance can be minimized with mean squared error loss once correspondences are known. Fitting TOCH to a sequence runs at approximately 1 fps on average while it takes ContactOpt over a minute to fit a single frame.

Conclusion

In this paper, we introduced TOCH, a spatio-temporal model of hand-object interactions. Our method encodes the hand with TOCH fields, an effective novel object-centric correspondence representation which captures the spatio-temporal configurations of hand and object even before and after contact occurs. TOCH reasons about hand-object configurations beyond plain contacts, and is naturally invariant to rotation and translation. Experiments demonstrate that TOCH outperforms previous methods on the task of 3D hand-object refinement. In future work, we plan to extend TOCH to model more general human-scene interactions.

Acknowledgements This work is supported by the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A. This work is funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180 (Emmy Noether Programme, project: Real Virtual Humans). Gerard Pons-Moll is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645. The project was made possible by funding from the Carl Zeiss Foundation.

References