CAMS: CAnonicalized Manipulation Spaces for Category-Level Functional Hand-Object Manipulation Synthesis

Juntian Zheng, Qingyuan Zheng, Lixing Fang, Yun Liu, Li Yi

Introduction

Human conducts hand-object manipulation (HOM) for certain functional purposes commonly in daily life, e.g. opening a laptop and using scissors to cut. Understanding how such manipulation happens and being able to synthesize realistic hand-object manipulation has naturally become a key problem in computer vision. A generative model that can synthesize human-like functional hand-object manipulation plays an essential role in various applications, including video games, virtual reality, dexterous robotic manipulation, and human-robot interaction.

This problem has only been studied with a very limited scope previously. Most existing works focus on the synthesis of a static grasp either with or without a functional goal. Recently, there have been works started focusing on dynamic manipulation synthesis . However, these works restrict their scope to rigid objects and do not consider the fact that functional manipulation might change the object geometry as well, such as in opening a laptop by hand. Moreover, these works usually require a strong input, including hand and object trajectories or a grasp reference, limiting their application scenarios.

To expand the scope of HOM synthesis, we propose a new task of category-level functional hand-object manipulation synthesis. Given a 3D shape from a known category as well as a sequence of functional goals, our task is to synthesize human-like and physically realistic hand-object manipulation to sequentially realize the goals as shown in Figure 1. Besides rigid objects, we also consider articulated objects, which support richer manipulations than a simple move. We represent a functional goal as a 6D pose for each rigid part of the object. We emphasize category-level for generalization to unseen geometry and for more human-like manipulations revealing the underlying semantics.

In this work, we choose to tackle the above task with a learning approach. We can learn from human demonstrations for HOM synthesis thanks to the recent effort in capturing category-level human-object manipulation dataset . The key challenges lie in three aspects. First, a synthesizer needs to generalize to a diverse set of geometry with complex kinematic structures. Second, humans can interact with an object in diverse ways. Faithfully capturing such distribution and synthesizing in a similar manner is difficult. Third, physically realistic synthesis requires understanding the complex dynamics between the hand and the object. Such understanding makes sure that the synthesized hand motion indeed drives the object state change without violating basic physical rules.

To address the above challenges, we choose to generate object motion through motion planning and learn a neural synthesizer to generate dynamic hand motion accordingly. Our key idea is to canonicalize the hand pose in an object-centric and contact-centric view so that the neural synthesizer only needs to capture a compact distribution. This idea comes from the following key observations. During functional hand-object manipulation, human hands usually possess a strong preference for the contact regions, and such preference is highly correlated to the object geometry, e.g. hand grasping the display edge while opening a laptop. From the contact point’s perspective, the finger pose also lies in a low-dimensional space. Representing hand poses from an object-centric view as a set of contact points and from a contact-centric view as a set of local finger embeddings could greatly reduce the learning complexity.

Specifically, given an input object plus several functional goals, we first interpolate per-part object poses between every adjacent functional goal, resulting in an object motion trajectory. Then we take a two-stage method to synthesize the corresponding hand motion. In the first stage, we introduce CAnonicalized Manipulation Spaces (CAMS) to plan the hand motion. CAMS is defined as a two-level space hierarchy. At the root level, all corresponding parts from the category of interest are scale-normalized and consistently oriented so that the distribution of possible contact points becomes concentrated. At the leaf level, each contact point would define a local frame. This local frame would simplify the distribution of the corresponding finger pose. With CAMS, we could represent a hand pose as an object-centric and contact-centric CAMS embedding. At the core of our method is a conditional variation auto-encoder, which learns to predict a CAMS embedding sequence given an object motion trajectory. In the second stage, we introduce a contact- and penetration-aware motion synthesizer to further synthesize an object motion-compatible hand motion given the CAMS embedding sequence.

To summarize, our main contributions include: i) A new task of functional category-level hand-object manipulation synthesis. ii) CAMS, a hierarchy of spaces canonicalizing category-level HOM enabling manipulation synthesis for unseen objects. iii) A two-stage motion synthesis method to synthesize human-like and physically realistic HOM. iv) State-of-the-art HOM synthesis results for both articulated and rigid object categories.

Related Work

Human motion synthesis, including motion prediction, interpolation, and completion, has attracted many interests these years . Conditional variational autoencoder (CVAE) was widely used for its generalizability across various scenes and human motions. These works have achieved great success in human motion synthesis, while their potential in HOM synthesis was overlooked. In particular, focused on modeling human-scene interaction when generating human motion. Such scene-aware or context-aware methods might also suit HOM synthesis.

2 Physics-Based Object Manipulation Synthesis

Generating high-quality human grasps remains challenging due to the complex geometry and complicated skeletal constraints. Physics-based methods were favored for in-hand manipulation synthesis since the generated motion was physically plausible. IBS presented a novel representation of hand-object interaction and leveraged reinforcement learning (RL) methods with execution success, and geometric measure rewards to generate successful grasping motion. D-Grasp developed its grasping policy based on the physical attributes of hand and object, including angles and velocities. Yang et al. concentrated on using chopsticks in diverse gripping styles, and it solved this rather difficult task by first optimizing physically valid gripping poses with predefined gripping styles and then utilizing the carefully designed hand-controlled policies to synthesize manipulation.

3 Data-Driven Object Manipulation Synthesis

Besides physics-based methods, a line of data-driven approaches could generate manipulation in a more natural and human-like manner and generalize to novel object instances. Most of the previous data-driven works focused on reconstruction and synthesis for the static grasp. Grasping Field and CPF reconstructed a static grasping field in 3D space for hand-object interaction based on an RGB image. GraspTTA predicted a static joint-angle configuration of the grasping hand from a given object point cloud and contributed a Test Time Adaptation (TTA) strategy for helping the method generalize to novel objects. Going beyond static grasping synthesis, ManipNet proposed several geometric sensors and managed to generate long-term complex manipulation sequences. TOCH achieves a data-driven approach of dynamic motion refinement in hand object motion synthesis.

Problem Formulation and Notations

In this paper, we focus on synthesizing functional manipulation for a specific task goal defined on a known category of articulated or rigid objects (e.g. opening a laptop). The input of our system can be split into three parts:

A novel object instance from a known object category C\mathcal{C}. The object instance is described by triangular object part meshes {Mk}k=1N\{\mathbf{M}_{k}\}_{k=1}^{N}, where N>1N>1 indicates an articulated object and N=1N=1 indicates a rigid object. We assume that the number of rigid parts NN for each object category is known and constant.

A goal sequence G={(Sj→Sj+1,tj)}j=0M\mathbf{G}=\{(\mathbf{S}^{j}\to\mathbf{S}^{j+1},\mathbf{t}^{j})\}_{j=0}^{M} that defines the goal of a manipulation task by the movement of object’s parts, in which tjt^{j} denotes the lasting time of the jj-th stage, and Sj={Skj∈SE(3)}k=1N\mathbf{S}^{j}=\{\mathbf{S}^{j}_{k}\in SE(3)\}_{k=1}^{N} denotes the 6D poses of each rigid part at the jj-th stage transition point. The whole manipulation procedure is subdivided into multiple stages according to the object’s state of movement (e.g. A goal sequence of the opening task for laptops consists of two stages: approaching and opening, as shown in Figure 1).

Method

Figure 2 shows the framework of our system. We first introduce CAnonicalized Manipulation Spaces (CAMS), a two-level space hierarchy that allows representing dynamic hand motions in an object-centric and contact-centric view (see Section 4.1). By embedding hand motions in CAMS, we obtain a more compact description for dynamic manipulations, which is more friendly to learning.

To learn to synthesize human-like manipulations, we propose a two-stage framework consisting of a CVAE-based planner module (see Section 4.2) and an optimization-based synthesizer module (see Section 4.3). Taking the object shape O\mathbf{O}, the manipulation goal G\mathbf{G} and the initial hand configuration H0\mathbf{H}_{0} as input, we first divide the whole manipulation process into several motion stages by the sequential goals. Then, the planner module generates stage-wise contact targets and time-continuous finger embeddings as guidance for HOM synthesis. Finally, the synthesizer module takes the generated contact targets and hand embeddings as inputs and leverages an optimization strategy to produce a human-like and physically realistic HOM animation.

It is challenging to model the space of hand manipulation from the given object shape due to the shape variety within a category and the huge diversity of human manipulation styles. For two object instances that have a non-negligible geometry difference, even if we apply the same manipulation style on them (e.g. same finger placement on two pairs of scissors with different sizes), the resulting hand pose can also be significantly different. Hence it is difficult for previous approaches that directly learn the MANO parameters to generalize to novel object instances, especially when object shapes are similar during training. To address this challenge, we propose to express the motion of each finger in a canonical reference frame centered at a contact point on the object, thus associating such motion to the local contact region rather than the whole object shape.

Based on this motion expression, we introduce CAnonicalized Manipulation Spaces with two-level canonicalization for manipulation representation. At the root level, the canonicalized contact targets (see Section 4.1.1) describe the discrete contact information. At the leaf level, the canonicalized finger embedding (see Section 4.1.2) transforms finger motion from global space into local reference frames defined on the contact targets.

At the root level of CAMS, the canonicalized contact targets describe the contact information between each finger and the object. The contact targets are first used to define the contact-centric reference frames for finger embeddings (Section 4.1.2), and then allow the contact optimization of the synthesizer to achieve grasps with accurate contacts (Section 4.3).

Formally, we define the contact targets as a sequence

1.2 Canonicalized Finger Embeddings

Given contact points Pi,j,k\mathbf{P}_{i,j,k} from contact targets, we can now build contact-centric spaces for finger embedding. We denote Ri,j,k\mathbf{R}_{i,j,k} as the contact reference frame centered at Pi,j,k\mathbf{P}_{i,j,k}, with orientation aligned to the corresponding rigid part of the object (as shown in 3). For a static hand pose represented by MANO parameters θ\theta and β\beta, we first calculate the corresponding MANO joint coordinates Jtip,Jdip,Jpip,Jmcp\mathbf{J}^{tip},\mathbf{J}^{dip},\mathbf{J}^{pip},\mathbf{J}^{mcp} of finger ii and Jroot\mathbf{J}^{root} for the hand wrist under reference frame Ri,j,k\mathbf{R}_{i,j,k}. Based on joint positions, we further compute the normalized directions

2 CAMS-CVAE: the Motion Planner

Given an object instance with task configuration, we present CAMS-CVAE, a CVAE-based generative motion planner for generating CAMS sequences, including contact targets and continuous-time finger embeddings. An overview of the model architecture is shown in Figure 4.

For a detailed calculation of the loss terms, please refer to our released code.

3 Optimization-Based Motion Synthesizer

Once a CAMS embedding sequence has been generated from the planner, our optimization-based synthesizer subsequently produces a complete HOM sequence. We simply use Bézier curve-based interpolation to generate the object trajectory and thus focus on synthesizing a hand trajectory that satisfies our goal (Section 3). Given the object shape and trajectory, as well as a CAMS embedding sequence, our synthesizer adopts a two-stage optimization method that first optimizes the MANO pose parameters to best fit the CAMS finger embedding (see Section 4.3.1) and then optimizes the contact effect to improve physical plausibility (see Section 4.3.2). In each frame, the synthesizer reads a binary flag fn\mathbf{f}_{n} indicating whether the finger is near enough (10cm) that the finger embedding will be used to guide the generated motion, and a binary flag fc\mathbf{f}_{c} indicating whether there is a contact between the finger and object.

Given the predicted contact targets C\mathbf{C}, the bidirectional finger embedding F1,F2\mathbf{F}_{1},\mathbf{F}_{2} and the binary flags fc,fn\mathbf{f}_{c},\mathbf{f}_{n}, we optimize the MANO parameters of hand pose and transition θ={θt}t=1T\theta=\{\theta_{t}\}_{t=1}^{T} by minimizing:

The first term of Eq.(2) is the tip transition loss, which constrains the tip position Jtip\mathbf{J}^{tip} of the finger in the canonicalized finger embedding. The second term of Eq.(2) is the joint orientation loss, which is used to optimize the direction vectors of four subsequent joints Ddip\mathbf{D}^{dip}, Dpip\mathbf{D}^{pip}, Dmcp\mathbf{D}^{mcp}, Droot\mathbf{D}^{root} of the finger in the canonicalized finger embedding. And the last term of (2) is a smoothness loss for improving temporal continuity.

3.2 Optimizing Contact and Penetration

After fitting MANO parameters θ\theta by finger embedding, we leverage another optimization-based method to handle the penetration and inaccurate contact issues of the hand pose. The optimization course refines θ\theta to physically realistic MANO parameters θ′\theta^{\prime} as the final result of synthesis.

We minimize the overall loss value defined as

To produce more accurate contacts, we iteratively conduct such an optimization process for several epochs. In each epoch we feed the current θ′\theta^{\prime} to the optimization course for further improvement.

Experiments

In this section, we apply our method to synthesize HOM, and evaluate the effect of our method with various metrics. We first introduce the experimental settings including dataset (Section 5.1), baselines (Section 5.3), and evaluation metrics (Section 5.2), while the experimental details are presented in our supplementary material. We then show that our method could handle shape diversity and manipulation diversity and achieve a surpassing synthesis performance compared with other methods in Section 5.5.

We utilize HOI4D Dataset in our experiment. HOI4D is a real-world dataset that contains dynamic HOM data spanning various rigid and articulated object categories. We select five object categories with different functionalities in our experiment, as shown in (Tab. 1). The selected HOM data contains both various manipulation types (e.g. opening a laptop with different human preferences) and complex manipulation processes (e.g. opening a small scissor by putting fingers through the hole) to improve the diversity and complexity of our synthesis results. To improve the usability of HOI4D under our setting, we applied some data cleaning and augmentation methods to the raw data (see our supplementary material for more details).

2 Evaluation Metrics

We report several evaluation metrics on our experiments. We use these metrics to quantize the human-likeness and physical plausibility of motion synthesis results.

Contact-Movement Consistency We evaluate whether the object’s movement can align with the contact forces produced by hand-object contacts, using the same physics model as in ManipNet. For articulated objects, we assume there are free opposite forces at the spin axis of the object. We calculate the proportion of frames that the contacts align with the object motion.

Articulation Consistency For articulated objects, we evaluate whether the hand pose can manipulate the object in a human-like manner. In particular, for each part of the object in each frame, we compute the torque of all contact points w.r.t. the object’s spin axis if a unit force is applied along the normal direction of the contact point. If the maximal attitude of the torque with the same direction of object rotation exceeds a threshold, we regard such frame as qualified. We calculate the proportion of qualified frames.

Penetration Rate We compute the mean penetration proportion of hand vertices for each sequence. We regard penetrations within a small threshold λ\lambda=5mm as not penetrated since small penetrations can be seen as contacts made by a soft hand in real.

Perceptual Score We collect human perceptual scores to judge the naturalness of the motion sequences. We ask people not familiar with motion synthesis to give discrete perceptual scores for the results and calculate a mean score for each category using each baseline method. The detailed approach is left to the supplementary material.

3 Baselines

As introduced in Section 2, there are only a few works about dynamic HOM generation using a learning-based method, while there are various techniques that could have the potential to be used in this task. Consequently, we design the following baselines.

GraspTTA: GraspTTA proposed a strong baseline in static grasp generation. We use it to generate static grasps of several key snapshots in manipulation and then refine the result at test time using the TTA loss (We do not refine the network at test time). The dynamic HOM animation is thus generated from interpolation based on these snapshots.

ManipNet: Benefiting from carefully designed geometric sensors, ManipNet has shown a strong ability to generalization on generic object manipulation synthesis tasks. Different from our setting, ManipNet assumes additional input of wrist trajectory. To compare it with other methods, we provide it with the wrist trajectory generated by CAMS as input (and thus it only differs at fingers).

Though both GraspTTA and ManipNet have their own mechanisms to improve synthesis quality (e.g. reduce penetration), they can also benefit from our optimization-based motion synthesizer (optimizing Lpenetr\mathcal{L}_{penetr} to reduce penetration). We combine all baselines with an optimization stage (denoted as “w/ opt”) and compare them with the CAMS-CVAE motion planner.

4 Ablation Studies

Remove Contact Optimization: To demonstrate the advantages of contact optimizations in motion synthesis, we train a baseline model with the contact optimizations removed (CAMS-).

Besides removing the contact optimization in the synthesizer, we also did several ablation studies using different representation spaces of fingers. These experiments are left to our supplementary material.

5 Comparison

Quantitative Results Table 2 and Table 3 show the quantitative results of our method and all baselines. Our method outperforms previous work on all tasks, even if contact optimization is applied to them.

An observation is that after applying the offline optimization stage, the performance gain of our method is significantly higher than the baselines. This can be explained by that the Lcontact\mathcal{L}_{contact} term in contact optimization takes the contact targets from our planner as input, and it is the key to producing gradients guiding the finger placement. Without intermediate contact target information, the optimization can only leverage the Lpenetr\mathcal{L}_{penetr} loss term, and it may push the finger out of the object in unpredictable directions.

Qualitative Results Besides Fig. 1, Fig. 5 shows that our method can generate diverse grasp modes on a single object instance. Fig. 6 shows that our method can generate reasonable poses given complex object shapes. We also show full result demonstrations in our video, including the generated whole HOM processes for different tasks, robustness to different object shapes and sizes, comparison between baseline methods, and diversity of generated manipulation styles.

Conclusion

In this work, we tackle a novel task of category-level functional hand-object manipulation synthesis. To generate human-like and physically realistic manipulation sequences, we design a two-level space hierarchy named CAnonicalized Manipulation Spaces (CAMS) and thus present a two-stage framework containing a planner and a synthesizer that leverage CAMS as an intermediate representation. Our method achieves state-of-the-art performance for both articulated and rigid object categories.

References