Complex In-Hand Manipulation via Compliance-Enabled Finger Gaiting and Multi-Modal Planning

Andrew S. Morgan, Kaiyu Hang, Bowen Wen, Kostas Bekris, Aaron M. Dollar

I INTRODUCTION

Within-Hand Manipulation can be characterized as the ability to reorient or reposition an object with respect to the hand frame . The quest for this capability has been studied in the robot manipulation community for decades–from planning and control with rigid hands to efforts with soft, compliant, or underactuated hands . Notably, the majority of these works constrain contacts to remain fixed or rolling during manipulation. This constraint limits the object’s workspace as the actuators can only operate on a single hand-object configuration manifold. Alternatively, finger gaiting, i.e., the process of repositioning contacts on the object during manipulation, can help alleviate such limitations and extend the object’s available workspace.

Finger gaiting is an inherently difficult task for a robot. Given a robot hand, the individual serial link fingers must work in proper unison without collision; maintaining stability while making and breaking contact with the object . The computational complexity of this problem has traditionally been very expensive–requiring the system to calculate and modulate forces and planned joint trajectories online during manipulation. We, conversely, are able to alleviate many of these complexities by leveraging compliance, i.e. safe modal transitions, in our system. The capabilities we present extend beyond what has been illustrated previously in the literature, some purely in simulation and some demonstrated on a real robot, but with support surfaces .

In this letter, we build off the observation that by leveraging the passive adaptive properties of an underactuated hand, we can convert traditional position, force, and stability control problems to a unified motion-only control paradigm, based on the compliant properties of the hand. Concretely, finger gaiting requires contacts to be constantly in motion. The presence and thus activation of certain contacts constrains what forces the hand can impart onto the object. This phenomena creates the abstraction of modes, which can be conceptualized as different families of motion manifolds that are subject to the system’s current set of constraints–typically in the form of gaining or losing contacts. Multi-modal planning within and between these constraints thus becomes a major focus of this work.

We in this work can constrain our planning approach according to the nature of the task. Formally, any orientation in SO(3)SO(3) can be achieved via a three action trajectory comprised of two orthogonal rotations. From this formulation, we design two core modal actions for the hand and a multi-modal planner that runs online and is constantly updated via a vision-based, low-latency 6D pose object tracker . This method continually calculates a trajectory from start to goal and when undesired scenarios arise, such as contact slip, the planner updates and suggests new actions accordingly. We incorporate this approach into an open-source and underactuated hand with four fingers . In the end, we showcase the efficacy of our system through various experiments: tracking the planned and executed trajectory of an object, evaluating the recovery potential given undesired perturbations, and finally, the ability to extend to novel object geometries.

The contributions in this letter are threefold. First, we develop a complete and fast planning solution for in-hand reorientation using two extrinsic rotation axes. Secondly, we describe the utility of compliance for switching between modes and how it “inflates” the contact switching region. And, finally, we present a simple yet effective robot system capable of complex finger gaiting capabilities, underscoring continued discussion in the community on the utility of compliance for in-hand manipulation tasks.

II RELATED WORK

Modeling robot manipulation has historically been arduous as the dynamics associated with contact are difficult to predict in novel scenarios. From an in-hand manipulation perspective, various levels of modeling have been investigated – from contact models and fingerpad curvature models , to hand kinematic models and whole hand-object system models . While these approaches have elucidated many powerful hand-object relationships, inaccurate parameterizations can often lead to task failure . To help alleviate such uncertainty, several end effectors leverage designs that are soft or underactuated . This property provides passive reconfigurability to the mechanism, which can absorb much of the “slack” traditionally required to be fully accounted for in system modeling and can further reduce grasp planning times . In this work, we leverage such mechanisms to unify our planning approach–creating the notion of safe regions for contact switching.

II-2 Extending In-hand Manipulation Capabilities

Models for in-hand manipulation are typically constrained to fixed or rolling contact scenarios. While this advantageously simplifies assumptions for control, it also limits the object’s available workspace to be dependent on the kinematic topology of the hand. Without relying on external contact, e.g, , or task-specific hands with roller-based fingers, e.g , two general approaches can help alleviate such constraint: sliding manipulation and finger gaiting. Leveraging the former is very difficult, as detecting and controlling its nonlinear conditions requires various levels of advanced sensing . The latter, alternatively, has been largely difficult due to computational considerations, but has been made more successful in recent years . The work in used finger gaits and tactile sensors to maintain grasp stability. used a 24-DOF hand with a motion capture system, in addition to “over 100 years” of simulated data to perform impressive and fast cube manipulation in the palm. leveraged the capabilities of a soft hand with 16 degrees of actuation for similar types of manipulation. While as impressive as these aforementioned works are, we are interested in extending beyond these capabilities to work against gravity, i.e. maintaining object stability without a support surface, and while utilizing a simple hand with very little onboard sensing.

II-3 Multi-Modal Planning for Manipulation

Problems in robot manipulation are frequently multi-modal, i.e. the seemingly continuous problems have an underlying discrete structure guided by constraints . These constraints are typically imposed by the nature of contact, and by discretizing planning in terms of these manifolds, the planner search space is confined according to physical constraints . Numerous sampling-based multi-modal planners have been described in the literature, which are able to generalize well, particularly in high-DOF scenarios. Though, sampling-based methods without informed exploration are vastly inefficient and suffer from time-complexities associated with over-exploration. Recent work has attempted to address this issue by using informed “leads” to guide exploration . In our work, we build off these observations and constrain our planner’s search according to physical properties of the SO(3)SO(3) rotation group. That is, our developed planner constrains the number and nature of mode switching to reduce planning time for continual updates during online execution.

III PRELIMINARIES

In this section, we introduce the preliminaries associated with multi-modal motion planning. We first discuss constrained manifolds and then develop notation and terminology for modal switching leveraging compliance.

III-B Safe Mode Transitions

Multi-modal planning requires a robot to transition between at least two constraint manifolds, which likely lie in different foliations, in order to reach the desired goal. Consider a robot starting in configuration qsq_{s}, located in manifold Mη1ζ1\mathcal{M}^{\zeta_{1}}_{\eta_{1}}. Let’s now assume the goal configuration qgq_{g} is located in a different mode and thus along a different manifold, i.e. Mη2ζ2\mathcal{M}^{\zeta_{2}}_{\eta_{2}} (Fig. 2(c)). In order for a robot to transition to qgq_{g}, a planner must find a candidate submanifold, Mη1∩η2\mathcal{M}_{\eta_{1}\cap\eta_{2}}, that enables a transition. Inside of this submanifold lies candidate configuration, q′q^{\prime}, for transitioning. We can define Mη1∩η2\mathcal{M}_{\eta_{1}\cap\eta_{2}} according to,

Note that it is possible for Mη1∩η2\mathcal{M}_{\eta_{1}\cap\eta_{2}} to be empty when there exists no q′q^{\prime} to transition Mη1ζ1\mathcal{M}^{\zeta_{1}}_{\eta_{1}} to Mη2ζ2\mathcal{M}^{\zeta_{2}}_{\eta_{2}}.

The success of our finger gating platform in this work largely leverages this concept. Explicitly defining ρ\rho, however, is difficult as it is system dependent, i.e. properties of the hand and the object determine how inflated the switching regions become. As a proof-of-concept analysis, we estimate ρ\rho via experimentation in Sec. VII and will leave computational investigation for future work.

IV ORIENTATION-BASED MOTION PLANNING

We formally define the problem of orientation-based motion planning and the minimum number of actions required. Then, we utilize this as a guarantee for the completeness our planner’s modal search transitions.

Proper rotations are sets of angles that can fully define any rotation in the group, SO(3)SO(3). Let’s assume we have an object with some rotation, RR. It is possible to transition RR to any rotation configuration via a combination of rotations along two orthogonal axes. Though there are six total combinations, we will instantiate our planner to focus on rotations along the z−z-axis (RzR_{z}) and x−x-axis (RxR_{x}) to serve as an example.

Theorem 1: Let the rotational configuration of the object be R∈SO(3)R\in SO(3). Then there exists at least one (ϕ,θ,ψ)∈[0,2π)×[0,π]×[0,2π)(\phi,\theta,\psi)\in[0,2\pi)\times[0,\pi]\times[0,2\pi) such that R=Rz(ϕ)Rx(θ)Rz(ψ)R=R_{z}(\phi)R_{x}(\theta)R_{z}(\psi) .

IV-B Planning Orientation Transitions

Following this realization and the preliminaries in Sec. III, we develop a complete multi-modal planning algorithm for finger gaiting that consists of three rotations along two orthogonal axes. If, for instance, all three orthogonal axes were available, the planning problem would be mostly trivial–simply transition the object along each axis as much as necessary to reach the goal. For the duration of this paper, we will refer to rotations about axes as being extrinsic, i.e. according to the hand frame of robot, and will adopt the notation that Rx(⋅)R_{x}(\cdot) and Rz(⋅)R_{z}(\cdot) are rotations of the object about the x−x- and z−z-axes of the hand frame, or roll and yaw, respectively.

As to accelerate computation, the proposed planning method is bi-directional in that it builds a search trajectory outward from both, the start and goal configurations. More formally, there exists rotations such that Rg=Rz2(ϕ)Rx(θ)Rz1(ψ)RsR_{g}=R_{z_{2}}(\phi)R_{x}(\theta)R_{z_{1}}(\psi)R_{s} where Rz1(ψ)R_{z_{1}}(\psi) is the first rotation about the z−z-axis, Rx(θ)R_{x}(\theta) is the rotation about the x−x-axis, and Rz2(ϕ)R_{z_{2}}(\phi) is the second rotation about the z−z-axis. Thus, the goal of this planner is to solve for ψ,θ,\psi,\theta, and ϕ\phi and enumerate a trajectory accordingly.

Alg. 1 begins with the backward pass by populating a z−z-axis rotation outward from RgR_{g} with step size, σR\sigma_{R}, which is user-defined (Fig. 3(a)). For our implementation, σR\sigma_{R} can generally range anywhere from 0.5-5.0°\degree and can be adaptive according to system architectures. Let’s say that these discretized rotations outward from RgR_{g} can sufficiently define the entire manifold, Mz2\mathcal{M}_{z_{2}}, i.e. an entire rotation of the z−z-axis about RgR_{g}. From these configurations in Mz2\mathcal{M}_{z_{2}}, we can build a KD-Tree, KDKD, which is a data structure that enables efficient querying of the rotations included in this manifold.

Now that the backwards direction is instantiated, the goal of the forward pass is to find rotations Rx(θ)Rz1(ψ)RsR_{x}(\theta)R_{z_{1}}(\psi)R_{s} such that they can “connect” to the manifold Mz2\mathcal{M}_{z_{2}} represented by KDKD within some distance threshold ρ\rho from (3) (Fig. 3(c)). As aforementioned, ρ\rho is the inflation factor for contact switching, and is generally system dependent. Here, we utilize a user-defined rotational distance function distR(⋅)dist_{R}(\cdot), where for our instantiation, we calculate the L2-norm of the difference vector along the roll, pitch, and yaw dimensions of two rotations. Though, we note that there are other metrics available, e.g. . Tangibly, this forward pass sequentially searches through candidate values for ψ\psi and θ\theta until the trajectory reaches the goal manifold Mz2\mathcal{M}_{z_{2}}, which is found by continually querying KDKD. Once this connection is found (Fig. 3(d)), values for ψ\psi, θ\theta, and ϕ\phi are recorded, and a trajectory for plan P\mathcal{P} is enumerated with timestamps along step size σR\sigma_{R} from start to goal.

V Object Control

Finger-gaiting is a dynamic and uncertain process; making and breaking contacts can often lead to slip which could further result in object ejection. To mitigate such occurrences, our solution is to continually replan and execute modal actions. More specifically, we rely on an online servoing-based framework so as to reach the desired goal. When the object enters a region of potential instability, such as being distant from the workspace center, a recovery phase comprised of translational modal actions is enacted to relocate the object.

Planning translations within-hand, presented in Alg. 2, performs a simple, greedy visual servoing control approach. Generally, the algorithm calculates the translational difference, γ\gamma, along each of the coordinate axes from start position, TsT_{s}, and goal position, TgT_{g}. The planner then chooses the direction of largest difference and creates a plan to take a single step in that direction .

V-2 Orientation and Translation Control

Our control approach combines both, Alg. 1 and Alg. 2. This method requires a goal threshold, τ\tau, a linear interpolant update parameter, λ\lambda, and a priori knowledge of the hand’s workspace center, TcenterT_{center}, for object recovery.

In executing Alg. 3, orientation control is accomplished first. A plan P\mathcal{P} is found according to the object’s current object rotation, RR. The robot then takes a single action along P\mathcal{P} via the robot.mode_action(⋅).mode\_action(\cdot) method. This system-specific function enacts a motor position step proportional to σR\sigma_{R} along the desired mode. From this action, an object pose displacement, δ\delta, is calculated by taking the distance between the original orientation, RR, and the new orientation, R′R^{\prime}, after the action. This value δ\delta serves as a differential update term along the orientation transition, which allows us to adaptively re-estimate σR\sigma_{R} according to an interpolant update parameter λ\lambda. Since, as aforementioned, there is inherent uncertainty in the task of finger gaiting, this adaptive update to the transition step size σR\sigma_{R} serves to aid in our plan’s accuracy for the true motion of the hand-object system. The algorithm continues while RR is outside of the goal distance threshold τR\tau_{R}. Notably, a recovery phase (lines 12-16) is developed inside of this loop and is enacted when the object position, TT, is distance τT\tau_{T} outside of TcenterT_{center}. At this point, the system takes greedy transition steps towards the center of the workspace for recovery.

Alg. 3 concludes by performing steps along P\mathcal{P} found from the TranslationPlan(⋅)TranslationPlan(\cdot) method so as to reposition the object within-hand after object reorientation. We implement this method last in an attempt to decouple orientational and translational planning for the final pose.

V-3 Generalization and Practical Algorithm Modifications

The authors present this solution as a generalized approach to planning for in-hand manipulation. Surely, the design and realization of the robot.mode_action(⋅).mode\_action(\cdot) method begs questions for system-specific scenarios, but can be found in a variety of ways: analytical modeling, dynamics learning, reinforcement learning, etc. While we do not focus on describing the underlying components of this function in this work, we largely leverage an energy-modeling technique outlined in for determining predicted object transitions.

Moreover, the proposed planner can be further accelerated by subtle variations to the aforementioned algorithms. First, KDKD only needs to be computed once since the goal pose is static. Also, initialized step sizes for σ\sigma can deviate between modes according to system-specific properties, e.g. when gravity affects the rotation for one mode more than another.

VI SYSTEM SETUP

We utilize the open-source Model Q, an underactuated hand from the Yale OpenHand Project. The hand is equipped with four two-link fingers and four total actuators. A single motor controls the actuation of an opposing set of fingers with joints comprised of flexures, guided by a differential. This enables passive reconfigurability between these two fingers and into/out-of the plane of grasping. These two coupled fingers are connected to an actuated rotary joint located within the palm, allowing 110°110\degree rotation about the z−z-axis of the hand (Fig. 4). The final two motors serve to actuate the remaining two fingers individually via tendons. Being underactuated, the hand’s joint configuration cannot be determined directly, as it is not equipped with joint encoders or tactile sensors. See for additional design information.

The hand is mounted on a 7-DOF Barrett WAM arm. External to the robot, an RGBD Intel Realsense D415 camera is calibrated to the robot’s environment. Its imagery is processed through the low-latency (60Hz) 6D pose object tracker (Fig. 4). The tracker is robust to finger occlusion, has approximately ±\pm1.5mmmm/ 2°\degree of noise along any axis, and is trained solely with synthetic object data. See for additional information.

VI-2 Mode Design

As described in Sec. IV-B, a series of rotations about two orthogonal axes allow an object to reach any orientation in SO(3)SO(3). Given the hand topology of the Model Q, an extrinsic z−z-axis rotation is easily possible via the palm’s rotary joint. Moreover, an extrinsic x−x-axis rotation is also possible via coordinated finger motions; grasping with the differentially actuated opposing pair and pushing the object from either individually driven fingers as validated by .

In addition to the two rotational motions, the hand topology furthermore allows for only two translational modes. The first can translate the object along the z−z-axis by squeezing the object towards the palm and regrasping, or by leveraging small amounts of coordinated slip, and simple y−y-axis translation is possible via a method described in (Fig. 5). Note that translation about the x−x-axis is not possible due to the actuation coupling of the rotary set of fingers. By combining these two translational modes, we are able to provide a recovery phase for the object (Fig. 1, 8), repositioning the object towards the middle of the workspace so that manipulation can safely continue. Note that x−x-axis translation is thus omitted.

VII EXPERIMENTS

The passive adaptive nature of compliant hands is particularly beneficial for the task of finger gaiting. Simply, small errors associated with state estimation, modeling, and control can likely be dismissed by the system “absorbing” the slack. It moreover determines the “inflation” of switching regions, denoted by ρ\rho. To help quantify and validate this adaptive nature, we evaluate how a regrasp action, i.e. transferring the grasp between sets of opposing fingers, affects the resultant object configuration.

Evaluated with the cube (Fig. 6), we perturb the object along each axis. Subsequently, the hand transfers the grasp to the opposite set of opposing fingers. In doing so, we record the amount of object reconfiguration along each axis via the object tracker. Notably, with the roll orientation of the object, we see a small amount of reconfiguration, but in the yaw orientation, we note the most reconfiguration (Fig. 7), due to the geometric properties of the cube. More specifically, from the start grasp (blue circle), to the final regrasped configuration (red square), the object tends to fall towards the minimum energy configuration of the system (red region) . We note that although there is a large variability in the object’s start orientation (±0.5\pm 0.5 radians), the hand is able to regrasp the object successfully in all cases. This underscores our safe modes discussion, and generally allows us to estimate how large ρ\rho can be. For safety, we define ρ=0.2\rho=0.2 radians.

VII-2 Single Trajectory Execution

We evaluate the repeatability of executing a single trajectory of the cube. Here, the goal of the task is to rotate the cube from face A to face B, and when completed, translate the object 1.8cm along the negative y−y-axis. The resultant executions, depicted in Fig. 8(a), shows the average trajectory of the object over 8 different trials. During this task, we note how closely trajectories follow along the roll and yaw dimensions, as these are controlled via our safe modes. Object recovery occurs at similar times in the trials, where the system transitions the object along the positive z−z-axis before completing the x−x-axis rotations (Fig. 8(b)). At the end of the orientation control sequence, the object is appropriately translated along the y−y-axis to reach its desired goal pose. The average trajectory data, depicted in solid black, does not extend past the 55 second mark due to one trial finishing earlier than others. The authors have, for clarity, extended the trajectory in the y−y-axis translation using a dotted line to illustrate the average final pose.

VII-3 Continuous Goal Trajectories and Robustness

We showcase the robustness of our method by evaluating both, an extended control trajectory and by disturbing the object pose during execution. As depicted in Fig. 9(a), we execute a trajectory starting on cube face F, transitioning through faces A,B,C,D,E and finally reaching the face F once again. In all cases, the object was able to reach the goal face within 8°\degree.

Moreover, we illustrate our ability to replan control trajectories when unmodeled perturbations occur (Fig. 9(b)). At different points during manipulation, we randomly perturb the object pose 13 total times along the least constrained dimensions, the x−x-axis rotation and the z−z-axis translation. Recovering and adapting to these occurrences over the course of a 205 second trajectory execution, the system completes the object reorientation task within 6°\degree of the goal face.

VII-4 Generalization to Object Geometry

Finally, we test the transferability of our method by varying object geometries (Fig. 6). In doing so, recovery properties, such as TcenterT_{center} from Alg. 3, were recalculated, and the 6D pose object tracker was retrained with new synthetic object data; other aforementioned parameters, including ρ\rho, remained the same. Depicted in Fig. 10, we evaluate simple, purely rotational trajectories about each of the objects. Notably, the sphere easily followed each of its guided reference trajectories, and was able to reach goal orientations within ±5°\pm 5\degree along any axis. The toy truck, further illustrated our capabilities, transitioning to its goal configuration in approximately 110 seconds with 4 recovery phases. And finally, the Stanford Bunny and the toy duck were the most difficult to manipulate due to their non-convex properties. With the bunny, fingers would often get stuck behind its ears while making and breaking contact. Consequently, this occurrence caused the hand-object system to get stuck in specific object orientations and required restart. Moreover, during duck manipulation, the size of the duck’s head affected the object’s center of mass. This made some actions, such as x−x-axis rotation, difficult as the motion was then more dynamic in nature. Only a simple, two-rotation trajectory could be completed with the duck, resulting in a shorter execution time. The precision of these two manipulations was decreased to ±10°\pm 10\degree along any rotation axis and were consequently not quite as robust as the other objects. We provide benchmarking baselines for our experimentation in Fig. 11 according to guidelines outlined . Please refer to the supplementary video for experimental evaluations.

VIII DISCUSSIONS AND FUTURE WORK

In this work, we developed a robust and complete solution to SO(3)SO(3) finger gait planning. While instantiated on an underactuated hand in this work, the methods described can generalize to other hand-object systems. We showcased our method with numerous tasks–highlighting the safety of our modes, the repeatability of our trajectories, the robustness of our planning and control approach, and our ability to generalize to new object scenarios. This work builds on a promising approach of vision-based control for compliant within-hand manipulation. The tasks completed in this letter extend beyond what has been possible in previous works.

Although we showcase novel capabilities, there are some limitations. First, hand design plays a crucial role in our manipulation capabilities. We conceptualize an altered hand that would enable rotation about the y−y-axis, thus shortening trajectories. While our method may have relied less on data and/or advanced sensing, we note that our manipulation actions were significantly slower than some other related works . Moreover, we discuss the “inflation” parameter ρ\rho quite extensively, which begs for further theoretical investigation. To this end, the development of geometry-focused modal actions and transitions is an interesting avenue. Lastly, the current tracking schema requires an object CAD model for training. We are interested in combining object-agnostic tracking methods together with the current object-agnostic controller for instant application to novel objects.

References