Learning to Estimate Pose and Shape of Hand-Held Objects from RGB Images

Mia Kokic, Danica Kragic, Jeannette Bohg

I Introduction

Observations of object poses during manipulation allow to infer contact interactions between objects and the hand. Contact interactions are central to understanding manipulation actions on a physical level. This level of understanding is crucial for applications such as object manipulation in virtual reality or when robots are learning to grasp objects from human demonstration .

While there exists a plethora of 2D images and videos of humans manipulating different objects , they often lack annotations that go beyond 2D bounding boxes and therefore do not allow a physical interpretation of manipulation actions. We aim to augment these datasets by lifting the data from 2D to 3D such that inferring physical properties of hand-object interactions becomes feasible. For this purpose, we propose an approach for estimating the shape and 6D pose of a hand-held object from an RGB image, see Fig. LABEL:fig-teaser

This problem raises a number of challenges. First, the lack of depth data in the observations makes this problem highly under-constrained. Second, there are significant occlusions of the object and/or the hand during manipulation. Third, no real-world, large-scale dataset exists that is annotated with object and hand poses during manipulation actions. To alleviate some of these challenges we propose a system that: a) can operate on real images of novel objects instances from a known category while being trained only on synthetic data and b) leverages an estimate of hand pose and configuration to facilitate estimation of the object pose and shape.

We develop a CNN termed HOPS-Net which estimates the object pose and outputs shape features used to retrieve the most similar object from a set of 3D meshes. It leverages information about the hand in two ways: i) the hand parameters form an input to the network together with the segmented object to infer its pose and shape and ii) we use a model of the human hand to refine the initial object pose estimate and generate a plausible grasp in a grasping simulator. We train the network on images of synthetic meshes held by a hand in various poses. To address the sim-to-real gap, we translate images of synthetic objects to multiple, realistic variants via image-to-image translation network. To our knowledge, this is the first application of such a network to address the sim-to-real gap for objects which for the same shape can have many different appearances.

We develop a system for 3D shape and 6D pose estimation of a hand-held objects from a single RGB image.

Our approach can deal with novel object instances from four categories while being trained only on synthetic data.

We incorporate a human hand into training and refinement which we show improves the shape and pose estimates. This holds true even if the estimate of the hand pose and configuration are noisy.

We evaluate the accuracy of our method on our dataset of realistic objects held by a hand and demonstrate the applicability on a dataset of real world egocentric activities .

II Related Work

Understanding hand-object interactions from vision is a long-standing problem that has been addressed from several angles. One line of work focuses on a hand alone, e.g., inferring hand pose and configuration from RGB or RGB-D data to either track the hand or infer manipulation actions . Another line of work follows the insight that the shape of an object influences the configuration of the hand and explores the hand in conjunction with the object . Some authors even attempt to predict forces and contacts between the hand and the object . However, scaling this approach to many objects would require large RGB or RGB-D datasets annotated with hand and object poses which is expensive to collect. Because of this, existing datasets contain either a small number of instances annotated with a 6D pose or a large number of instances with coarse annotations such as grasp types or bounding boxes, rendering them insufficient for in-depth understanding of hand-object interactions. In our work, we propose an approach for estimating both the hand and object pose from RGB which could potentially allow us to automatically annotate these datasets and use them for further analysis of hand-object interactions on a physical level.

II-B Object Shape and Pose Estimation

To estimate object shape from RGB images, most recent works adopt a metric learning approach where the features of an image are matched against those of 3D mesh models rendered under multiple viewpoints. Shape estimation is based on finding the nearest neighbor among features of the rendered meshes and retrieving it . Another approach is to map 3D voxel grids and an image into one low-dimensional embedding space. 3D model retrieval is then performed by embedding a real RGB image and finding the nearest neighbor among voxel grids . We adopt this approach since it allows us to use all the available information about the shape instead of just renderings from different viewpoints. We embed the 2D and 3D instances in the same latent space and force the 2D and 3D features of the same object to be close in this space and distant otherwise.

Research on 6D pose estimation of objects can be grouped based on two aspects: i) input modality (2D, 2.5D, 3D, multi-modal) and ii) prior object knowledge (known, unknown). When the object is known and the input is an RGB image, the main challenge is to resolve ambiguities that stem from object symmetries as well as the orientation representations (e.g., Euler angles or quaternions). State-of-the-art methods use CNNs for directly regressing position and orientation or implicitly learning pose descriptors . The latter has an advantage over direct regression since it can handle ambiguous representations and poses caused by symmetric views. The initial predictions are often refined using Iterative Closest Point (ICP) if depth data is available . In our work, we directly regress the position and orientation from RGB images but resolve ambiguities by constructing a continuous, category-specific representation of the object orientation that is invariant to the symmetries of this specific object category.

When the object is new but belongs to a known category, the biggest challenge is to develop a method that can cope with large intra-class variations and can generalize to novel instances. To do this, the datasets for training have to contain many 3D shapes annotated with pose labels. This is often done by manually aligning meshes of objects with objects in RGB images which is time-consuming and expensive. A way to address this is to train the model on synthetic data, e.g., on the renderings of meshes in various poses. Pose annotations of the resulting images are for free. The caveat is that the synthetic images is so different in appearance from real images that models trained on it typically do not generalize well to reality.

II-C Bridging the Sim-to-Real Gap

There are several strategies for making synthetic data more realistic. For example, photo-realistic rendering and domain randomization techniques augment synthetic images with different backgrounds, random lighting conditions, etc. This is often sufficient to reduce the discrepancy if the gap between synthetic and real data is already relatively small, e.g., when synthetic objects are already textured. In our work, we circumvent the problem of manually annotating the real data to train the model. We use only images of synthetic objects for which the annotations are automatically available but translate it into realistic looking images with a variant of a Generative Adversarial Network.

III Problem Statement

Our method estimates the 3D shape and 6D pose of a novel hand-held object instance from a single RGB image. We consider four categories of objects: bottle, mug, knife and bowl. We follow the insight that information about the hand helps estimation of object pose and shape. Our category-specific CNN HOPS-Net takes as input (i) an RGB image of a segmented object and (ii) the estimated hand pose and configuration. It outputs the position and orientation of the object in the camera frame as well as shape features which we use to retrieve the most similar looking mesh from the training data.

III-A2 Hand

III-A3 Object

Features are denoted by ψ(⋅)\psi(\cdot), estimated variables by ∼.

III-B System Overview

IV Pose Estimation and Shape Retrieval

The model is trained with the objective of minimizing the sum of the shape, position and orientation losses.

For estimating object shape, we adopt a metric learning approach that generates a joint embedding space for shape and image features. Our model is based on a Siamese network that is used for dimensionality reduction in weakly supervised metric learning . It consists of two twin networks that have the same structure and share weights.

For training, the network takes two data points and outputs their features which are used to compute the loss. The loss ensures that similar data points are mapped close to each other in the embedding space and distant otherwise. In our model, the first data point is a crop of the segmented object XMc\mathbf{X}_{M}^{c} as well as the hand pose and its configuration HP\mathbf{HP}. The second data point is a random mesh MV\mathbf{M_{V}}. Each data point is pre-processed to obtain features of the same dimension, and a Siamese network network outputs ψ(XMc,HP)\psi(\mathbf{X}_{M}^{c},\mathbf{HP}) and ψ(MV)\psi(\mathbf{M_{V}}), which are used to compute the shape loss LM\mathcal{L}_{\mathbf{M}}. This loss is defined as follows:

where dd is a Euclidean distance between the features ψ(XMc,HP)\psi(\mathbf{X}_{M}^{c},\mathbf{HP}) and ψ(MV)\psi(\mathbf{M_{V}}), mm is a margin as defined in and yy is a binary valued function which is equal to 11 if XMc\mathbf{X}_{M}^{c} and MV\mathbf{M_{V}} belong to the same object, and otherwise.

IV-B From a 2D Image to 6D Pose

As discussed in Sec. II, computing the orientation loss over the whole rotation matrix is challenging due to ambiguities in the representation and object symmetry. However, with a representation that is invariant to symmetry we can facilitate the learning process. Following the work of , we construct a category-specific representation R\mathbf{R} which consists of columns u\mathbf{u} of a rotation matrix and is defined depending on the type of symmetry that an object possesses, see Fig. 2.

V Data Generation

To train HOPS-Net, we need realistic RGB images of objects that are annotated with poses and shapes as well as information on the hand. In Sec. V-A we describe how we generate realistic images from textureless meshes and in Sec. V-B how we generate pose and shape labels.

A naive way to generate the data for training would be to render projections of grasped 3D meshes in various poses. However, training the network on synthetic images would not generalize well to real images due to the large visual discrepancy them. This discrepancy can be reduced by translating the synthetic images to real images and annotating them. To do this, we use a variant of CycleGAN - an image-to-image translation network which maps images from the source domain YY to the target domain XX and vice versa. The crucial benefit of CycleGAN is that it can learn from unpaired data i.e. without explicit information on which data point from YY matches a data point from XX. To constrain this problem, it uses a cycle consistency loss which ensures that the forward and backward mappings are bijective.

CycleGAN can only handle one-to-one mappings which is unsuitable for our scenario where an image of a synthetic object can be mapped to many real variants, e.g., an image of a synthetic bottle can be mapped to many real bottles of different colors and textures. Therefore, we propose to use an augmented CycleGAN which can produce one-to-many and many-to-many mappings between domains. It learns these mappings by augmenting YY and XX with latent spaces ZxZ_{x} and ZyZ_{y} respectively. The latent spaces have a Gaussian prior over their elements and as such can capture variations in both domains (see Fig. 3). We constrain the model to learn one-to-many mappings between synthetic and realistic images by learning a function i:Y×Zx↦Xi:Y\times Z_{x}\mapsto X. Cycle consistency is achieved with two losses, one for recovering YY and one for recovering ZxZ_{x}. For details see .

To train the AugCGAN, for each category we construct the source domain YY by projecting the gray, textureless 3D meshes in various poses onto a 2D image. The target domain XX is a collection of images from ImageNet that are pre-processed with Mask R-CNN to generate object segments. In both domains objects are cropped and randomly rotated. The AugCGAN outputs “real” images of synthetic objects that are used to generate a dataset for HOPS-Net. We denote an instance of the source domain with Y\mathbf{Y}, the target domain with X\mathbf{X} and the translated “real” image with X^\hat{\mathbf{X}}. The resulting model allows us to generate realistic training data for HOPS-Net without requiring laborious, manual annotations.

V-B Generating Annotations for HOPS - Net

We use the GraspIt! simulation environment to generate a dataset of hand-held objects. Thereby, the generated images X^M\hat{\mathbf{X}}_{M} and X^Mc\hat{\mathbf{X}}_{M}^{c} show the segmented object often severely occluded by the hand. Each image is annotated with hand configurations {HS\{\mathbf{HS}, HP}\mathbf{HP}\} as well as object pose and shape {p,O,M}\{\mathbf{p},\mathbf{O},\mathbf{M}\}. We refer to the dataset generated in this manner as the “real” occluded dataset.

To generate one data point, we import a mesh M\mathbf{M} into GraspIt! and render the object in pose w\mathbf{w} to yield the image YM\mathbf{Y}_{M}. We then crop the image and run it through the AugCGAN to obtain X^Mc\hat{\mathbf{X}}_{M}^{c}. We also save the coordinates of the crop center so that we can generate X^M\hat{\mathbf{X}}_{M} by positioning X^Mc\hat{\mathbf{X}}_{M}^{c} into the original image. We then import the hand parameterized with HP\mathbf{HP} which is randomly sampled from a set HP\mathcal{HP} containing configurations of stable grasps. The grasps are generated off-line with the EigenGrasp planner . We manually constrain the set by selecting grasps that satisfy two conditions: i) all fingertips are in contact with the object surface and ii) generated grasps follow a real world distribution of task-specific grasps for a certain category, e.g., mugs are grasped either from the top or from the side with the thumb and index finger close to the opening since these grasps allow lifting or pouring.

When testing on real world images, we can expect noisy hand estimates. To better account for that, we augment the dataset by adding Gaussian noise to all parameters in HP\mathbf{HP} and render an image of the resulting hand-object configuration. We then mask out the hand by computing the pixel-wise difference between the image with and without the hand. All pixels showing a difference are coloured white. This yields X^M\hat{\mathbf{X}}_{M}. We also crop this image to obtain X^Mc\hat{\mathbf{X}}_{M}^{c}. If the hand occludes less than 50%50\% of the object, we save the images as a training data point and compute HS\mathbf{HS} from HP\mathbf{HP} through forward kinematics of the human hand model. Otherwise, we proceed with the next instance.

VI Training and Inference

VI-B Inference

VI-C Pose Refinement

VII Evaluation

We quantitatively evaluate the accuracy of shape retrieval and pose estimation on the “real” occluded dataset. To evaluate how much information on the hand helps object pose estimation, we perform an ablation study for training the model with and without the hand. We show qualitative results on a subset of images “in the wild” from a dataset of manipulation actions .

To evaluate the accuracy of object pose estimation, we first compute the average pairwise distance between the mesh surface points in the ground truth and estimated pose (also known as average distance (ADD) metric ),

Although in this way we take care of some of the ambiguities in symmetric objects, we also report the results with ADD-S metric , where the average pairwise distance between the mesh surface points in the ground truth and estimated pose is computed using the closest point distance instead of the distance between corresponding points:

VII-B Quantitative Evaluation

VII-B2 Pose Estimation

For the ablation study we consider four different cases of training data: i) “real” images only, ii) “real” occluded without the HP,HS\mathbf{HP,HS}, iii) “real” occluded with the noisy HP,HS\mathbf{HP,HS} and iv) “real” occluded with the noisy HP,HS\mathbf{HP,HS} and refinement. We test it on the “real” occluded set with the HP,HS\mathbf{HP,HS}. The training and testing data are split 70:3070:30 for each category.

VII-C Qualitative Evaluation on GUN71 dataset

Since there are no applicable real-world datasets annotated with object and hand poses, we provide only qualitative results on GUN-71 dataset . We show two images per category from for which the HandNet and Mask R-CNN were successful, see Fig. 7. The results show qualitatively reasonable object poses and shapes even in the case of large occlusions and transparency. The failure cases in Fig. 8 show that when the hand configuration is not correctly estimated, the accuracy of the estimated object pose degrades. This is because HOPS-Net hinges on hand cues when there is uncertainty about the orientation, e.g., openings and bottoms of the bottles are sometimes difficult to discern and the correctly estimated hand configuration can help to resolve this.

VIII Discussion and Conclusion

We proposed an approach for 3D shape and 6D pose estimation of a novel hand-held object from an RGB image. We designed a CNN, HOPS-Net, that takes as input the object segment together with a hand pose and configuration, and outputs a position, orientation and shape features used to retrieve the most similar mesh. It uses a hand to facilitate object shape and pose estimation and furthermore to refine the initial predictions and obtain plausible grasps in a grasping simulator. The model was trained only on synthetic data which were made more realistic via an image-to-image translation network. The synthetic data also includes occlusions from the hand generated by grasping the objects with a model of the human hand in the simulator. The limitation of our method is that it hinges on existing approaches for pre-processing which are often inaccurate. In future, we plan on jointly learning to estimate hand and object. Furthermore, we plan on adding more object categories to scale the system.

References