DeepIM: Deep Iterative Matching for 6D Pose Estimation

Yi Li, Gu Wang, Xiangyang Ji, Yu Xiang, Dieter Fox

Introduction

Localizing objects in 3D from images is important in many real world applications. For instance, in a robot manipulation task, the ability to recognize the 6D pose of objects, i.e., 3D location and 3D orientation of objects, provides useful information for grasp and motion planning. In a virtual reality application, 6D object pose estimation enables virtual interactions between human and objects. While several recent techniques have used depth cameras for object pose estimation, such cameras have limitations with respect to frame rate, field of view, resolution, and depth range, making it very difficult to detect small, thin, transparent, or fast moving objects. Unfortunately, RGB-only 6D object pose estimation is still a challenging problem, since the appearance of objects in the images changes according to a number of factors, such as lighting, pose variations, and occlusions between objects. Furthermore, a robust 6D pose estimation method needs to handle both textured and textureless objects.

Traditionally, the 6D pose estimation problem has been tackled by matching local features extracted from an image to features in a 3D model of the object (Lowe, 1999; Rothganger et al., 2006; Collet et al., 2011). By using the 2D-3D correspondences, the 6D pose of the object can be recovered. Unfortunately, such methods cannot handle textureless objects well since only few local features can be extracted for them. To handle textureless objects, two classes of approaches were proposed in the literature. Methods in the first class learn to estimate the 3D model coordinates of pixels or keypoints of the object in the input image. In this way, the 2D-3D correspondences are established for 6D pose estimation (Brachmann et al., 2014; Rad and Lepetit, 2017; Tekin et al., 2017). Methods in the second class convert the 6D pose estimation problem into a pose classification problem by discretizing the pose space (Hinterstoisser et al., 2012b) or into a pose regression problem (Xiang et al., 2018). These methods can deal with textureless objects, but they are not able to achieve highly accurate pose estimation, since small errors in the classification or regression stage directly lead to pose mismatches. A common way to improve the pose accuracy is pose refinement: Given an initial pose estimation, a synthetic RGB image can be rendered and used to match against the target input image. Then a new pose is computed to increase the matching score. Existing methods for pose refinement use either hand-crafted image features (Tjaden et al., 2017) or matching score functions (Rad and Lepetit, 2017).

In this work, we propose DeepIM, a new refinement technique based on a deep neural network for iterative 6D pose matching. Given an initial 6D pose estimation of an object in a test image, DeepIM predicts a relative SE(3) transformation that matches a rendered view of the object against the observed image, or in other words, it predicts the relative rotation and translation that can refine the initial 6D pose estimation. By iteratively re-rendering the object based on the improved pose estimates, the two input images to the network become more and more similar, thereby enabling the network to generate more and more accurate pose estimates. Fig. 1 illustrates the iterative matching procedure of our network for pose refinement.

This work makes the following main contributions. i) We introduce a deep network for iterative, image-based pose refinement that does not require any hand-crafted image features and automatically learns an internal refinement mechanism. ii) We propose a disentangled representation of the SE(3) transformation between object poses to achieve accurate pose estimates. This representation also enables our approach to refine pose estimates of unseen objects. iii) We have conducted extensive experiments on the LINEMOD (Hinterstoisser et al., 2012b) and the Occlusion LINEMOD (Brachmann et al., 2014) datasets to evaluate the accuracy and various properties of DeepIM. These experiments show that our approach achieves large improvements over state-of-the-art RGB-only methods on both datasets. Furthermore, initial experiments demonstrate that DeepIM is able to accurately match poses for textureless objects (T-LESS (Hodan et al., 2017)) and for unseen objects (Wu et al., 2015). The rest of the paper is organized as follows. After reviewing related works in Section 2, we describe our approach for pose matching in Section 3. Experiments are presented in Section 4, and Section 5 concludes the paper.

Related work

We review representative works on 6D pose estimation in the literature.

Traditionally, object pose estimation using RGB images is tackled by matching local features (Lowe, 1999; Rothganger et al., 2006; Collet et al., 2011). In this paradigm, a 3D model of an object is first reconstructed and local features of the object are attached to the 3D model. Keypoint-based features such as SIFT (Lowe, 1999) or SURF (Bay et al., 2008) are widely used. Given an input image, local features extracted from the image are matched against features on the 3D model. By filtering out incorrect matches using robust estimation techniques such as RANSAC (Nistér, 2005), the 6D pose of the object can be recovered using the 2D-to-3D correspondences between the local features. Local-feature matching based methods can handle partial occlusions between objects as long as the features on the visual part of the object are sufficient to determine the 6D pose. However, these methods cannot handle textureless objects well, since rich texture on the object is required in order to detect these features robustly.

In contrast, template-matching based methods are capable of handling textureless objects (Jurie and Dhome, 2001; Liu et al., 2010; Gu and Ren, 2010; Hinterstoisser et al., 2012a). In this paradigm, templates of an object are first constructed, where examples of templates are renderings of the object from the 3D object model or Histogram of Oriented Gradients (HOG) (Dalal and Triggs, 2005) templates from different viewpoints. Then these templates are matched against the input image to determine the location and orientation of the target object in the input image. The drawback of template-matching based methods is that they are not robust to occlusions between objects. When the target object is heavily occluded, the matching score is usually low which may result in incorrect pose estimation.

Recent approaches apply machine learning, especially deep learning, for 6D pose estimation using RGB images (Brachmann et al., 2014; Krull et al., 2015). Learning techniques are employed to detect object keypoints for matching or learn better feature representations for pose estimation. The state-of-the-art methods (Rad and Lepetit, 2017; Kehl et al., 2017; Tekin et al., 2017; Xiang et al., 2018; Tremblay et al., 2018) augment deep learning based object detection or segmentation methods (Girshick, 2015; Long et al., 2015; Liu et al., 2016; Redmon et al., 2016) for 6D pose estimation. For example, (Rad and Lepetit, 2017; Tjaden et al., 2017; Tremblay et al., 2018) utilize deep neural networks to detect keypoints on the objects, and then compute the 6D pose by solving the PnP problem. (Kehl et al., 2017; Xiang et al., 2018) employ deep neural networks to detect objects in the input image, and then classify or regress the detected object to its pose. A recent work (Sundermeyer et al., 2018) uses an autoencoder to map the object in the image to a vector and search for the most similar vector in a pre-generated codebook for pose estimation. Overall, learning-based methods achieve better performance than traditional methods, largely due to the ability of learning a powerful feature representation for pose estimation.

2 Depth based 6D Pose Estimation

From another point of view, the 6D pose estimation problem can be tackled using depth images. Given a 3D model of an object and an input depth image, the problem is formulated as aligning the two point clouds computed from the 3D model and the depth image, respectively, which is also known as the geometric registration problem. Roughly speaking, geometric registration methods can be classified as local refinement methods and global registration methods. The most well-known local refinement algorithm is the Iterative Closest Point (ICP) algorithm (Besl and McKay, 1992) and its variants (Rusinkiewicz and Levoy, 2001; Salvi et al., 2007; Tam et al., 2013). Given an initial pose estimation, the ICP algorithm iterates between finding the correspondences between points and refining the pose estimation using the new correspondences. In general, local refinement algorithms are sensitive to the initial pose. If the initial pose estimation is not close enough, the algorithm may converge to a local mimimum.

Global registration methods (Mellado et al., 2014; Theiler et al., 2015; Zhou et al., 2016; Yang et al., 2016) solve a more challenging problem by not assuming an initial pose estimate. A common strategy is to utilize iterative model fitting frameworks such as RANSAC. In each iteration, a set of point correspondences are sampled, and an alignment is computed and evaluated using the sampled correspondences. The limitation of most global registration methods is that they are computationally expensive. Also, the registration quality heavily depends on the quality of the 3D model and the scanned point cloud. In order to improve the registration performance, features on point clouds are also introduced for matching. These include point pairs (Mian et al., 2006; Hinterstoisser et al., 2016), spin-images (Johnson and Hebert, 1999), and point-pair histograms (Rusu et al., 2009; Tombari et al., 2010). Similar to the trend in image-based matching, recent approaches (Wang et al., 2019) propose to learn point features for registration, such as applying deep neural networks to point clouds (Qi et al., 2017).

3 RGB-D based 6D Pose Estimation

When both RGB images and depth images are available, they can be combined to improve 6D pose estimation. A common strategy is to estimate an initial pose of an object based on the color image, and then refine the pose using depth-based local refinement algorithms such as ICP (Hinterstoisser et al., 2012b; Michel et al., 2017; Zeng et al., 2017).

For example, Hinterstoisser et al. (2012b) renders the 3D model of an object into templates of color images, and then matches these templates against the input image to estimate an initial pose. The final pose estimation is obtained via ICP refinement on the initial pose. Brachmann et al. (2014), Brachmann et al. (2016), Michel et al. (2017) regress each pixel on the object in the input image to the 3D coordinate of that pixel on the 3D model. When depth images are available, the 3D coordinate regression establishes correspondences between 3D scene points and 3D model points, from which the 6D pose can be computed by solving a least-squares problem. PoseCNN (Xiang et al., 2018) introduces an end-to-end neural network for 6D object pose estimation using RGB images only. Given an initial pose from the network, a customized ICP method is applied to refine the pose. A recent work (Wang et al., 2019) introduces a neural network that combines RGB images and depth images for 6D pose estimation, and an iterative pose refinement network using point clouds as input.

4 RGB vs. RGB-D

Overall, the performance of RGB-based methods is still not comparable to that of the RGB-D based methods. We believe that this performance gap is largely due to the lack of an effective pose refinement procedure using RGB images only. Manhardt et al. (2018) which is published at the same time as ours introduces a method to refine 6D object poses with only RGB images, but there is still a large performance gap between Manhardt et al. (2018) and depth-based methods. Our work is complementary to existing 6D pose estimation methods by providing a novel iterative pose matching network for pose refinement on RGB images.

The approaches most related to ours are the object pose refinement network in Rad and Lepetit (2017) and the iterative hand pose estimation approaches in Carreira et al. (2016); Oberweger et al. (2015). Compared to these techniques, our network is designed to directly regress to relative SE(3) transformations. We are able to do this due to our disentangled representation of rotation and translation and the reference frame we used for rotation, which also allows our approach to match unseen objects. As shown in Mousavian et al. (2017), the choice of reference frame is important to achieve good pose estimation results. Our work is also related to recent visual servoing methods based on deep neural networks (Saxena et al., 2017; Costante and Ciarfuglia, 2018) that estimate the relative camera pose between two image frames, while we focus on 6D pose refinement of objects. Recent works (Garon et al., 2016; Garon and Lalonde, 2017) that focus on tracking could predict the transformation of the object pose between previous frame and current frame and have the potential to be used for pose refinement.

DeepIM Framework

In this section, we describe our deep iterative matching network for 6D pose estimation. Given an observed image and an initial pose estimate of an object in the image, we design the network to directly output a relative SE(3) transformation that can be applied to the initial pose to improve the estimate. We first present our strategy of zooming in the observed image and the rendered image that are used as inputs of the network. Then we describe our network architecture for pose matching. After that, we introduce a disentangled representation of the relative SE(3) transformation and a new loss function for pose regression. Finally, we describe our procedure for training and testing the network.

It can be difficult to extract useful features for matching if objects in the input image are very small. To obtain enough details for pose matching, we zoom in the observed image and the rendered image before feeding them into the network, as shown in Fig. 2. Specifically, in the ii-th stage of the iterative matching, given a 6D pose estimate p(i−1)\mathbf{p}^{(i-1)} from the previous step, we render a synthetic image using the 3D object model viewed according to p(i−1)\mathbf{p}^{(i-1)}.

We additionally generate one foreground mask for the observed image and rendered image. The four images are cropped using an enlarged bounding box according to the observed mask and the rendered mask, where we make sure the enlarged bounding box has the same aspect ratio as the input image and is centered at the 2D projection of the origin of the 3D object model.

In more detail, given the rendered mask mrend\mathbf{m}_{\text{rend}} and the observed mask mobs\mathbf{m}_{\text{obs}}, the cropping patch is computed as

where u∗,d∗,l∗,r∗u_{*},d_{*},l_{*},r_{*} denotes the upper, lower, left, right bound of foreground mask of observed or rendered images, xc,ycx_{c},y_{c} represent the 2D projection of the center of the object in imgrend\mathbf{img}_{\text{rend}}, rr represent the aspect ratio of the origin image (width/height), λ\lambda denotes the expand ratio, which is fixed to 1.4 in the experiment in order to make the expanded patch is roughly twice than the nested one. Then this patch is bilinearly sampled to the size of the original image, which is 480×640480\times 640 in this paper. By doing so, not only the object is zoomed in without being distorted, but also the network is provided with the information about where the center of the object lies.

2 Network Structure

Fig. 3 illustrates the network architecture of DeepIM. The observed image, the rendered image, and the two masks, are concatenated into an eight-channel tensor input to the network (3 channels for observed/rendered image, 1 channel for each mask). We use the FlowNetSimple architecture from Dosovitskiy et al. (2015) as the backbone network, which is trained to predict optical flow between two images. We tried using the VGG16 image classification network (Simonyan and Zisserman, 2014) as the backbone network, but the results were very poor, confirming the intuition that a representation related to optical flow is very useful for pose matching (Wang et al., 2017).

The pose estimation branch takes the feature map after 10 convolution layers from FlowNetSimple as input. It contains two fully-connected layers each with dimension 256, followed by two additional fully-connected layers for predicting the quaternion of the 3D rotation and the 3D translation, respectively.

During training, we also add two auxiliary branches to regularize the feature representation of the network and increase training stability and performance, see Sec. 4.4 and Table. 2 for more details. One branch is trained for predicting optical flow between the rendered image and the observed image, and the other branch for predicting the foreground mask of the object in the observed image.

3 Disentangled Transformation Representation

The representation of the coordinate frames and the relative SE(3) transformation Δp\mathbf{\Delta p} between the current pose estimate and the target pose has important ramifications for the performance of the network. Ideally, we would like (1) the individual components of these transformations to be maximally dis-entangled, thereby not requiring the network to learn unnecessarily complex geometric relationships between translations and rotations, and (2) the transformations to be independent of the intrinsic camera parameters and the actual size and coordinate system of an object, thereby enabling the network to reason about changes in object appearance rather than accurate distance estimates.

The most obvious choice are camera coordinates to represent object poses and transformations. Denote the relative rotation and translation as [RΔ∣tΔ][\mathbf{R_{\Delta}}|\mathbf{t}_{\Delta}] (We denote R∗\mathbf{R}_{*} as rotation and and t∗\mathbf{t}_{*} as translation in this paper). Given a source object pose [Rsrc∣tsrc][\mathbf{R}_{\text{src}}|\mathbf{t}_{\text{src}}], the transformed target pose would be as follows:

where fxf_{x} and fyf_{y} denote the focal lengths of the camera. The scale change vzv_{z} is defined to be independent of the absolute object size or distance by using the ratio between the distances of the rendered and observed object. We use logarithm for vzv_{z} to make sure that a value of zero corresponds to no change in scale or distance. Considering the fact that fxf_{x} and fyf_{y} are constant for a specific dataset, we simply fix it to 1 in training and testing the network.

Our representation of the relative transformation has several advantages. First, rotation does not influence the estimation of translation, so that the translation no longer needs to offset the movement caused by rotation around the camera center. Second, the intermediate variables vxv_{x}, vyv_{y}, vzv_{z} represent simple translations and scale change in the image space. Third, this representation does not require any prior knowledge of the object. Using such a representation, the DeepIM network can operate independently of the actual size of the object, its internal model coordinate framework, and the camera intrinsics. It only has to learn to transform the rendered image such that it becomes more similar to the observed image.

4 Matching Loss

5 Training and Testing

In training, we assume that we have 3D object models and images annotated with ground truth 6D object poses. By adding noises to the ground truth poses as the initial poses, we can generate the required observed and rendered inputs to the network along with the pose target output that is the pose difference between the ground truth pose and the noisy pose. Then we can train the network to predict the relative transformation between the initial pose and the target pose.

During testing, we find that the iterative pose refinement can significantly improve the accuracy. To see, let p(i)\mathbf{p}^{(i)} be the pose estimate after the ii-th iteration of the network. If the initial pose estimate p(0)\mathbf{p}^{(0)} is relatively far from the correct pose, the rendered image imgrend(p(0))\mathbf{img}_{\text{rend}}(\mathbf{p}^{(0)}) may have only little viewpoint overlap with the observed image imgobs\mathbf{img}_{\text{obs}}. In such cases, it is very difficult to accurately estimate the relative pose transformation Δp(0)\mathbf{\Delta p}^{(0)} directly. This task is even harder if the network has no priori knowledge about the object to be matched. In general, it is reasonable to assume that if the network improves the pose estimate p(i+1)\mathbf{p}^{(i+1)} by updating p(i)\mathbf{p}^{(i)} with Δp(i)\mathbf{\Delta p}^{(i)} in the ii-th iteration, then the image rendered according to this new estimate, imgrend(p(i+1))\mathbf{img}_{\text{rend}}(\mathbf{p}^{(i+1)}) is also more similar to the observed image imgobs\mathbf{img}_{\text{obs}} than imgrend(p(i))\mathbf{img}_{\text{rend}}(\mathbf{p}^{(i)}) was in the previous iteration, thereby providing input that can be matched more accurately.

However, we found that, if we train the network to regress the relative pose in a single step, the estimates of the trained network do not improve over multiple iterations in testing. To generate a more realistic data distribution for training similar to testing, we perform multiple iterations during training as well. Specifically, for each training image and pose, we apply the transformation predicted from the network to the pose and use the transformed pose estimate as another training example for the network in the next iteration. By repeating this process multiple times, the training data better represents the test distribution and the trained network also achieves significantly better results during iterative testing (such an approach has also proven useful for iterative hand pose matching (Oberweger et al., 2015) and image alignment (Lin and Lucey, 2017)).

Experiments

We conduct extensive experiments on the LINEMOD dataset (Hinterstoisser et al., 2012b) and the Occlusion LINEMOD dataset (Brachmann et al., 2014) to evaluate our DeepIM framework for 6D object pose estimation. We test different properties of DeepIM and show that it surpasses other RGB-only methods by a large margin. We also show that our network can be applied to pose matching of unseen objects during training.

We use the pre-trained FlowNetSimple (Dosovitskiy et al., 2015) to initialize the weights in our network. Weights of the new layers are randomly initialized, except for the additional weights in the first conv layer that deals with the input masks and the fully-connected layer that predicts the translation, which are initialized with zeros. Other than predicting the pose transformation, the network also predicts the optical flow and the foreground mask. Including the two additional losses could slightly increase the pose estimation performance and make the training more stable. Specifically, we use the optical flow loss LflowL_{\text{flow}} as in FlowNet (Dosovitskiy et al., 2015) and the sigmoid cross-entropy loss as the mask loss LmaskL_{\text{mask}}. Two deconvolutional blocks in FlowNet are inherited to produce the feature map used for the mask and the optical flow prediction, whose spatial scale is 0.0625. Two 1×11\times 1 convolutional layers with output channel 1 (mask prediction) and 2 (flow prediction) are appended after this feature map. The predictions are then bilinearly up-sampled to the original image size (480×640480\times 640) to compute losses.

The overall loss is L=αLpose+βLflow+γLmaskL=\alpha L_{\text{pose}}+\beta L_{\text{flow}}+\gamma L_{\text{mask}}, where we use α=0.1\alpha=0.1, β=0.25\beta=0.25, γ=0.03\gamma=0.03 throughout the experiments (except some of our ablation studies). Each training batch contains 16 images. We train the network with 4 GPUs where each GPU processes 4 images. We generate 4 items for each image as described in Sec. 3.1: two images and two masks. The observed mask is randomly dilated with no more than 10 pixels to avoid over-fitting.

The Distribution of Rendered Pose during Training:

The rendered image imgrend\mathbf{img}_{\text{rend}} and mask mrend\mathbf{m}_{\text{rend}} are randomly generated during training without using prior knowledge of the initial poses in the test set. Specifically, given a ground truth pose p^\mathbf{\hat{p}}, we add noises to p^\mathbf{\hat{p}} to generate the rendered poses. For rotation, we independently add a Gaussian noise N(0,152)\mathcal{N}(0,15^{2}) to each of the three Euler angles of the rotation. If the angular distance between the new pose and the ground truth pose is more than 45°45\degree, we discard the new pose and generate another one in order to make sure the initial pose for refinement is within 45°45\degree of the ground truth pose during training. For translation, considering the fact that RGB-based pose estimation methods usually have larger standard deviation on depth estimation, the following Gaussian noises are added to the three components of the translation: Δx∼N(0,0.012),Δy∼N(0,0.012),Δz∼N(0,0.052)\Delta x\sim\mathcal{N}(0,0.01^{2}),\Delta y\sim\mathcal{N}(0,0.01^{2}),\Delta z\sim\mathcal{N}(0,0.05^{2}), where the standard deviations are 1 cm, 1 cm and 5 cm, respectively.

Synthetic Training Data:

Real training images provided in existing datasets may be highly correlated or lack images in certain situations such as occlusions between objects. Therefore, generating synthetic training data is essential to enable the network to deal with different scenarios in testing. In generating synthetic training data for the LINEMOD dataset, considering the fact that the elevation variation is limited in this dataset, we calculate the elevation range of the objects in the provided training data. Then we rotate the object model with a randomly generated quaternion and repeat it until the elevation is within this range. The translation is randomly generated using the mean and the standard deviation computed from the training set. During training, the background of the synthetic image is replaced by a randomly chosen indoor image from the PASCAL VOC dataset as shown in Fig. 6.

For the Occlusion LINEMOD dataset, multiple objects are rendered into one image in order to introduce occlusions among objects. The number of objects ranges from 3 to 8 in these synthetic images. As in the LINEMOD dataset, the quaternion of each object is also randomly generated to ensure that the elevation range is within that of training data in the Occlusion LINEMOD dataset. The translations of the objects in the same image are drawn according to the distributions of the objects in the YCB-Video dataset (Xiang et al., 2018) by adding a small Gaussian noise.

For the YCB-Video dataset, synthetic images are generated on the fly. Other than the target object, we also render another object close to it to introduce partial occlusion.

The real training images may also lack variations in light conditions exhibited in the real world or in the testing set. Therefore, we add a random light condition to each synthetic image in both the LINEMOD dataset and the Occlusion LINEMOD dataset.

2 Testing Implementation Details

The mask prediction branch and the optical flow branch are removed during testing. Since there is no ground truth segmentation of the object in testing, we use the tightest bounding box of the rendered mask mrend\mathbf{m}_{\text{rend}} instead, so the network searches the neighborhood near the estimated pose to find the target object to match. Unless specified, we use the pose estimates from PoseCNN (Xiang et al., 2018) as the initial poses. Our DeepIM network runs at 12 fps per object using an NVIDIA 1080 Ti GPU with 2 iterations during testing.

Pose Initialization during inference:

Our framework takes an input image and an initial pose estimation of an object in the image as inputs, and then refine the initial pose iteratively. In our experiments, we have tested two pose initialization methods.

The first one is PoseCNN (Xiang et al., 2018), a convolutional neural network designed for 6D object pose estimation. PoseCNN performs three tasks for 6D pose estimation, i.e., semantic labeling to classify image pixels into object classes, localizing the center of the object on the image to estimate the 3D translation of the object, and 3D rotation regression. In our experiments, we use the 6D poses from PoseCNN as initial poses for pose refinement.

To demonstrate the robustness of our framework on pose initialization, we have implemented a simple 6D pose estimation method for pose initialization, where we extend the Faster R-CNN framework designed for 2D object detection (Ren et al., 2015) to 6D pose estimation. Specifically, we use the bounding box of the object from Faster R-CNN to estimate the 3D translation of the object. The center of the bounding box is treated as the center of the object. The distance of the object is estimated by maximizing the overlap of the projection of the 3D object model with the bounding box. To estimate the 3D rotation of the object, we add a rotation regression branch to Faster R-CNN as in PoseCNN. In this way, we can obtain a 6D pose estimation for each detected object from Faster R-CNN.

In our experiments on the LINEMOD dataset described in Sec. 4.4, we have shown that, although the initial poses from Faster R-CNN are much worse than the poses from PoseCNN, our framework is still able to refine these poses using the same weights. The performance gap between using the two different pose initialization methods is quite small, which demonstrates the ability of our framework in using different methods for pose initialization.

3 Evaluation Metrics

We use the following three evaluation metrics for 6D object pose estimation. i) The 5°, 5cm metric considers an estimated pose to be correct if its rotation error is within 5° and the translation error is below 5cm. ii) The 6D Pose metric (Hinterstoisser et al., 2012b) computes the average distance between the 3D model points transformed using the estimated pose and the ground truth pose. For symmetric objects, we use the closest point distance in computing the average distance. An estimated pose is correct if the average distance is within 10% of the 3D model diameter. iii) The 2D Projection metric computes the average distance of the 3D model points projected onto the image using the estimated pose and the ground truth pose. An estimated pose is correct if the average distance is smaller than 5 pixels.

Proposed in Shotton et al. (2013). The 5°, 5cm metric considers an estimated pose to be correct if its rotation error is within 5° and the translation error is below 5cm. We also provided the results with 2°, 2cm and 10°, 10cm in Table 6 to give a comprehensive view about the performance.

For symmetric objects such as eggbox and glue in the LINEMOD dataset, we compute the rotation error and the translation error against all possible ground truth poses with respect to symmetry and accept the result when it matches one of these ground truth poses.

D Pose:

Hinterstoisser et al. (2012b) use the average distance (ADD) metric to compute the averaged distance between points transformed using the estimated pose and the ground truth pose as in Eq. 5:

where mm is the number of points on the 3D object model, M\mathcal{M} is the set of all 3D points of this model, p=[R∣t]\mathbf{p}=[\mathbf{R}|\mathbf{t}] is the ground truth pose and p^=[R^∣t^]\mathbf{\hat{p}}=[\mathbf{\hat{R}}|\mathbf{\hat{t}}] is the estimated pose. Here the number of points mm can be different from the number of points nn used in Eq. 4 as the point clouds used for training is a subset randomly sampled from the original point clouds to reduce the time to compute the loss during training. Rx+t\mathbf{R}\mathbf{x}+\mathbf{t} indicates transforming the point with the given SE(3) transformation (pose) p\mathbf{p}. Following (Brachmann et al., 2016), we compute the distance between all pairs of points from the model and regard the maximum distance as the diameter dd of this model. Then a pose estimation is considered to be correct if the computed average distance is within 10% of the model diameter. In addition to using 0.1d0.1d as the threshold, we also provided pose estimation accuracy using thresholds 0.02d0.02d and 0.05d0.05d in Table 6. We use 0.1d0.1d as the threshold of 6D Pose metric in the following paper if not specified.

For symmetric objects, we use the closest point distance in computing the average distance for 6D pose evaluation as in Hinterstoisser et al. (2012b):

In the YCB-Video Dataset, we use the metric ADD and ADD-S described in Xiang et al. (2018). After getting the ADD and ADD-S distance described in Eq. 5 and Eq. 6, we vary the threshold from 0 to 10 cm and accumulate the area under the accuracy curves.

D Projection:

focuses on the matching of pose estimation on 2D images. This metric is considered to be important for applications such as augmented reality. We compute the error using Eq. 7 and accept a pose estimation when the 2D projection error is smaller than a predefined threshold:

where K\mathbf{K} denotes the intrinsic parameter matrix of the camera and K(Rx+t)\mathbf{K}(\mathbf{R}\mathbf{x}+\mathbf{t}) indicates transforming a 3D point according to the SE(3) transformation and then projecting the transformed 3D point onto the image. In addition to using 5 pixels as the threshold, we also show our results with the thresholds 2 pixels and 10 pixels. We use 5 pixels as the threshold of Proj. 2D metric in the following paper if not specified.

For symmetric objects such as eggbox and glue in the LINEMOD dataset, we compute the 2D projection error against all possible ground truth poses and accept the result when it matches one of these ground truth poses.

4 Experiments on the LINEMOD Dataset

The LINEMOD dataset contains 15 objects. We train and test our method on 13 of them as other methods in the literature. We follow the procedure in (Brachmann et al., 2016) to split the dataset into the training and test sets, with around 200 images for each object in the training set and 1,000 images in the test set. Fig. 9 shows a subset of objects used in LINEMOD dataset. These objects are textureless and thus difficult for pose estimation methods using only local features.

For every image, we generate 10 random poses near the ground truth pose, resulting in 2,000 training samples for each object in the training set. Furthermore, we generate 10,000 synthetic images for each object where the pose distribution is similar to the real training set. For each synthetic image, we generate 1 random pose near its ground truth pose. Thus, we have a total of 12,000 training samples for each object in training. The background of a synthetic image is replaced with a randomly chosen indoor image from PASCAL VOC (Everingham et al., 2010). We train the networks for 8 epochs with initial learning rate 0.0001. The learning rate is divided by 10 after the 4th and 6th epoch, respectively.

Ablation study on iterative training and testing:

Table 1 shows the results that use different numbers of iterations during training and testing. The networks with train_iter=1train\_iter=1 and train_iter=2train\_iter=2 are trained with 32 and 16 epochs respectively to keep the total number of updates the same as train_iter=4train\_iter=4. The table shows that without iterative training (train_iter=1train\_iter=1), multiple iteration testing does not improve, potentially even making the results worse (test_iter=4test\_iter=4). We believe that the reason is due to the fact that the network is not trained with enough rendered poses close to their ground truth poses. The table also shows that one more iteration during training and testing already improves the results by a large margin. The network trained with 2 iterations and tested with 2 iterations is slightly better than the one trained with 4 iterations and tested with 4 iterations. This may be because the LINEMOD dataset is not sufficiently difficult to generate further improvements by using 3 or 4 iterations. Since it is not straightforward to determine how many iterations to use in each dataset, we use 4 iterations during training and testing in all other experiments.

Ablation study on the zoom in strategy, network structures, transformation representations, and loss functions:

Table 3 summarizes the ablation studies on various aspects of DeepIM. The “zoom” column indicates whether the network uses full images as its input or zoomed in bounding boxes up-sampled to the original image size. Comparing rows 5 and 7 shows that the higher resolution achieved via zooming in provides very significant improvements.

“Regressor”: We train the DeepIM network jointly over all objects, generating a pose transformation independent of the specific input object (labeled “shared” in “regressor” column). Alternatively, we could train a different 6D pose regressor for each individual object by using a separate fully connected layer for each object after the final FC256 layer shown in Fig. 3. This setting is labeled as “sep.” in Table 3. Comparing rows 3 and 7 shows that both approaches provide nearly indistinguishable results. But the shared network provides some efficiency gains.

“Network”: Similarly, instead of training a single network over all objects, we could train separate networks, one for each object as in Rad and Lepetit (2017). Comparing row 1 to 7 shows that a single, shared network provides better results than individual ones, which indicates that training on multiple objects can help the network learn a more general representation for matching. We also present an ablation study of mask prediction and flow prediction in Table 2. It shows that when trained with these two auxiliary branches, the network could achieve the highest performance.

“Coordinate”: This column investigates the impact of our choice of coordinate frame to reason about object transformations, as described in Fig. 5. The row labeled “camera” provides results when choosing the camera frame of reference as the representation for the object pose, rows labeled “model” move the center of rotation to the object model and choose the object model coordinate frame to reason about rotations, and the “disentangled” rows provide our disentangled approach of moving the center into the object model while keeping the camera coordinate frame for rotations. Comparing rows 2 and 3 shows that reasoning in the camera rotation frame provides slight improvements. Furthermore, it should be noted that only our “disentangled” approach is able to operate on unseen objects. Comparing rows 4 and 5 shows the large improvements our representation achieves over the common approach of reasoning fully in the camera frame of reference.

“Loss”: The traditional loss for pose estimation is specified by the distance (“Dist”) between the estimated and ground truth 6D pose coordinates, i.e., angular distance for rotation and euclidean distance for translation. Comparing rows 6 and 7 indicates that our point matching loss (“PM”) provides significantly better results especially on the 6D pose metric, which is the most important measure for reasoning in 3D space.

Application to different initial pose estimation networks:

Table 4 provides results when we initialize DeepIM with two different pose estimation networks. The first one is PoseCNN (Xiang et al., 2018), and the second one is a simple 6D pose estimation method based on Faster R-CNN (Ren et al., 2015). Specifically, we use the bounding box of the object from Faster R-CNN to estimate the 3D translation of the object. The center of the bounding box is treated as the center of the object. The distance of the object is estimated by maximizing the overlap of the projection of the 3D object model with the bounding box. To estimate the 3D rotation of the object, we add a rotation regression branch to Faster R-CNN as in PoseCNN. As we can see in Table 4, our network achieves very similar pose estimation accuracy even when initialized with the estimates from the extension of Faster R-CNN, which are not as accurate as those provided by PoseCNN (Xiang et al., 2018).

Comparison with the state-of-the-art 6D pose estimation methods:

Table 5 shows the comparison with the best color-only techniques on the LINEMOD dataset. DeepIM achieves very significant improvements over all prior methods, even those that also deploy refinement steps (BB8 (Rad and Lepetit, 2017) and SSD-6D (Kehl et al., 2017)).

Detailed Results on the LINEMOD Dataset:

Table 6 shows our detailed results on all the 13 objects in the LINEMOD dataset. The network is trained and tested with 4 iterations and 8 epochs. Initial poses are estimated by PoseCNN (Xiang et al., 2018).

5 Experiments on the Occlusion LINEMOD Dataset

The Occlusion LINEMOD dataset proposed in Brachmann et al. (2014) shares the same images used in the LINEMOD dataset (Hinterstoisser et al., 2012b), but annotated 8 objects in one video that are heavily blocked by other objects.

For every real image, we generate 10 random poses as described in Sec. 4.4. Considering the fact that most of the training data lacks occlusions, we generated about 20,000 synthetic images with multiple objects in each image. By doing so, every object has around 12,000 images which are partially occluded, and a total of 22,000 images for each object in training. We perform the same background replacement and training procedure as in the LINEMOD dataset.

Comparison with the state-of-the-art methods:

The comparison between our method and other RGB-only methods is shown in Fig. 8. We only show the plots with accuracies on the 2D Projection metric because these are the only results reported in Rad and Lepetit (2017) and (Tekin et al., 2017) (results for eggbox and glue use a symmetric version of this accuracy). It can be seen that our method greatly improves the pose accuracy generated by PoseCNN and surpasses all other RGB-only methods by a large margin. It should be noted that BB8 (Rad and Lepetit, 2017) achieves the reported results only when using ground truth bounding boxes during testing. Our method is even competitive with the results that use depth information and ICP to refine the estimates of PoseCNN. Fig. 9 shows some pose refinement results from our method on the Occlusion LINEMOD dataset.

Detailed Results on the Occlusion LINEMOD Dataset:

Table 7 shows our results on the Occlusion LINEMOD dataset. We can see that DeepIM can significantly improve the initial poses from PoseCNN. Notice that the diameter here is computed using the extents of the 3D model following the setting of (Xiang et al., 2018) and other RGB-D based methods. Some qualitative results are shown in Figure 7.

6 Experiments on the YCB-Video Dataset

The YCB-Video Dataset, which is proposed in (Xiang et al., 2018), annotates 21 YCB objects (Calli et al., 2015) in 92 video sequences (133,827 frames). It is a challenging dataset as the objects have varied sizes (diameter from 10 cm to 40 cm), different types of symmetries, and a large variety of occlusions and lighting conditions. We split the dataset as (Xiang et al., 2018), with 80 video sequences for training and 2,949 keyframes in the remaining 12 videos for testing.

As images in one video are similar to those in nearby frames, we use 1 image out of every 10 images in the training set for training. Training batches consist of captured real images from the dataset (1/8) and synthetic images which are partially occluded and generated on the fly (7/8). The network is trained with 8 epochs and we decrease the learning rate after 4 and 6 epochs. We found that with large training sets and enough epochs it was not necessary to include the flow prediction and the masks in the input, so we removed those branches and the corresponding loss from this experiment. For different categories, they share the same network but use separate regressors to achieve the best performance.

Evaluation Metric:

We follow the PoseCNN (Xiang et al., 2018) paper when evaluating the results which uses accuracy under curve of ADD (Eq. 5) and ADD-S (Eq. 6 for each object. We also report the results of ADD(-S) and AUC ADD(-S) metric which is similar to the one we used in LINEMOD (Brachmann et al., 2014). More specifically, we use ADD when the object is not symmetric and use ADD-S when the object is symmetric. Then we compute the averaged accuracy as the final result.

Symmetric Objects:

As described in Sec. 4.1, we only keep rendered poses that have an angular distance less than 45 degrees from ground truth poses during training, which means we don’t need to take special care of objects which have a symmetry angle of more than 90 degrees. However, object 024_bowl in the YCB-Video dataset is rotational symmetric. To deal with this kind of symmetry, rather than using the ground truth pose p^\mathbf{\hat{p}} provided by the dataset to compute the loss, we choose the distance to the closest pose p∗\mathbf{p}^{*} among all poses that look the same as the ground truth pose:

Here, Q\mathcal{Q} denotes the set of poses whose corresponding rendered images are the same as the one rendered using the ground truth pose. We assume that the rotation axis goes through the origin of the model frame so that no translation needs to be considered. In the experiment, we calibrate the rotation axis manually and use bisection search to locate the closest ground truth pose. Table. 8 compares networks trained with and without this strategy, showing that this training loss is useful.

Comparison with state-of-the-art methods:

Table 10 compares our results with two state-of-the-art methods: PoseCNN (Xiang et al., 2018) and DenseFusion (Wang et al., 2019). As can be seen, DeepIM greatly refines the initial pose provided by PoseCNN and is on par with those refined with ICP on many objects despite not using any depth or point cloud data. Notice that DeepIM produces low numbers on symmetric objects, like 024_bowl, under ADD metric. This is because the ADD metric cannot well represent the performance on symmetric objects as such objects have multiple correct poses but only one of these poses are labeled as the ground truth in the dataset. Table 9 shows the result compared with PoseCNN (Xiang et al., 2018) and PoseRBPF (Deng et al., 2019) using the ADD(-S) metrci which can avoid such problems. Fig. 10 visualizes some pose refinement results from our method on the YCB-Video dataset.

Tracking in the YCB-Video Dataset:

Considering the similarity between pose refinement and object tracking, it is natural to use DeepIM to track objects in videos. Therefore, we conducted an experiment testing DeepIM’s ability to track objects in the YCB-Video dataset. Provided with the ground truth pose of an object in the first frame of each video, DeepIM can perform tracking by using the refined pose estimate from the previous frame as the initial pose of the next frame. Rather than doing inference only on key frames, we applied DeepIM to all images in the test video so that the object poses were close between successive frames.

In order to determine when DeepIM loses track of an object due to heavy occlusion, we follow a simple strategy: we count the tracking as “lost” if the last iteration of the last 10 frames has an average rotation greater than 10 degrees or an average translation greater than 1 cm. Once the tracking is marked as lost, the network will be re-initialzed with PoseCNN’s prediction. This strategy is designed with the intuition that successful tracking should have a small offset at the last iteration. Re-initialization happens every 340 frames on average. Table 9 and Table 10 shows our numerical results. Notice that the results of tracking are better than PoseCNN+DeepIM in most cases and are comparable to the results refined with ICP which uses depth information. Also note that the performance on object 036_wood_block is bad because the model of the wooden block is different from the object used in the actual dataset video, which makes it nearly impossible to match the model with the image.

Tracking YCB objects in real scenes:

To demonstrate our framework’s generalization, we use our network to track objects in real scenes. This means we don’t have any prior knowledge about the lighting conditions, background, or camera parameters. Similar to tracking on the YCB-Video dataset, we use DeepIM to refine poses predicted from the previous frame. Thanks to the disentangled representation, we did not have to calibrate the camera to get its intrinsic matrix. Fig. 11 shows some tracking results using our method in the real world environment in real time.

Using Depth information:

Other than using RGB images to do pose refinement, DeepIMcan be easily extended to utilize depth information to improve its performance. Here we append the depth images of the observed image and the rendered image with the two zero-initialized additional channels in the first convolution (one for the rendered depth and the other for the observed depth). To provide the network with information of the center of the object, we normalize the depth images by subtract them from the depth of the object’s center. The results are shown in Table. 10.

Failure cases:

In Fig. 12 we show 10 instances that the network fails to refine to a correct pose. They can be grouped into 5 categories: 1) discrepancy between object models and images. This can be caused by bad light conditions or an inaccurate object model; 2) few patterns to match. This usually happens when only certain featureless side-views are visible or the object is heavily occluded; 3) objects’ shapes are unusu al and difficult to learn; 4) the initial pose is too far away from the correct pose; 5) objects with tiny key components.

7 Application to Unseen Objects and Unseen Categories

As stated in Sec. 3.3, we designed the disentangled pose representation such that it is independent of the coordinate frame and the size of a specific 3D object model. In other words, the transformation predicted from the network does not need to have prior knowledge about the model itself. Therefore, the pose transformations correspond to operations in the image space. This opens the question whether DeepIM can refine the poses of objects that are not included in the training set. From the experiment results we found that our network can perform accurate refinement on these unseen models. See Fig. 13 for example results. We also tested our framework on refining the poses of unseen object categories, where the training categories and the test categories are completely different.

In this experiment, we explore the ability of the network in refining poses of objects that has never been seen in training. ModelNet (Wu et al., 2015) contains a large number of 3D models in different object categories. Here, we tested our network on three of them: airplane, car and chair. For each of these categories, we train a network on no more than 200 3D models and test its performance on 70 unseen 3D models from the same category. Similar to the way that we generate synthetic data as described in Sec 4.1, we generate 50 poses for each model as the target poses and train the network for 4 epochs. We use uniform gray texture for each model and add a light source which has a fixed relative position to the object to reflect the norms of the object. The initial pose used in training and testing is generated in the same way as we did in previous experiments as described in Sec. 4.1. The results are show in Table 11.

Test on Unseen Categories:

We also tested our framework on refining the poses of unseen object categories, where the training categories and the test categories are completely different. We train the network on 8 categories from ModelNet (Wu et al., 2015): airplane, bed, bench, car, chair, piano, sink, toilet with 30 models in each category and 50 image pairs for each model. The network was trained with 4 iterations and 4 epochs. Then we tested the network on 7 other categories: bathtub, bookshelf, guitar, range hood, sofa, wardrobe, and tv stand. The results are shown in Table. 12. It shows that the network indeed has learned some general features for pose refinement across different object categories.

Conclusion

In this work we introduce DeepIM, a novel framework for iterative pose matching using color images only. Given an initial 6D pose estimation of an object, we have designed a new deep neural network to directly output a relative pose transformation that improves the pose estimate. The network automatically learns to match object poses during training. We introduce an disentangled pose representation that is also independent of the object size and the coordinate frame of the 3D object model. In this way, the network can even match poses of unseen objects, as shown in our experiments. Our method significantly outperforms state-of-the-art 6D pose estimation methods using color images only and provides performance close to methods that use depth images for pose refinement, such as using the iterative closest point algorithm. Example visualizations of our results on LINEMOD, ModelNet, T-LESS can be found here: https://rse-lab.cs.washington.edu/projects/deepim.

This work opens up various directions for future research. For instance, we expect that a stereo version of DeepIM could further improve pose accuracy. Furthermore, DeepIM indicates that it is possible to produce accurate 6D pose estimates using color images only, enabling the use of cameras that capture high resolution images at high frame rates with a large field of view, providing estimates useful for applications such as robot manipulation.

References