HybridPose: 6D Object Pose Estimation under Hybrid Representations

Chen Song, Jiaru Song, Qixing Huang

Introduction

Estimating the 6D pose of an object from an RGB image is a fundamental problem in 3D vision and has diverse applications in object recognition and robot-object interaction. Advances in deep learning have led to significant breakthroughs in this problem. While early works typically formulate pose estimation as end-to-end pose classification or pose regression , recent pose estimation methods usually leverage keypoints as an intermediate representation , and align predicted 2D keypoints with ground-truth 3D keypoints. In addition to ground-truth pose labels, these methods incorporate keypoints as an intermediate supervision, facilitating smooth model training. Keypoint-based methods are built upon two assumptions: (1) a machine learning model can accurately predict 2D keypoint locations; and (2) these predictions provide sufficient constraints to regress the underlying 6D pose. Both assumptions easily break in many real-world settings. Due to object occlusions and representational limitations of the prediction network, it is often impossible to accurately predict 2D keypoint coordinates from an RGB image alone.

In this paper, we introduce HybridPose, a novel 6D pose estimation approach that leverages multiple intermediate representations to express the geometric information in the input image. In addition to keypoints, HybridPose integrates a prediction network that outputs edge vectors between adjacent keypoints. As most objects possess a (partial) reflection symmetry, HybridPose also utilizes predicted dense pixel-wise correspondences that reflect the underlying symmetric relations between pixels. Compared to a unitary representation, this hybrid representation enjoys a multitude of advantages. First, HybridPose integrates more signals in the input image: edge vectors encode spacial relations among object parts, and symmetry correspondences incorporate interior details. Second, HybridPose offers more constraints than using keypoints alone for pose regression, enabling accurate pose prediction even if a significant fraction of predicted elements are outliers (e.g., because of occlusion). Finally, it can be shown that symmetry correspondences stabilize the rotation component of pose prediction, especially along the normal direction of the reflection plane (details are provided in the supp. material).

Given the intermediate representation predicted by the first module, the second module of HybridPose performs pose regression. In particular, HybridPose employs trainable robust norms to prune outliers in predicted intermediate representation. We show how to combine pose initialization and pose refinement to maximize the quality of the resulting object pose. We also show how to train HybridPose effectively using a training set for the pose prediction module, and a validation set for the pose regression module.

We evaluate HybridPose on two popular benchmark datasets, Linemod and Occlusion Linemod . In terms of accuracy (under the ADD(-S) metric), HybridPose leads to improvements from state-of-the-art methods that merely utilize keypoints. On Occlusion Linemod , HybridPose achieves an accuracy of 47.5%, which beats DPOD , the current state-of-the-art method on this benchmark dataset.

Despite the gain in accuracy, our approach is efficient and runs at 30 frames per second on a commodity workstation. Compared to approaches which utilize sophisticated network architecture to predict one single intermediate representation (such as Pix2Pose ), HybridPose achieves better performance by using a relative simple network to predict hybrid representations.

Related Works

Intermediate representation for pose. To express the geometric information in an RGB image, a prevalent intermediate representation is keypoints, which achieves state-of-the-art performance . The corresponding pose estimation pipeline combines keypoint prediction and pose regression initialized by the PnP algorithm . Keypoint predictions are usually generated by a neural network, and previous works use different types of tensor descriptors to express 2D keypoint coordinates. A common approach represents keypoints as peaks of heatmaps , which becomes sub-optimal when keypoints are occluded, as the input image does not provide explicit visual cues for their locations. Alternative keypoint representations include vector-fields and patches . These representations allow better keypoint predictions under occlusion, and eventually lead to improvement in pose estimation accuracy. However, keypoints alone are a sparse representation of the object pose, whose potential in improving estimation accuracy is limited.

Besides keypoints, another common intermediate representation is the coordinate of every image pixel in the 3D physical world, which provides dense 2D-3D correspondences for pose alignment, and is robust under occlusion . However, regressing dense object coordinates is much more costly than keypoint prediction. They are also less accurate than keypoints due to the lack of corresponding visual cues. In addition to keypoints and pixel-wise 2D-3D correspondences, depth is another alternative intermediate representation in visual odometry settings, which can be estimated together with pose in an unsupervised manner . In practice, the accuracy of depth estimation is limited by the representational power of neural networks.

Unlike previous approaches, HybridPose combines multiple intermediate representations, and exhibits collaborative strength for pose estimation.

Multi-modal input. To address the challenges for pose estimation from a single RGB image, several works have considered inputs from multiple sensors. A popular approach is to leverage information from both RGB and depth images . In the presence of depth information, pose regression can be reformulated as the 3D point alignment problem, which is then solved by the ICP algorithm . Although HybridPose utilizes multiple intermediate representations, all intermediate representations are predicted from an RGB image alone. HybridPose handles situations in which depth information is absent.

Edge features. Edges are known to capture important image features such as object contours , salient edges , and straight line segments . Unlike these low-level image features, HybridPose leverages semantic edge vectors defined between adjacent keypoints. This representation, which captures correlations between keypoints and reveals underlying structure of object, is concise and easy to predict. Such edge vectors offer more constraints than keypoints alone for pose regressions and have clear advantages under occlusion. Our approach is similar to , which predicts directions between adjacent keypoints to link keypoints into a human skeleton. However, we predict both the direction and the magnitude of edge vectors, and use these vectors to estimate object poses.

Symmetry detection from images. Symmetry detection has received significant attention in computer vision. We refer readers to for general surveys, and for recent advances. Traditional applications of symmetry detection include face recognition , depth estimation , and 3D reconstruction . In the context of object pose estimation, people have studied symmetry from the perspective that it introduces ambiguities for pose estimation (c.f. ), since symmetric objects with different poses can have the same appearance in image. Several works have explored how to address such ambiguities, e.g., by designing loss functions that are invariant under symmetric transformations.

Robust regression. Pose estimation via intermediate representation is sensitive to outliers in predictions, which are introduced by occlusion and cluttered backgrounds . To mitigate pose error, several works assign different weights to different predicted elements in the 2D-3D alignment stage . In contrast, our approach additionally leverages robust norms to automatically filter outliers in the predicted elements.

Besides the reweighting strategy, some recent works propose to use deep learning-based refiners to boost the pose estimation performance . use point matching loss and achieve high accuracy. predicts pose updates using contour information. Unlike these works, our approach considers the critical points and the loss surface of the robust objective function, and does not involve a fixed pre-determined iteration count used in recurrent network based approaches.

Approach

As illustrated in Figure 2, HybridPose consists of a prediction module and a pose regression module.

Prediction module (Section 3.2). HybridPose utilizes three prediction networks fθKf_{\theta}^{\mathcal{K}}, fϕEf_{\phi}^{\mathcal{E}}, and fγSf_{\gamma}^{\mathcal{S}} to estimate a set of keypoints K={pk}\mathcal{K}=\{\boldsymbol{p}_{k}\}, a set of edges between keypoints E={ve}\mathcal{E}=\{\boldsymbol{v}_{e}\}, and a set of symmetry correspondences between image pixels S={(qs,1,qs,2)}\mathcal{S}=\{(\boldsymbol{q}_{s,1},\boldsymbol{q}_{s,2})\}. K\mathcal{K}, E\mathcal{E}, and S\mathcal{S} are all expressed in 2D. θ\theta, ϕ\phi, and γ\gamma are trainable parameters.

The keypoint network fθKf_{\theta}^{\mathcal{K}} employs an off-the-shelf prediction network . The other two prediction networks, fϕEf_{\phi}^{\mathcal{E}}, and fγSf_{\gamma}^{\mathcal{S}}, are introduced to stabilize pose regression when keypoint predictions are inaccurate. Specifically, fϕEf_{\phi}^{\mathcal{E}} predicts edge vectors along a pre-defined graph of keypoints, which stabilizes pose regression when keypoints are cluttered in the input image. fγSf_{\gamma}^{\mathcal{S}} predicts symmetry correspondences that reflect the underlying (partial) reflection symmetry. A key advantage of this symmetry representation is that the number of symmetry correspondences is large: every image pixel on the object has a symmetry correspondence. As a result, even with a large outlier ratio, symmetry correspondences still provide sufficient constraints for estimating the plane of reflection symmetry for regularizing the underlying pose. Moreover, symmetry correspondences incorporate more features within the interior of the underlying object than keypoints and edge vectors.

Pose regression module (Section 3.3). The second module of HybridPose optimizes the object pose (RI,tI)(R_{I},\boldsymbol{t}_{I}) to fit the output of the three prediction networks. This module combines a trainable initialization sub-module and a trainable refinement sub-module. In particular, the initialization sub-module performs SVD to solve for an initial pose in the global affine pose space. The refinement sub-module utilizes robust norms to filter out outliers in the predicted elements for accurate object pose estimation.

Training HybridPose (Section 3.4). We train HybridPose by splitting the dataset into a training set and a validation set. We use the training set to learn the prediction module, and the validation set to learn the hyper-parameters of the pose regression module. We have tried training HybridPose end-to-end using one training set. However, the difference between the prediction distributions on the training set and testing set leads to sub-optimal generalization performance.

2 Hybrid Representation

This section describes three intermediate representations used in HybridPose.

Besides outliers in predicted keypoints, another limitation of keypoint-based techniques is that when the difference (direction and distance) between adjacent keypoints characterizes important information of the object pose, inexact keypoint predictions incur large pose error.

Symmetry correspondences. The third intermediate representation consists of predicted pixel-wise symmetry correspondences that reflect the underlying reflection symmetry. In our experiments, HybridPose extends the network architecture of FlowNet 2.0 that combines a dense pixel-wise flow and the semantic mask predicted by PVNet. The resulting symmetry correspondences are given by predicted pixel-wise flow within the mask region. Compared to the first two representations, the number of symmetry correspondences is significantly larger, which provides rich constraints even for occluded objects. However, symmetry correspondences only constrain two degrees of freedom in the rotation component of the object pose (c.f. ). It is necessary to combine symmetry correspondences with other intermediate representations.

A 3D model may possess multiple reflection symmetry planes. For these models, we train HybridPose to predict symmetry correspondences with respect to the most salient reflection symmetry plane, i.e., one with the largest number of symmetry correspondences on the original 3D model.

Summary of network design. In our experiments, fθK(I)f_{\theta}^{\mathcal{K}}(I), fϕE(I)f_{\phi}^{\mathcal{E}}(I), and fγSf_{\gamma}^{\mathcal{S}} are all based on ResNet , and the implementation details are discussed in Section 4.1. Trainable parameters are shared across all except the last convolutional layer. Therefore, the overhead of introducing the edge prediction network fϕE(I)f_{\phi}^{\mathcal{E}}(I) and the symmetry prediction network fγSf_{\gamma}^{\mathcal{S}} is insignificant.

3 Pose Regression

Initialization sub-module. This sub-module leverages constraints between (RI,tI)(R_{I},\boldsymbol{t}_{I}) and predicted elements and solves (Ri,tI)(R_{i},\boldsymbol{t}_{I}) in the affine space, which are then projected to SE(3)SE(3) in an alternating optimization manner. To this end, we introduce the following difference vectors for each type of predicted elements:

Following EPnP , we compute x\boldsymbol{x} as

Refinement sub-module. Although (5) combines hybrid intermediate representations and admits good initialization, it does not directly model outliers in predicted elements. Another limitation comes from (1) and (2), which do not minimize the projection errors (i.e., with respect to keypoints and edges), which are known to be effective in keypoint-based pose estimation (c.f. ).

Benefited from having an initial object pose (Rinit,tinit)(R^{\mathit{init}},\boldsymbol{t}^{\mathit{init}}), the refinement sub-module performs local optimization to refine the object pose. We introduce two difference vectors that involve projection errors: ∀k,e,s,\forall k,e,s,

To prune outliers in the predicted elements, we consider a generalized German-Mcclure (or GM) robust function

With this setup, HybridPose solves the following non-linear optimization problem for pose refinement:

where βK\beta_{\mathcal{K}}, βE\beta_{\mathcal{E}}, and βS\beta_{\mathcal{S}} are separate hyper-parameters for keypoints, edges, and symmetry correspondences. Σk\Sigma_{k} and Σe\Sigma_{e} denote the covariance information attached to the keypoint and edge predictions. ∥x∥A=(xTAx)12\|\boldsymbol{x}\|_{A}=(\boldsymbol{x}^{T}A\boldsymbol{x})^{\frac{1}{2}}. When covariances of predictions are unavailable, we simply set Σk=Σe=I2\Sigma_{k}=\Sigma_{e}=I_{2}. The above optimization problem is solved by Gauss-Newton method starting from RinitR^{\mathit{init}} and tinit\boldsymbol{t}^{\mathit{init}}.

In the supp. material, we provide a stability analysis of (9), and show how the optimal solution of (9) changes with respect to noise in predicted representations. We also show collaborative strength among all three intermediate representations. While keypoints significantly contribute to the accuracy of t\boldsymbol{t}, edge vectors and symmetry correspondences can stablize the regression of RR.

4 HybridPose Training

This section describes how to train the prediction networks and hyper-parameters of HybridPose using a labeled dataset T={I,(KIgt,EIgt,SIgt,(RIgt,tIgt))}\mathcal{T}=\{I,(\mathcal{K}_{I}^{gt},\mathcal{E}_{I}^{gt},\mathcal{S}_{I}^{gt},(R_{I}^{gt},\boldsymbol{t}_{I}^{gt}))\}. With II, KIgt\mathcal{K}_{I}^{gt}, EIgt\mathcal{E}_{I}^{gt}, SIgt\mathcal{S}_{I}^{gt}, and (RIgt,tIgt)(R_{I}^{gt},\boldsymbol{t}_{I}^{gt}), we denote the RGB image, labeled keypoints, edges, symmetry correspondences, and ground-truth object pose, respectively. A popular strategy is to train the entire model end-to-end, e.g., using recurrent networks to model the optimization procedure and introducing loss terms on the object pose output as well as the intermediate representations. However, we found this strategy sub-optimal. The distribution of predicted elements on the training set differs from that on the testing set. Even by carefully tuning the trade-off between supervisions on predicted elements and the final object pose, the pose regression model, which fits the training data, generalizes poorly on the testing data.

Our approach randomly divides the labeled set T=Ttrain∪Tval\mathcal{T}=\mathcal{T}_{train}\cup\mathcal{T}_{val} into a training set and a validation set. Ttrain\mathcal{T}_{train} is used to train the prediction networks, and Tval\mathcal{T}_{val} trains the hyper-parameters of the pose regression model. Implementation and training details of the prediction networks are presented in Section 4.1. In the following, we focus on training the hyper-parameters using Tval\mathcal{T}_{val}.

Initialization sub-module. Let RIinitR_{I}^{\mathit{init}} and tIinit\boldsymbol{t}_{I}^{\mathit{init}} be the output of the initialization sub-module. We obtain the optimal hyper-parameters αE\alpha_{E} and αS\alpha_{S} by solving the following optimization problem:

Since the number of hyper-parameters is rather small, and the pose initialization step does not admit an explicit expression, we use the finite-difference method to compute numerical gradient, i.e., by fitting the gradient to samples of the hyper-parameters around the current solution. We then apply back-track line search for optimization.

The refinement module solves an unconstrained optimization problem, whose optimal solution is dictated by its critical points and the loss surface around the critical points. We consider two simple objectives. The first objective forces ∂fI∂c(0,β)≈0\frac{\partial f_{I}}{\partial\boldsymbol{c}}(\boldsymbol{0},\beta)\approx 0, or in other words, the ground-truth is approximately a critical point. The second objective minimizes the condition number \kappa(\frac{\partial^{2}f_{I}}{\partial^{2}\boldsymbol{c}}(\boldsymbol{0},\beta))=\lambda_{\max}\big{(}\frac{\partial^{2}f_{I}}{\partial^{2}\boldsymbol{c}}(\boldsymbol{0},\beta)\big{)}/\lambda_{\min}\big{(}\frac{\partial^{2}f_{I}}{\partial^{2}\boldsymbol{c}}(\boldsymbol{0},\beta)\big{)}. This objective regularizes the loss surface around each optimal solution, promoting a large converge radius for fI(c,β)f_{I}(\boldsymbol{c},\beta). With this setup, we formulate the following objective function to optimize β\beta:

where γ\gamma is a constant hyperparemeter. The same strategy used in (10) is then applied to optimize (11).

Experimental Evaluation

This section presents an experimental evaluation of the proposed approach. Section 4.1 describes the experimental setup. Section 4.2 quantitatively and qualitatively compares HybridPose with other 6D pose estimation methods. Section 4.3 presents an ablation study to investigate the effectiveness of symmetry correspondences, edge vectors, and the refinement sub-module.

Datasets. We consider two popular benchmark datasets that are widely used in the 6D pose estimation problem, Linemod and Occlusion Linemod . In comparsion to Linemod, Occlusion Linemod contains more examples where the objects are under occlusion. Our keypoint annotation strategy follows that of , i.e., we choose ∣K∣=8|\mathcal{K}|=8 keypoints via the farthest point sampling algorithm. Edge vectors are defined as vectors connecting each pair of keypoints. In total, each object has ∣E∣=∣K∣⋅(∣K∣−1)2=28|\mathcal{E}|=\frac{|\mathcal{K}|\cdot(|\mathcal{K}|-1)}{2}=28 edges. We further use the algorithm proposed in to annotate Linemod and Occlusion Linemod with reflection symmetry labels.

Following the convention described in , we select 15% of Linemod examples as the training data, and the rest 85% as well as all of Occlusion Linemod examples for testing. To avoid overfitting, we use the same synthetic data generation scheme introduced in PVNet . A previous version of this paper uses a different dataset split, which is inconsistent with baseline approaches. This problem has been fixed now. Please refer to our GitHub issues page for related discussions.

Implementation details. We use ResNet with pre-trained weights on ImageNet to build the prediction networks fθKf_{\theta}^{\mathcal{K}}, fϕEf_{\phi}^{\mathcal{E}}, and fγSf_{\gamma}^{\mathcal{S}}. The prediction networks take an RGB image II of size (3,H,W)(3,H,W) as input, and output a tensor of size (C,H,W)(C,H,W), where (H,W)(H,W) is the image resolution, and C=1+2∣K∣+2∣E∣+2C=1+2|\mathcal{K}|+2|\mathcal{E}|+2 is the number of channels in the output tensor.

The first channel in the output tensor is a binary segmentation mask MM. If M(x,y)=1M(x,y)=1, then (x,y)(x,y) corresponds to a pixel on the object of interest in the input image II. The segmentation mask is trained using the cross-entropy loss.

The 2∣K∣2|\mathcal{K}| channels afterwards in the output tensor give xx and yy components of all ∣K∣|\mathcal{K}| keypoints. A voting-based keypoint localization scheme is applied to extract the coordinates of 2D keypoints from this 2∣K∣2|\mathcal{K}|-channel tensor and the segmentation mask MM.

The next 2∣E∣2|\mathcal{E}| channels in the output tensor give the xx and yy components of all ∣E∣|\mathcal{E}| edges, which we denote as EdgeEdge. Let ii (0≤i<∣E∣0\leq i<|\mathcal{E}|) be the index of an edge. Then

is a set of 2-tuples containing pixel-wise predictions of the ithi^{th} edge vector in EdgeEdge. The mean of EdgeiEdge_{i} is extracted as the predicted edge.

The final 2 channels in the output tensor define the xx and yy components of symmetry correspondences. We denote this 2-channel “map” of symmetry correspondences as SymSym. Let (x,y)(x,y) be a pixel on the object of interest in the input image, i.e. M(x,y)=1M(x,y)=1. Assuming Δx=Sym(0,x,y)\Delta x=Sym(0,x,y) and Δy=Sym(1,x,y)\Delta y=Sym(1,x,y), we consider (x,y)(x,y) and (x+Δx,y+Δy)(x+\Delta x,y+\Delta y) to be symmetric with respect to the reflection symmetry plane.

The architecture described above achieves good performance in terms of detection accuracy. Nevertheless, it should be emphasized that the framework of HybridPose can incorporate future improvements in keypoint, edge vector, and symmetry correspondence detection techniques. Besides, Hybridpose can be extended to handling multiple objects within an image. One approach is to predict instance-level rather than semantic-level segmentation masks by methods such as Mask R-CNN . Intermediate representations are then extracted from each instance, and fed to the pose regression module in 3.3.

Evaluation protocols. We use two metrics to evaluate the performance of HybridPose:

1. ADD(-S) first calculates the distance between two point sets transformed by predicted pose and ground-truth pose respectively, and then extracts the mean distance. When the object possesses symmetric pose ambiguity, the mean distance is computed from the closest points between two transformed sets. ADD(-S) accuracy is defined as the percentage of examples whose calculated mean distance is less than 10% of the model diameter.

2. In the ablation study, we compute and report the the angular rotation error ∥log⁡(RgtTRI)2∥\|\frac{\log(R_{gt}^{T}R_{I})}{2}\| and the relative translation error ∥tI−tgt∥d\frac{\|\boldsymbol{t}_{I}-\boldsymbol{t}_{gt}\|}{d} between the predicted pose (RI,tI)(R_{I},\boldsymbol{t}_{I}) and the ground-truth pose (Rgt,tgt)(R_{gt},\boldsymbol{t}_{gt}), where dd is object diameter.

2 Analysis of Results

As shown in Table 1, Table 2, and Figure 3, HybridPose leads to accurate pose estimation. On Linemod and Occlusion Linemod, HybridPose has an average ADD(-S) accuracy of 91.3 and 47.5, respectively. The result on Linemod outperforms all except one state-of-the-art approaches that regress poses from intermediate representations. The result on Occlusion-Linemod outperforms all state-of-the-art approaches.

Baseline comparison on Linemod. HybridPose outperforms PVNet , the backbone model we use to predict keypoints. The improvement is consistent across all object classes, which demonstrates clear advantage of using a hybrid as opposed to unitary intermediate representation. HybridPose shows competitive results against DPOD , winning on six object classes. The advantage of DPOD comes from data augmentation and explicit modeling of dense correspondences between input and projected images, both of which cater to situations without object occlusion. A detailed analysis reveals that the classes of objects on which HybridPose exhibits sub-optimal performance are among the smallest objects in Linemod. It suggests that pixel-based descriptors used in our pipeline are limited by image resolution.

Baseline comparison on Occlusion Linemod. HybridPose outperforms all baselines. In terms of ADD(-S), our approach improves PVNet from 40.8 to 47.5, representing a 16.4% enhancement, which clearly shows the advantage of HybridPose on occluded objects, where predictions of invisible keypoints can be noisy, and visible keypoints may not provide sufficient constraints for pose regression alone. HybridPose also outperforms DPOD, the state-of-the-art model on this dataset.

Running time. On a desktop with 16-core Intel(R) Xeon(R) E5-2637 CPU and GeForce GTX 1080 GPU, HybridPose takes 0.6 second to predict the intermediate representations, 0.4 second to regress the pose. Assuming a batch size of 30, this gives an an average processing speed of around 30 fps, enabling real-time analysis.

3 Ablation Study

Table 3 summarizes the performance of HybridPose using different predicted intermediate representations on the Linemod dataset.

With keypoints. As a baseline approach, we estimate object poses by only utilizing keypoint information. This gives a mean absolute rotation error of 1.357°, and a mean relative translation error of 0.061.

With keypoints and symmetry. Adding symmetry correspondences to keypoints leads to some performance gain in rotation. On the other hand, the translation error remains almost the same. One explanation is that symmetry correspondences only constrain two degrees of freedom in a total of three rotation parameters, and provide no constraint on translation parameters (see (3)).

Full model. Adding edge vectors to keypoints and symmetry correspondences leads to salient performance gain in both rotation and translation estimations. One explanation is that edge vectors provide more constraints on both translation and rotation (see (2)). Edge vectors provide more constraints on translation than keypoints as they represent adjacent keypoints displacement and provide gradient information for regression. Unlike symmetry correspondences, edge vectors constrain 3 degrees of freedom on rotation parameters which further boosts the performance of rotation estimation.

Conclusions and Future Work

In this paper, we introduce HybridPose, a 6D pose estimation approach that utilizes keypoints, edge vectors, and symmetry correspondences. Experiments show that HybridPose enjoys real-time prediction and outperforms current state-of-the-art pose estimation approaches in accuracy. HybridPose is robust to occlusion. In the future, we plan to extend HybridPose to include more intermediate representations such as shape primitives, normals, and planar faces. Another possible direction is to enforce consistency across different representations in a similar way to as a self-supervision loss in network training.

Acknowledgement

We would like to acknowledge the support of this research from NSF DMS-1700234, a Gift from Snap Research, and a hardware donation from NVIDIA.

References