PVNet: Pixel-wise Voting Network for 6DoF Pose Estimation

Sida Peng, Yuan Liu, Qixing Huang, Hujun Bao, Xiaowei Zhou

Introduction

Object pose estimation aims to detect objects and estimate their orientations and translations relative to a canonical frame xiang2014beyond. Accurate pose estimations are essential for a variety of applications such as augmented reality, autonomous driving and robotic manipulation. For instance, fast and robust pose estimation is crucial in Amazon Picking Challenge correll2018analysis, where a robot needs to pick objects from a warehouse shelf. This paper focuses on the specific setting of recovering the 6DoF pose of an object, i.e., rotation and translation in 3D, from a single RGB image of the object. This problem is quite challenging from many perspectives, including object detection under severe occlusions, variations in lighting and appearance, and cluttered background objects.

Traditional methods lowe1999object; lepetit2005monocular; hinterstoisser2012model have shown that pose estimation can be achieved by establishing the correspondences between an object image and the object model. They rely on hand-crafted features, which are not robust to image variations and background clutters. Deep learning based methods su2015render; kehl2017ssd; xiang2017posecnn; bui2018regression train end-to-end neural networks that take an image as input and output its corresponding pose. However, generalization remains as an issue, as it is unclear that such end-to-end methods learn sufficient feature representations for pose estimation.

Some recent methods pavlakos20176; rad2017bb8; tekin2018real use CNNs to first regress 2D keypoints and then compute 6D pose parameters using the Perspective-n-Point (PnP) algorithm. In other words, the detected keypoints serve as an intermediate representation for pose estimation. Such two-stage approaches achieve state-of-the-art performance, thanks to robust detection of keypoints. However, these methods have difficulty in tackling occluded and truncated objects, since part of their keypoints are unseen. Although CNNs may predict these unseen keypoints by memorizing similar patterns, generalization remains difficult.

We argue that addressing occlusion and truncation requires dense predictions, namely pixel-wise or patch-wise estimates for the final output or intermediate representations. To this end, we propose a novel framework for 6D pose estimation using a Pixel-wise Voting Network (PVNet). The basic idea is illustrated in Figure 1. Instead of directly regressing image coordinates of keypoints, PVNet predicts unit vectors that represent directions from each pixel of the object towards the keypoints. These directions then vote for the keypoint locations based on RANSAC fischler1981random. This voting scheme is motivated from a property of rigid objects that once we see some local parts, we are able to infer the relative directions to other parts.

Our approach essentially creates a vector-field representation for keypoint localization. In contrast to coordinate or heatmap based representations, learning such a representation enforces the network to focus on local features of objects and spatial relations between object parts. As a result, the location of an invisible part can be inferred from the visible parts. In addition, this vector-field representation is able to represent object keypoints that are even outside the input image. All these advantages make it an ideal representation for occluded or truncated objects. Xiang et al. xiang2017posecnn proposed a similar idea to detect objects and here we use it to localize keypoints.

Another advantage of the proposed approach is that the dense outputs provide rich information for the PnP solver to deal with inaccurate keypoint predictions. Specifically, RANSAC-based voting prunes outlier predictions and also gives a spatial probability distribution for each keypoint. Such uncertainties of keypoint locations give the PnP solver more freedom to identify consistent correspondences for predicting the final pose. Experiments show that the uncertainty-driven PnP algorithm improves the accuracy of pose estimation.

We evaluate our approach on LINEMOD hinterstoisser2012model, Occlusion LINEMOD brachmann2014learning and YCB-Video xiang2017posecnn datasets, which are widely-used benchmark datasets for 6D pose estimation. Across all datasets, PVNet exhibits state-of-the-art performances. We also demonstrate the capability of our approach to handle truncated objects on a new dataset called Truncation LINEMOD which is created by randomly cropping images of LINEMOD. Furthermore, our approach is efficient, which runs 25 fps on a GTX 1080ti GPU, to be used for real-time pose estimation.

In summary, this work has the following contributions:

We propose a novel framework for 6D pose estimation using a pixel-wise voting network (PVNet), which learns a vector-field representation for robust 2D keypoint localization and naturally deals with occlusion and truncation.

We propose to utilize an uncertainty-driven PnP algorithm to account for uncertainties in 2D keypoint localizations, based on the dense predictions from PVNet.

We demonstrate significant performance improvements of our approach compared to the state of the art on benchmark datasets (ADD: 86.3% vs. 79% and 40.8% vs. 30.4% on LINEMOD and OCCLUSION, respectively). We also create a new dataset for evaluation on truncated objects.

Related work

Given an image, some methods aim to estimate the 3D location and orientation of the object in a single shot. Traditional methods mainly rely on template matching techniques huttenlocher1993comparing; gu2010discriminative; hinterstoisser2012gradient; zhu2014single, which are sensitive to cluttered environments and appearance changes. Recently, CNNs have shown significant robustness to environment variations. As a pioneer, PoseNet kendall2015posenet introduces a CNN architecture to directly regress a 6D camera pose from a single RGB image, a task similar to object pose estimation. However, directly localizing objects in 3D is difficult due to a lack of depth information and the large search space. To overcome this problem, PoseCNN xiang2017posecnn localizes objects in the 2D image and predicts their depths to obtain the 3D location. However, directly estimating the 3D rotation is also difficult, since the non-linearity of the rotation space makes CNNs less generalizable. To avoid this problem, tulsiani2015viewpoints; su2015render; liu2016ssd; sundermeyer2018implicit discretize the rotation space and cast the 3D rotation estimation into a classification task. Such discretization produces a coarse result and a post-refinement is essential to get an accurate 6DoF pose.

Keypoint-based methods.

Instead of directly obtaining the pose from an image, keypoint-based methods adopt a two-stage pipeline: they first predict 2D keypoints of the object and then compute the pose through 2D-3D correspondences with a PnP algorithm. 2D keypoint detection is relatively easier than 3D localization and rotation estimation. For objects of rich textures, traditional methods lowe1999object; rothganger20063d; bay2006surf detect local keypoints robustly, so the object pose is estimated both efficiently and accurately, even under cluttered scenes and severe occlusions. However, traditional methods have difficulty in handling texture-less objects and processing low-resolution images lepetit2005monocular. To solve this problem, recent works define a set of semantic keypoints and use CNNs as keypoint detectors. rad2017bb8 uses segmentation to identify image regions that contain objects and regresses keypoints from the detected image regions. tekin2018real employs the YOLO architecture redmon2017yolo9000 to estimate the object keypoints. Their networks make predictions based on a low-resolution feature map. When global distractions occur, such as occlusions, the feature map is interfered oberweger2018making and the pose estimation accuracy drops. Motivated by the success of 2D human pose estimation newell2016stacked, another category of methods pavlakos20176; oberweger2018making outputs pixel-wise heatmaps of keypoints to address the issue of occlusion. However, since heatmaps are fix-sized, these methods have difficulty in handling truncated objects, whose keypoints may be outside the input image. In contrast, our method makes pixel-wise predictions for 2D keypoints using a more flexible representation, i.e., vector field. The keypoint locations are determined by voting from the directions, which are suitable for truncated objects.

Dense methods.

In these methods, every pixel or patch produces a prediction for the desired output, and then casts a vote for the final result in a generalized Hough voting scheme liebelt2008independent; sun2010depth; glasner2011aware. brachmann2014learning; michel2017global use a random forest to predict 3D object coordinates for each pixel and produce 2D-3D correspondence hypotheses using geometric constraints. To utilize the powerful CNNs, kehl2016deep; doumanoglou2016recovering densely sample image patches and use networks to extract features for the latter voting. However, these methods require RGB-D data. In the presence of RGB data alone, brachmann2016uncertainty uses an auto-context regression framework tu2010auto to produce pixel-wise distributions of 3D object coordinates. Compared with sparse keypoints, object coordinates provide dense 2D-3D correspondences for pose estimation, which is more robust to occlusion. But regressing object coordinates is more difficult than keypoint detection due to the larger output space. Our approach makes dense predictions for keypoint localization. It can be regarded as a hyprid of keypoint-based and dense methods, which combines advantages of both methods.

Proposed approach

In this paper, we propose a novel framework for 6DoF object pose estimation. Given an image, the task of pose estimation is to detect objects and estimate their orientations and translations in 3D. Specifically, 6D pose is represented by a rigid transformation (R;t)(R;{\bf t}) from the object coordinate system to the camera coordinate system, where RR represents the 3D rotation and t\bf t represents the 3D translation.

Inspired by recent methods pavlakos20176; rad2017bb8; tekin2018real, we estimate the object pose using a two-stage pipeline: we first detect 2D object keypoints using CNNs and then compute 6D pose parameters using the PnP algorithm. Our innovation is in a new representation for 2D object keypoints as well as a modified PnP algorithm for pose estimation. Specifically, our method uses a Pixel-wise Voting Network (PVNet) to detect 2D keypoints in a RANSAC-like fashion, which robustly handles occluded and truncated objects. The RANSAC-based voting also gives a spatial probability distribution of each keypoint, allowing us to estimate the 6D pose with an uncertainty-driven PnP.

Figure 2 overviews the proposed pipeline for keypoint localization. Given an RGB image, PVNet predicts pixel-wise object labels and unit vectors that represent the direction from every pixel to every keypoint. Given the directions to a certain object keypoint from all pixels belonging to that object, we generate hypotheses of 2D locations for that keypoint as well as the confidence scores through RANSAC-based voting. Based on these hypotheses, we estimate the mean and covariance of the spatial probability distribution for each keypoint.

In contrast to directly regressing keypoint locations from an image patch rad2017bb8; tekin2018real, the task of predicting pixel-wise directions enforces the network to focus more on local features of objects and alleviates the influence of cluttered background. Another advantage of this approach is the ability to represent keypoints that are occluded or outside the image. Even if a keypoint is invisible, it can be correctly located according to the directions estimated from other visible parts of the object.

More specifically, PVNet performs two tasks: semantic segmentation and vector-field prediction. For a pixel p{\bf p}, PVNet outputs the semantic label that associates it with a specific object and the unit vector vk(p){\bf v}_{k}({\bf p}) that represents the direction from the pixel p\bf p to a 2D keypoint xk{\bf x}_{k} of the object. The vector vk(p){\bf v}_{k}({\bf p}) is defined as

Given semantic labels and unit vectors, we generate keypoint hypotheses in a RANSAC-based voting scheme. First, we find the pixels of the target object using semantic labels. Then, we randomly choose two pixels and take the intersection of their vectors as a hypothesis hk,i{\bf h}_{k,i} for the keypoint xk{\bf x}_{k}. This step is repeated NN times to generate a set of hypotheses {hk,i∣i=1,2,...,N}\{{\bf h}_{k,i}|i=1,2,...,N\} that represent possible keypoint locations. Finally, all pixels of the object vote for these hypotheses. Specifically, the voting score wk,iw_{k,i} of a hypothesis hk,i{\bf h}_{k,i} is defined as

The resulting hypotheses characterize the spatial probability distribution of a keypoint in the image. Figure 2(e) shows an example. Finally, the mean μk{\boldsymbol{\mu}_{k}} and the covariance Σk{\bf\Sigma}_{k} for a keypoint xk{\bf x}_{k} are estimated by:

which are used latter for uncertainty-driven PnP described in Section 3.2.

The keypoints need to be defined based on the 3D object model. Many recent methods rad2017bb8; tekin2018real; oberweger2018making use the eight corners of the 3D bounding box of the object as the keypoints. An example is shown in Figure 3(a). These bounding box corners are far away from the object pixels in the image. The longer distance to the object pixels results in larger localization errors, since the keypoint hypotheses are generated using the vectors that start at the object pixels. Figure 3(b) and (c) show the hypotheses of a bounding box corner and a keypoint selected on the object surface, respectively, which are generated by our PVNet. The keypoint on the object surface usually has a much smaller variance in the localization.

Therefore, the keypoints should be selected on the object surface in our approach. Meanwhile, these keypoints should spread out on the object to make the PnP algorithm more stable. Considering the two requirements, we select KK kepoints using the farthest point sampling (FPS) algorithm. First, we initialize the keypoint set by adding the object center. Then, we repeatedly find a point on the object surface, which is farthest to the current keypoint set, and add it to the set until the size of the set reaches KK. The empirical results in Section 5.3 show that this strategy produces better results than using the bounding box corners. We also compare the results using different numbers of keypoints. Considering both accuracy and efficiency, we suggest K=8K=8 according to the experiment results.

Multiple instances.

Our method can handle multiple instances based on the strategy proposed in xiang2017posecnn; papandreou2018personlab. For each object class, we generate the hypotheses of the object centers and their voting scores using our proposed voting scheme. Then, we find the modes among the hypotheses and mark these modes as centers of different instances. Finally, the instance masks are obtained by assigning pixels to the nearest instance center they vote for.

2 Uncertainty-driven PnP

Given 2D keypoint locations for each object, its 6D pose can be computed by solving the PnP problem using an off-the-shelf PnP solver, e.g., the EPnP lepetit2009epnp used in many previous methods tekin2018real; rad2017bb8. However, most of them ignore the fact that different keypoints may have different confidences and uncertainty patterns, which should be considered when solving the PnP problem.

As introduced in Section 3.1, our voting-based method estimates a spatial probability distribution for each keypoint. Given the estimated mean μk{\boldsymbol{\mu}_{k}} and covariance matrix Σk{\bf\Sigma}_{k} for k=1,⋯ ,Kk=1,\cdots,K, we compute the 6D pose (R,t)(R,{\bf t}) by minimizing the Mahalanobis distance:

Implementation details

Assuming there are CC classes of objects and KK keypoints for each class, PVNet takes as input the H×W×3H\times W\times 3 image, processes it with a fully convolutional architecture, and outputs the H×W×(K×2×C)H\times W\times(K\times 2\times C) tensor representing unit vectors and H×W×(C+1)H\times W\times(C+1) tensor representing class probabilities. We use a pretrained ResNet-18 he2016deep as the backbone network, and we make three revisions on it. First, when the feature map of the network has the size H/8×W/8H/8\times W/8, we do not downsample the feature map anymore by discarding the subsequent pooling layers. Second, to keep the receptive fields unchanged, the subsequent convolutions are replaced with suitable dilated convolutions YuKoltun2016. Third, the fully connected layers in the original ResNet-18 are replaced with convolution layers. Then, we repeatedly perform skip connection, convolution and upsampling on the feature map, until its size reaches H×WH\times W, as shown in Figure 2(b). By applying a 1×11\times 1 convolution on the final feature map, we obtain the unit vectors and class probabilities.

We implement hypothesis generation, pixel-wise voting and density estimation using CUDA. The EPnP lepetit2009epnp used to initialize the pose is implemented in OpenCV bradski2000opencv. To obtain the final pose, we use the iterative solver Ceres ceres-solver to minimize the Mahalanobis distance (5). For symmetric objects, there are ambiguities of keypoint locations. To eliminate the ambiguities, we rotate the symmetric object to a canonical pose during training, as suggested by rad2017bb8.

To prevent overfitting, we add synthetic images to the training set. For each object, we render 10000 images whose viewpoints are uniformly sampled. We further synthesize another 10000 images using the “Cut and Paste” strategy proposed in dwibedi2017cut. The background of each synthetic image is randomly sampled from SUN397 xiao2010sun. We also apply online data augmentation including random cropping, resizing, rotation and color jittering during training. We set the initial learning rate as 0.001 and halve it every 20 epochs. All models are trained for 200 epochs.

Experiments

is a standard benchmark for 6D object pose estimation. This dataset exhibits many challenges for pose estimation: cluttered scenes, texture-less objects, and lighting condition variations.

Occlusion LINEMOD brachmann2014learning

was created by additionally annotating a subset of the LINEMOD images. Each image contains multiple annotated objects, and these objects are heavily occluded, which poses a great challenge for pose estimation.

Truncation LINEMOD

To fully evaluate our method on truncated objects, we create this dataset by randomly cropping images in the LINEMOD dataset. After cropping, only 40% to 60% of the area of the target object remain in the image. Some examples are shown in Figure 5.

Note that, in our experiments, the Occlusion LINEMOD and Truncation LINEMOD are used for testing only. Our model tested on these two datasets is only trained on the LINEMOD dataset.

YCB-Video xiang2017posecnn

is a recently proposed dataset. The images are collected from the YCB object set calli2015ycb. This dataset is challenging due to the varying lighting conditions, significant image noise and occlusions.

2 Evalutation metrics

We evaluate our method using two common metrics: 2D projection metric brachmann2016uncertainty and average 3D distance of model points (ADD) metric hinterstoisser2012model.

This metric computes the mean distance between the projections of 3D model points given the estimated and the ground truth pose. A pose is considered as correct if the distance is less than 5 pixels.

ADD metric.

With the ADD metric hinterstoisser2012model, we transform the model points by the estimated and the ground truth poses, respectively, and compute the mean distance between the two transformed point sets. When the distance is less than 10% of the model’s diameter, it is claimed that the estimated pose is correct. For symmetric objects, we use the ADD-S metric xiang2017posecnn, where the mean distance is computed based on the closest point distance. We denote these two metrics as ADD(-S) and use the one appropriate to the object. When evaluating on the YCB-Video dataset, we compute the ADD(-S) AUC proposed in xiang2017posecnn. The ADD(-S) AUC is the area under the accuracy-threshold curve, which is obtained by varying the distance threshold in evaluation.

3 Ablation studies

We conduct ablation studies to compare different keypoint detection methods, keypoint selection schemes, numbers of keypoints and PnP algorithms, on the Occlusion LINEMOD dataset. Table 1 summarizes the results of ablation studies.

To compare PVNet with tekin2018real, we re-implement the same pipeline as tekin2018real but use PVNet to detect the keypoints which include 8 bounding box corners and the object center. The result is listed in the column “BBox 8” in Table 1. The column “Tekin” shows the original result of tekin2018real, which directly regresses coordinates of keypoints via a CNN. Comparing the two columns demonstrates that pixel-wise voting is more robust to occlusion.

To analyze the keypoint selection schemes discussed in Section 3.1, we compare the pose estimation results based on different keypoint sets: “BBox 8” that includes 8 bounding box corners plus the center and “FPS 8” that includes 8 surface points selected by the FPS algorithm plus the center. Comparing “BBox 8” with “FPS 8” in Table 1 shows that the proposed FPS scheme results in better pose estimation.

When exploring the influence of the keypoint number on pose estimation, we train PVNet to detect 4, 8 and 12 surface keypoints plus the object center, respectively. All the three sets of keypoints are selected by the FPS algorithm as described in Section 3.1. Comparing columns “FPS 4”, “FPS 8” and “FPS 12” shows that the accuracy of pose estimation increases with the keypoint number. But the gap between “FPS 8” and “FPS 12” is negligible. Considering efficiency, we use “FPS 8” in all the other experiments.

To validate the benefit of considering the uncertainties in solving the PnP problem, we replace the EPnP lepetit2009epnp used in “FPS 8” with the uncertainty-driven PnP. The results are shown in the last column “FPS 8 + Un” in Table 1, which demonstrate that considering uncertainties of keypoint locations improves the accuracy of pose estimation.

The configuration “FPS 8 + Un” is the final configuration for our approach, which is denoted by “OURS” in the following experiments.

4 Comparison with the state-of-the-art methods

We compare with the state-of-the-art methods which take RGB images as input and output 6D object poses.

In Table 2, we compare our method with rad2017bb8; tekin2018real on the LINEMOD dataset in terms of the 2D projection metric. rad2017bb8; tekin2018real detect keypoints by regression, while our method uses the proposed voting-based keypoint localization. BB8 rad2017bb8 trains another CNN to refine the predicted pose and the refined results are shown in a separate column. Our method achieves the state-of-the-art performance on all objects without the need of a separate refinement stage.

Table 3 shows the comparison of our methods with rad2017bb8; liu2016ssd; tekin2018real in terms of the ADD(-S) metric. Note that we compute the ADD-S metric for the eggbox and the glue, which are symmetric, as suggested in xiang2017posecnn. Comparing to these methods without using refinement, our method outperforms them by a large margin of at least 30.32%. SSD-6D kehl2017ssd significantly improves its own performance using edge alignment to refine the estimated pose. Nevertheless, our method still outperforms it by 7.27%.

Robustness to occlusion.

We use the model trained on the LINEMOD dataset for testing on the Occlusion LINEMOD dataset. Table 4 and Table 5 summarize the comparison with tekin2018real; xiang2017posecnn; oberweger2018making on the Occlusion LINEMOD dataset in terms of the 2D projection metric and the ADD(-S) metric, respectively. For both metrics, our method achieves the best performance among all methods. In particular, our method outperforms other methods by a margin of 10.37% in terms of the ADD(-S) metric. Some qualitative results are shown in Figure 4. The improved performance demonstrates that the proposed vector-field representation enables PVNet to learn the relationship between parts of the object, so that the occluded keypoints can be robustly recovered by the visible parts.

Robustness to truncation.

We evaluate our method on the Truncation LINEMOD dataset. Note that, the model used for testing is only trained on the LINEMOD dataset. Table 6 shows quantitative results in terms of the 2D projection and ADD(-S) metrics. We also test the released model from tekin2018real, but it does not obtain reasonable results as it is not designed for this case.

Figure 5 shows some qualitative results. Even the objects are partially visible, our method robustly recovers their poses. We show two failure cases in the last column of Figure 5, where the visible parts do not provide enough information to infer the poses. This phenomenon is particularly obvious for small objects, such as duck and ape, which have lower accuracies of the pose estimation.

Performance on the YCB-Video dataset.

In Table 7, we compare our method with xiang2017posecnn; oberweger2018making on the YCB-Video dataset in terms of the 2D projection and the ADD(-S) AUC metrics. Our method again achieves the state-of-the-art performance and surpasses Oberweger oberweger2018making which is specially designed for dealing with occlusion. The results of PoseCNN were obtained from Oberweger oberweger2018making.

5 Running time

Given a 480×640480\times 640 image, our method runs at 25 fps on a desktop with an Intel i7 3.7GHz CPU and a GTX 1080 Ti GPU, which is efficient for real-time pose estimation. Specifically, our implementation takes 10.9 ms for data loading, 3.3 ms for network forward propagation, 22.8 ms for the RANSAC-based voting scheme, and 3.1 ms for the uncertainty-driven PnP.

Conclusion

We introduced a novel framework for 6DoF object pose estimation, which consists of the pixel-wise voting network (PVNet) for keypoint localization and the uncertainty-driven PnP for final pose estimation. We showed that predicting the vector fields followed by RANSAC-based voting for keypoint localization gained a superior performance than direct regression of keypoint coordinates, especially for occluded or truncated objects. We also showed that considering the uncertainties of predicted keypoint locations in solving the PnP problem further improved pose estimation. We reported the state-of-the-art performances on all three widely-used benchmark datasets and demonstrated the robustness of the proposed approach on a new dataset of truncated objects.

References