CPPF: Towards Robust Category-Level 9D Pose Estimation in the Wild

Yang You, Ruoxi Shi, Weiming Wang, Cewu Lu

Introduction

Estimation of 3D position, orientation and scale of novel objects, namely, category-level 9D pose estimation in the wild, is of great importance in many fields, such as robotics and human-object interactions . There are many prior works exploring this direction, but with limitations, though. Some past works have explored the instance-level 6D pose estimation. However, they require exact object models and their sizes beforehand, which is often not realizable in real-world scenarios. NOCS introduces normalized object coordinate space to give a consistent representation across intra-class objects. Although it is able to achieve category-level pose estimations, it requires real-world pose annotations, which are tedious and limited by size. Besides, the 3D object scales predicted by NOCS are simple heuristics and prone to object occlusions, which is inevitable in real world. We doubt if one can leverage a sim-to-real approach that generalize 9D pose estimations from synthetic objects to real world, since ground-truth pose annotations in real-world scenarios are hard to acquire. Gao et al. tries to solve this problem via comparing the appearance of objects of synthetic objects and real objects, and finds a pose that minimize the difference. Though this method does not require real 9D pose labels, it is erroneous and inferior to NOCS in terms of orientation error, due to the domain gap between synthetic and real RGB images.

Embracing these challenges, in this paper, we propose Category-level PPF (CPPF) that votes for category-level 9D poses. We draw some inspirations from traditional point pair features (PPFs), and formulate the problem of pose estimation as a voting process, where each point pair would generate several offsets or relative angles towards ground-truth 9D poses. Next, the pose with the most votes is cast as our final prediction. In contrast to traditional instance-level PPFs, where each pair is matched against an offline database, our method is much faster and able to generalize to unseen objects. To overcome the difficulty in orientation voting, where false positives are generated, an auxiliary binary classification task is introduced.

In order to segment out the point cloud of a real-world target object, we leverage an off-the-shelf or fine-tuned instance segmentation model. We argue that instance segmentation labels are much cheaper and easier to annotate than 9D poses.

Besides, we develop a two-stage coarse-to-fine voting method, to robustly estimate the object pose when predictions of instance segmentation are not accurate. This method could eliminate noisy point pairs that do not contribute to the voting of object pose. Furthermore, our additional experiments show that when only object bounding boxes are available, our method could still achieve robust and appealing results.

We evaluate our method on the publicly available real-world dataset released by Wang et al. on category-level object pose estimation. Results show that our method beats sim-to-real state of the arts and is comparable to real-world training methods. In addition, we show that when only bounding box detections are provided, our method still gives decent 9D pose predictions. To further evaluate the generalization ability of our method to real-world scenarios, we directly apply our method on SUN RGB-D dataset with zero-shot transfer, which contains much more diverse and complex scenes. Our method also outperforms baselines by a large margin. In summary, our contribution is:

We propose a novel category-level voting scheme to extract 9D pose of objects. An auxiliary task is introduced to remove the ambiguity in orientation voting. A coarse-to-fine voting algorithm is proposed to eliminate noisy point pairs with robust pose predictions.

We introduce a novel sim-to-real pipeline with carefully designed point pair features to achieve generalizable sim-to-real transfer, with synthetic models only.

Extensive experiments show that our sim-to-real method is on par with current state-of-the-art methods, which utilize real-world training data. Besides, our model is robust to segmentation errors and could give accurate pose predictions with only bounding box detections.

Related Works

There are many pose detection methods that requires only point clouds. Drost et al. propose point pair features to match against an object database using a voting scheme. Later, some researchers improves upon Drost’s work: Hinterstoisser et al. leverage a smart sampling scheme to restrict the searching range; Vidal et al. improve the matching process by considering neighborhoods that potentially affected due to noise. It also improves the post-processing step such that the retrieved pose is more consistent with the observed camera view. Shi et al. propose a method on generating object poses from stable geometric groups. There are also many works taking RGB(-D) images as input. Kehl et al. extend the popular SSD paradigm to cover the full 6D pose space. Branchmann et al. learn to classify each pixel into a set of normalized coordinates and then generates a set of candidates by RANSAC. Grabner et al. render depth images from 3D models using the predicted poses and match learned image descriptors of RGB images against those of rendered depth images using a CNN-based multi-view metric learning approach. Rios et al. use a discriminative learning approach to match the object pose in images against a database. DeepIM leverages a FlowNet to output relative pose between real and rendered image patches. The pose is refined in an iterative way. Kehl et al. learn to auto-encode RGB-D patches and match them in a codebook to vote for the final pose. Gao et al. directly regress 6D poses from object point clouds. These methods lack the ability to generalize to unseen objects, and most of them do not scale well.

Recently, a few works focus on category-level object estimation, where an unseen object’s pose is to be detected. NOCS learns to regress objects’ normalized coordinates establishing 2D-3D relationships, so that object poses can be solved in closed form. However, it requires real-world pose annotations for training. SPD improves the predictions of canonical object models by deforming categorical shape priors. Then, CASS use a variational auto-encoder to capture pose-independent features, along with pose-dependent ones, to directly predict the 6D poses. FS-Net proposes a decoupled rotation mechanism that uses two decoders to decode the category-level rotation information. For translation and size estimation, it uses a residual estimation network. DualPoseNet leverages two parallel decoders either make a pose prediction explicitly, or implicitly do so by reconstructing the input point cloud in its canonical pose. The explicit prediction is then refined with the implicit one. Chen et al. propose to render synthetic models and compare the appearance with real images under different poses. This method, though achieves sim-to-real transfer, is inferior to our method, due to the domain gap between synthetic and real RGB images. Their method, however, is prone to occlusion and noise.

2 Sim-to-Real Transfer

Sim-to-Real is a common strategy in many fields like object reconstruction, pose estimation and reinforcement learning for robots. ShapeHD and MarrNet render realistic ShapeNet models for 3D object reconstruction from a single RGB image. PoseCNN renders different objects into random background to synthesis images for training of object poses. Many works follow this paradigm, and training on synthetic RGB images have been a common practice in object pose estimation. In reinforcement learning, domain adaption methods are usually leveraged, and visual/physical realistic simulation of real environment plays an important role. Most existing methods explore the domain transfer in color space, while few works focus on the domain gap between synthetic and real point clouds. This is due to the fact that in real scenarios, objects are often occluded with noisy backgrounds.

Preliminaries: Point Pair Features

In this section, we briefly discuss the original point pair features (PPFs) proposed by Drost et al. , which can be leveraged for instance-level retrieval.

Given two points p1\mathbf{p}_{1} and p2\mathbf{p}_{2} with normals n1\mathbf{n}_{1} and n2\mathbf{n}_{2}, set d=p2−p1\mathbf{d}=\mathbf{p}_{2}-\mathbf{p}_{1} and define the so-called point pair features FF:

In the offline phase, the global model description is created. Such a global model description contains all the pre-calculated PPFs for the object of interest.

In the online phase, a set of reference points in the scene is selected. All other points in the scene are paired with the reference points to create point pair features. These features are matched to the model features contained in the global model description, and a set of potential matches is retrieved. Every potential match votes for an object pose by using an efficient voting scheme where the pose is parametrized relative to the reference point. Specifically, for each match, the pose can be retrieved by aligning the PPF in the scene to that in the offline database. For more details, we refer the reader to Drost et al. . Though Drost PPF has been successful in many scenarios, it can not do a category-level pose estimation, and it is not scalable as the number of objects goes large.

Methodology

To address the problem raised by instance-level PPFs, we propose Category-level PPF (CPPF) - a brand-new method for detecting category-level object poses. Compared with Drost PPF, our method is free of an offline database. Instead, we directly predict the necessary statistics to vote for object centers, orientations and scales, using a neural network with augmented point pair features as the input.

We assume the target object is first segmented out by some off-the-shelf or fine-tuned instance segmentation model. The segmentation does not need to be exact, and we will introduce our coarse-to-fine voting process to robustly estimate 9D poses when the segmentation is inaccurate. Furthermore, in Section 6.2, we show that instance segmentation can be replaced by coarse bounding box masks.

Denote the object center as o\mathbf{o}, for each point pair p1\mathbf{p}_{1} and p2\mathbf{p}_{2}, we predict the following two offsets:

The proof is simple and omitted, observing that both inner product and L2 norm are invariant to rotations.

Once μ\mu and ν\nu are fixed, the object center is determined up to one degree-of-freedom ambiguity. Specifically, the object center will lie on a circle, with center c=p1+μ⋅p1p2→∥p1p2→∥2\mathbf{c}=\mathbf{p}_{1}+\mu\cdot\frac{\overrightarrow{\mathbf{p}_{1}\mathbf{p}_{2}}}{\|\overrightarrow{\mathbf{p}_{1}\mathbf{p}_{2}}\|_{2}} and radius ν\nu, demonstrated in Figure 2(a).

Inspired by Canonical Voting , we can enumerate every 2πK\frac{2\pi}{K} degree and generate multiple votes along the circle. Though there is a circle ambiguity for a single pair, when there are enough pairs, the ground truth location will emerge with the largest vote count, shown in Figure 2(b).

Another nice property of our voting scheme is that symmetric objects are naturally handled without any special treatment. Previous methods like NOCS require a special treatment of symmetric objects because of the ambiguity of the normalized space. They map the same input features to different outputs due to symmetry. In contrast, in our model, input features of symmetric point pairs are exactly the same (due to the SE(3)SE(3) invariance of PPF), and the output offsets for these pairs are also identical, so that our model learns a proper functional mapping. This is illustrated in Figure 3.

1.2 Voting for Orientations

Denoting the up orientation as e1\mathbf{e}_{1} and right orientation as e2\mathbf{e}_{2}, we predict the following two relative angles:

These two angles are also invariant to arbitrary rotations.

Analogous to center voting, there is also an ambiguity of one degree of freedom, shown in Figure 4(a). Likewise, we also generate a set of proposals for each point pair with a constant degree interval, and then select the predicted orientation as the one with the largest voting count. Because the orientation is continuous, in practice, we uniformly enumerate a set of orientations from unit sphere. For each orientation, we count the number of vote candidates that fall into a fixed solid angle around the orientation. The orientation with the largest count is identified as the final prediction. This is illustrated in Figure 4(b).

Unfortunately, for orientations, a fake peak with opposite direction sometimes appears when the object is symmetric in structure. When the relative angle α\alpha is about π2\frac{\pi}{2} for a majority of point pairs, the opposite orientation to ground-truth (i.e., -e1\mathbf{e}_{1}) also receives a lot of point votes, giving a false positive.

To eliminate false positives, an auxiliary task is introduced. For each point pair p1,p2\mathbf{p}_{1},\mathbf{p}_{2}, we calculate p1\mathbf{p}_{1}’s normal n1\mathbf{n}_{1} (normal ambiguity removed by ensuring n1⋅p1p2→<0\mathbf{n}_{1}\cdot\overrightarrow{\mathbf{p}_{1}\mathbf{p}_{2}}<0). Then we do a binary classification on the following two auxiliary variables:

1.3 Voting for Scales

1.4 Coarse-to-Fine Voting Process

In previous sections, we describe the overall voting process for 9D poses (i.e., translation, rotation, scale). However, the generated object pose (especially orientation) might be inaccurate due to noisy points when the instance segmentation is not precise. In order to filter out these noisy points, we propose a coarse-to-fine voting algorithm. Specifically, we first vote for object centers with all the points and then back-trace these votes, keeping only the points that generate enough votes close to the voted object center. Once the noisy points are removed, we vote for the object pose again with the filtered points. A formal description of this algorithm is illustrated in Algorithm 1.

2 Sim-to-Real Transfer

One big advantage of our method is that we only need to train on the synthetic models, and then generalize to real-world scenarios. We achieve this by using the depth map during both training and testing phases, and only local point features are leveraged as input. We find that depth or point clouds are much more accurate in sim-to-real generalizations. In contrast, color information is harder to transfer in real-world scenarios, because light conditions are really hard to tune in order to generalize. Color information is only leveraged when there are several ambiguous poses that cannot be distinguished from point clouds.

For each category, we choose several synthetic topology-correct models from ShapeNetCore55 similar to that in NOCS . Then we use OpenGL to render each model’s depth map from a sampled perspective. All points from the back faces get culled in order to simulate self-occlusion. Notice that compared with NOCS, our method does not need to choose a random background and paste synthetic models onto it. The only requirement is the synthetic model itself.

Another problem of sim-to-real transfer is the different sampling density in simulated and real scenarios. To mitigate the domain gap, during both training and testing, we voxelize input point clouds with a predefined resolution to obtain a constant sampling density. Besides, we observe that both simulated and real point clouds have a grid artifact. This is due to the rasterization of image pixels. To solve this issue, we randomly jitter the simulated and real point clouds, leading to an improvement on the final results.

2.2 Use Color Information to Disambiguate Poses

For most categories, our rendered depth images (i.e., point clouds) generalize well to real-world. However, for laptop, there are two ambiguous poses. The laptop lid and keyboard base are hard to discriminate with point clouds only, even for humans. To solve this problem, we train an additional network that takes RGB inputs to segment the lid and keyboard. The training data for this network only contains rendered synthetic laptop images with Blender , so that our model is still free of real training images. When testing, we calculate the normal of laptop keyboard by RANSAC plane detection on the predicted segmentation. If the voting result is inconsistent with this normal, we replace it with the normal, otherwise the result is unchanged.

Implementation Details

In practice, we convert the scalar regression problem into a multi-class classification problem by using a list of anchors, and find that this gives us a better result. We use 32 bins for translation and 36 bins for rotation. In both training and inference stage, we uniformly sample a fixed number of points (20,000 in training, 100,000 in testing) per model/image and predict their corresponding statistics (i.e., μ,ν,α,β,γ\mu,\nu,\alpha,\beta,\bm{\gamma}). In the center voting process, the accumulation 3D grid has a resolution of 0.4 cm except for laptop which is 1 cm. The range of 3D grid is the tightest axis-aligned bounding box of input. In the orientation voting process, the orientation grid has a resolution of 1.5 degrees.

We use SPRIN , which is an SO(3)SO(3) invariant network, and we modified input features to make it SE(3)SE(3) invariant. The input to the network is the point pair and the set of kk nearest neighbors of each point. Denoting kk neighbors of point p\mathbf{p} as {p(1),⋯ ,p(k)}\{\mathbf{p}^{(1)},\cdots,\mathbf{p}^{(k)}\}, the sides and angles of triangles formed by 1k∑1kp(n),p(n)\frac{1}{k}\sum_{1}^{k}\mathbf{p}^{(n)},\mathbf{p}^{(n)} and p\mathbf{p} are fed to SPRIN to extract rotation invariant point embeddings. Besides, the normal for each point is also estimated and leveraged from its kk nearest neighbors. The network profession is illustrated in Figure 5. We train each category separately, with Adam optimizer, using learning rate 1e-3, for 200 epochs. The batch size is 1.

Experiments

In this section, we evaluate our method on two datasets. NOCS REAL275 and SUN RGB-D . Both datasets provide RGB-D frames with annotated 9D bounding boxes.

We follow NOCS to report both intersection over union and 6D pose average precision. Intersection over union (IoU) is calculated between the predicted and ground-truth bounding boxes with threshold of 50%, while 6D pose average precision is calculated by measuring the average precision of object instances for which the error is less than mm cm for translation and n∘n^{\circ} for rotation. We follow NOCS to set a detection threshold of 10% bounding box overlap between prediction and ground truth to ensure that most objects are included in the evaluation. Notice that the original 3D box mAP computation code provided by NOCS is buggy. Instead, we use the correct code from Objectron .

1 NOCS REAL275 with Instance Mask

Wang et al. captures 8K real RGB-D frames (4300 for training, 950 for validation and 2750 for testing) of 18 different real scenes (7 for training, 5 for validation, and 6 for testing) using a Structure Sensor. We use the 2750 testing scenes for evaluation. We use the instance segmentation masks from NOCS for a fair comparison.

We compare our method with a set of real-world training methods: NOCS , CASS , SPD , FS-Net , DualPoseNet ; and a set of methods requiring synthetic training data only: Chen et al. , Gao et al. . The original mAP results reported by Chen et al. uses the IoU matches computed by NOCS which is potentially unfair, we fix this by using the matches computed by Chen et al.’s method itself. We also augment Gao et al.’s method to additionally regress 3D box scales. NOCS , Chen et al. , Gao et al. and DualPoseNet ’s results are given by running the official code provided by the authors, while the others are borrowed from the original papers. Notice that DualPoseNet uses its own instance segmentation masks other than those provided by NOCS, which may result in a higher mAP than the actual.

The results are given in Table 1. We see that our method achieves an mAP of 16.9, 44.9 and 50.8 for (5∘, 5 cm), (10∘, 5 cm) and (15∘, 5 cm) respectively. It outperforms the best sim-to-real baseline by 9.1, 27.8 and 24.3, which is a quite large margin. Our method is also comparable to those methods that are trained on real-world pose annotations. More detailed analysis and comparison is illustrated in Figure 6. Some qualitative comparisons are given in Figure 7. This experiment shows that our proposed method generalize well to real-world data with only synthetic training data.

2 NOCS REAL275 with Bounding Box Masks

Though our method does not require pose annotations for real-world data, it does need to first segment out the target object with a real-world trained instance segmentation model. Can we relax this requirement? Thanks to our robust coarse-to-fine voting algorithm 1 which filters out noisy points, we find that when only bounding boxes are given, our method still achieves significant results. Notice that current state of the arts all require pixel-wise instance segmentation as input. Quantitative results are given in Table 1. Qualitative pose predictions are shown in Figure 8.

Moreover, we also tried to get rid of detection priors completely for bowls, more details in the supplementary.

3 SUN RGB-D in the Wild

SUN RGB-D is a scene understanding benchmark which provides 58,657 9D bounding boxes with accurate object orientations for 10,000 images. We use all the chairs in validation split for evaluation, which contains 2,699 images. In order to make the problem more challenging, we randomly rotate the SUN RGB-D scenes while the original point clouds are aligned with gravity. In addition, we require that all the algorithms cannot see any training data but use existing instance segmentation models (i.e., trained on MSCOCO but not fine-tuned on SUN RGB-D).

We compare our method with two baselines: direct back-projection and Gao et al. . Direct back-projection is a simple baseline that directly back project the detected instance into an axis-aligned bounding box, while Gao et al. is the same as in Section 6.1. Since this is an extremely difficult task, we only evaluate orientation errors along the up axis. Results are listed in Table 2. More results are given in the supplementary.

4 Ablation Study and Running Time

In this section, we conduct various ablation studies on our model. Results are reported on REAL275 test set.

As the number of pair samples increase, the voting results for orientation and translation become more accurate, while getting saturated for 100,000 point pairs. The size of discrete orientation bins also decides the accumulation accuracy during the voting process. Quantitative results are given in Table 3.

The auxiliary classification task helps our model get rid of potentially flipped orientations, and Table 4 verifies this.

Recall, in Algorithm 1, we filter out point pairs that do not contribute to the proposed object center. This makes our model robust to the noisy points from the imperfect instance segmentation. Quantitative results are given in Table 4.

Direct regression on relevant statistics are worse than classification. This may due to the fact that regression does not constrain the value into a valid range and produces more noisy outputs. Quantitative results are shown in Table 4.

It takes 171ms, 229ms and 13ms per image for the voting of centers, orientations and scales respectively on a single 1080Ti GPU. Our model is efficient thanks to the highly parallelized voting process.

Conclusion

In this paper, we propose a category-level voting algorithm to predict 9D poses in the wild. To overcome the difficulty of false positives during the voting step, an auxiliary orientation classification task is introduced. Our model is trained on synthetic objects and generalizes well to real scenes. Results show that our method is superior to previous sim-to-real methods, even with bounding box masks.

Acknowledgements

This work was supported by the National Key Research and Development Project of China (No. 2021ZD0110700), the National Natural Science Foundation of China under Grant 51975350, Shanghai Municipal Science and Technology Major Project (2021SHZDZX0102), Shanghai Qi Zhi Institute, and SHEITC (2018-RGZN-02046). This work was also supported by the Shanghai AI development project (2020-RGZN-02006) and “cross research fund for translational medicine” of Shanghai Jiao Tong University (zh2018qnb17, zh2018qna37, YG2022ZD018).

References

Supplementary

Appendix A Direct 9D Pose Estimation and Segmentation

In our experiments, we show that our model can work with both segmentation and bounding box masks. Can we locate objects with our voting scheme directly (i.e., without any preprocessing instance detection pipeline)? For some categories like laptop, this is hard since the laptop base can be mixed with the floor in view of point clouds. However, for bowls, this is possible, where the pose estimation is indeed zero-shot without seeing any real-world data during the whole pipeline. We modify Algorithm 1 by sampling point pairs from the whole scene, while keeping the coarse-to-fine procedure. We also use a threshold to generate candidate object locations instead of a simple argmax (line 9 in Algorithm 1).

Surprisingly, as a by-product, our method is able to infer the instance-level mask without even seeing any segmentation labels. This is done by counting the contribution of each point, where any point that votes more than vv times within a small radius ϵ\epsilon of the true object center (line 15 in Algorithm 1), is considered lying on the target object. Qualitative 9D pose and segmentation results are given in Figure 9. Quantitative results are listed in Table 5.

Appendix B More Quantitative Results on SUN RGB-D Dataset

We conduct more zero-shot pose estimation experiments on SUN RGB-D, with the instance masks annotated by SUN RGB-D. Notice that our model is trained solely on ShapeNet synthetic models, and then directly tested on SUN RGB-D frames. Results are listed in Table 6. Qualitative results are given in Figure 10.

Appendix C More Ablation Studies

In our sim-to-real pipeline, point clouds are first voxelized and jittered in order to generate the same distribution for both synthetic and real objects. Table 7 gives the ablation studies on these two techniques, where we see that both voxelization and random jittering helps improve the final detection result.

Appendix D Detailed results on each category

We plot detailed comparisons of each category on NOCS REAL275 test set. Rotation AP is given in Figure 11 and Translation AP is shown in Figure 12. Our model achieves the best result on most categories, and outperforms previous self-supervised methods by a large margin.