Category-Level 6D Object Pose Estimation in the Wild: A Semi-Supervised Learning Approach and A New Dataset
Yang Fu, Xiaolong Wang
Introduction
Estimating the 6D object pose is one of the core problems in computer vision and robotics. It predicts the full configurations of rotation, translation and size of a given object, which has wide applications including Virtual Reality (VR) , scene understanding , and . There are two directions in 6D object pose estimation. One is performing instance-level 6D pose estimation, where a model is trained to estimate the pose of one exact instance with an existing 3D model . However, learning instance-level model restricts its generalization ability to unseen objects. To achieve generalization to unseen instance, another direction is recently proposed to perform category-level 6D pose estimation using one model . However, the large appearance and shape variance across instances largely increase the difficulty in learning.
To overcome this limitation, Wang et al. take the initial step to collect real-world dataset and annotations for category-level 6D pose estimation. Combining synthetic data with free ground-truth annotations, they show the learned model can be generalized to unseen objects within the same category. While this result is encouraging, the generalization ability of the model is still limited by the number and the diversity of the data due to the challenges in annotating 6D object poses. Specifically, only 8,000 images across 13 scenes are collected and annotated from the dataset proposed in . Thus it is still very challenging to generalize 6D pose estimation on diverse objects in complex scenes.
In this paper, we propose to generalize category-level 6D object pose estimation in the wild. To achieve this goal, we introduce a new dataset and a new semi-supervised learning approach. Our key insight is that, while annotating the 6D object pose is challenging, collecting the RGBD videos for these objects without labels is much easier and more affordable. On the other hand, there are infinite ground-truth annotations for synthetic data which come for free. We propose to leverage the benefits from both sides. We first collect a rich object-centric video dataset with diverse backgrounds and object instances using an RGBD camera. We train our model jointly using synthetic data with the free ground-truth 6D pose annotations and the unlabeled real-world RGBD videos via a silhouette matching objective. In this way, our pose estimation model can be generalized to in-the-wild data with minimum human labor. We collect an RGBD video dataset for 6D object pose estimation in the wild, namely Wild6D. Each video in the dataset shows multiple views of one or multiple objects (see examples in Fig. 1). In total, there are 5166 videos ( 1.1 million images) over 1722 object instances and 5 categories, which is significantly (300x) larger than the previous 6D object pose estimation dataset . For evaluation, we annotate 486 videos over 162 objects.
Given this dataset, we design a novel model for semi-supervised learning, called Rendering for Pose estimation Network (RePoNet). The RePoNet is composed of two branches of networks with a Pose Network to estimate the 6D object pose and a Shape Network to estimate the 3D object shape. Given an RGBD image of an object, our Pose Network first estimates the Normalized Object Coordinate Space (NOCS) map , which will be integrated with lower-layer features to regress the object 9D pose (rotation, translation, and size parameters). Meanwhile, the Shape Network takes a category-level 3D shape prior as an input, and estimates the object shape for the current input instance. During training, we utilize both the synthetic data with the full ground-truths and the real-world data from Wild6D with foreground segmentation masks (obtained by Mask R-CNN ). Given the inputs with synthetic data, we apply the regression losses on both the NOCS map, 6D pose parameters, and object shape. Additionally, given the estimated pose and object shape, we perform differentiable rendering to obtain an object mask projected in 2D. A silhouette matching objective is proposed to compare the difference between the projected mask and the ground-truth mask. With the real RGBD data, although we do not have the 3D ground-truths, we can still learn with the silhouette matching objective by comparing the projected mask against the foreground segmentation. In this way, the gradients are back-propagated through the 6D pose and object shape to adjust both Pose Network and Shape Network for real RGBD data. During inference, we first apply our Pose Network to estimate the NOCS map with a given image. With the NOCS map output, the object pose can be computed by solving the Umeyama algorithm . Some estimation results on Wild6D can be found in Fig. 1.
In our experiments, we first evaluate our approach with the dataset proposed in . Our RePoNet shows a improvement over state-of-the-art approaches when trained in a fully-supervised setting (using 3D ground-truths from both real and synthetic data). By using the real-world data without its 3D ground-truth with semi-supervised learning, our performance can still be on par with previous approaches using full annotations. We then experiment with our Wild6D dataset and consistently show a large performance gain over the baselines on in-the-wild objects. We highlight our main contributions as follows:
A new large-scale dataset Wild6D of object-centric RGBD videos in-the-wild for category-level 6D object pose estimation.
A semi-supervised approach with RePoNet which leverages both synthetic data and real-world data without 3D ground-truths for category-level 6D object pose estimation.
Our approach outperforms baselines by a large margin, especially when deployed on in-the-wild objects. We commit to release our dataset, annotation and code for benchmarking future research.
Related Work
Category-level 6D Object Pose Estimation. Recently, researchers have proposed to learn a single model for one category of objects instead of just one instance on pose estimation. For example, Wang et al. propose a canonical shape representation of different objects called Normalized Object Coordinate Space (NOCS) to handle the instance variations. The object pose and size are calculated by the Umeyama algorithm with predicted NOCS map and the observed points. Follow-up work using the NOCS representation has focused on improving the shape priors and incorporating direct regression of object pose and size . While these approaches can be deployed on unseen object instances, the generalization ability of these models is still limited by the scale and the diversity of real-world annotated data in the REAL275 dataset . Only 13 scenes with 18 real objects in total are presented in REAL275. In contrast, we propose a semi-supervised learning approach with RePoNet, which leverages a new large-scale unlabeled dataset Wild6D for training. Both our method and the new dataset are the key components that lead to our goal of generalizing category-level 6D pose estimation in the wild.
Large-scale 3D Object Datasets. Various large-scale 3D datasets have been proposed for different tasks including reconstruction and pose estimation. For example, the Objectron dataset are collected and studied with large-scale object-centric videos and 3D annotations . However, there are no depth images provided in Objectron, which might lead to ambiguities in 6D object pose estimation. Similarly, the recent proposed CO3D dataset with diverse instances in different categories is also collected without recording the depth. While we can perform 3D reconstruction or generate the depth map using COLMAP , the error in predicted depth makes it difficult to be used for 6D pose estimation. Different from these datasets, the 3DScan dataset records RGBD videos of different categories of objects. However, most objects are heavily occluded by hands or only partially visible with large objects like cars. The diversity of objects and scenes is also relatively low given limited human labor. Thus 3DScan is not suitable for pose estimation tasks. In this paper, we propose Wild6D, a dataset with RGBD videos containing diverse objects taken with diverse backgrounds. While the training set is not labeled, we provide 6D pose annotations for the test videos. To the best of our knowledge, Wild6D is the largest RGBD dataset for 6D object pose estimation in the wild.
Rendering in Object Pose Estimation. One effective refinement method for 6D object pose estimation is via rendering .
For instance, Iwase et al. propose to utilize differentiable rendering and learn the texture of a 3D model. However, it still depends on a pre-defined CAD model with a high-quality texture map for instance-level pose estimation and it is not feasible for category-level. Additionally, object pose estimation can be conducted via Analysis-by-Synthesis , but they usually require gradient descent for optimization leading to relatively slower inference speed. Critically, these works only focus on the instance-level setting and none of them can generalize to different instances. Inspired by recent work on 3D reconstruction , our RePoNet learns the object shape and pose (including both NOCS map and R/T/S parameters) simultaneously using differentiable rendering to provide a loss function during learning , which allows the gradients to backprop through the whole network for training.
While Manhardt et al. also utilize differentiable rendering to improve category-level 6D pose, it is only applied in test time to adjust the pose, instead of training the network end-to-end. We show substantial improvement over all baselines training with RePoNet.
Wild6D Dataset
Existing Category-level 6D Object Pose Datasets. NOCS is the most common dataset for category-level 6D object pose estimation, which consists of the synthetic CAMERA25 dataset and the real-world REAL275 dataset. The synthetic CAMERA25 dataset contains 300,000 RGBD images of 1,085 object instances from 6 categories for training and evaluation. The REAL275 dataset shares the same categories with CAMERA25 but is more challenging under the real-world background. It contains 4,300 RGBD images of 7 scenes for training and 2,750 images of 6 scenes for testing. Due to the limited scale and diversity of real data, the learned models from this dataset cannot generalize to in-the-wild scenarios. Objectron is a large-scale dataset for object pose estimation and tracking. Different from NOCS , it consists of a collection of short object-centric video clips captured in the real world. However, no depth maps are provided in Objectron, which might lead to ambiguities in pose estimation.
Wild6D Collection. To achieve a categorical 6D pose for real objects, we collect a new large-scale RGBD dataset, named Wild6D. Each video in the Wild6D is recorded via the iPhone front camera showing multiple views of objects where RGB images and the corresponding depth images and point cloud are captured simultaneously. The videos are captured by different turkers with their own iPhones to guarantee the diversity of instances and background scenes. Three videos are taken for each object under different scenes. In total, Wild6D consists of 5,166 videos (1.1 million images) over 1722 different object instances and 5 categories, i.e., bottle, bowl, camera, laptop, and mug. Among this data, we split 486 videos of 162 instances to use them as the test set. Table 1 summarizes the statistics of Wild6D comparing previous datasets: Wild6D significantly improves the number of images, object instances, and scene complexity.
Wild6D Annotation. To annotate more than 10,000 images in the testing set efficiently, we propose a tracking-based annotation pipeline. Inside a video, we manually annotate the 6D object poses every 50 frames as keyframes. Given the annotation of the keyframe, we implement TEASER++ together with colored ICP to achieve the registration between the keyframe and the following frame and compute the transformation matrix. The ground-truth object pose of the following frame can be obtained by applying the transformation matrix to the keyframe annotation. Following this pipeline, we can obtain accurate ground-truths by only labeling around 5 keyframes per video.
Proposed Method
We propose the Rendering for Pose estimation network (RePoNet) using both synthetic data and large-scale unlabeled real-world data in a semi-supervised manner.
RePoNet overview. The RePoNet is composed of two networks including the Pose Network to directly estimate the object 6D pose parameters (with the NOCS map as an intermediate representation) and the Shape Network to reconstruct the object shape. The outputs from both networks can go through a differentiable rendering module which outputs a segmentation mask.
One important contribution of this framework is to make the whole procedure with RePoNet differentiable, using the Shape Network, and a ConvNet to connect NOCS map to 6D pose parameters in the Pose Network. This allows the gradients to backprop through the network end-to-end. We will explain the architecture and objective details in the following subsections.
Inference for 6D pose. For 6D object pose estimation, only the NOCS map output of Pose Network is required for inference. The object pose can be computed by solving the Umeyama algorithm with the NOCS map. The 6D pose parameters and are only used to perform differentiable rendering during training. The reason is that the directly estimated pose parameters from conv module is accurate when fitting the training data, but not as accurate as using NOCS map when generalizing to novel test instances.
1.2 Shape Network.
The Shape Network (top part of Fig. 3) aims to reconstruct the 3D shape of the input object from the given shape prior via mesh deformation. The major usage of this network is for performing differentiable rendering to provide training signals. The Shape Network is not necessary during inference time for 6D object pose estimation.
1.3 Differentiable rendering.
2 Learning Objectives
We define multiple objectives for training the RePoNet, including the loss for 6D pose parameters, NOCS regression loss, shape reconstruction loss using synthetic ground-truths for supervision, and a silhouette matching loss which can be applied to both synthetic data and real data without 6D pose annotations. We introduce each loss as the following.
Disentangled 6D Pose Loss. Inspired by , we implement disentangled pose loss via individually supervising the rotation , translation and scale . Instead of directly computing parametric distances based on rotation matrix, we employ the variant of Point-Matching lossWe predict the quaternion representation of rotation matrix and follow the strategy in to deal with the symmetry of the object. . Additionally, we decouple the translation into the 2D location of the 3D centroid projection on the image plane and the object’s distance to camera . can be approximated as the bounding box center of the given object . Given the camera intrinsics , the translation can be calculated via Eq. 3,
Therefore, given the ground-truth pose annotations , the objective function of 6D pose is formulated as Eq. 4,
Therefore, for the fully-supervised training with annotated synthetic data, the total loss function is described in Eq. 8, where the s are balance parameters.
Meanwhile, under the semi-supervised setting with both annotated synthetic data and unlabeled data, the total loss function is then formulated as:
Experiments
Semi-supervised setting. We use the training data of CAMERA25 along with the corresponding annotations and jointly train the model with images of REAL275 or Wild6D without any 6D pose annotations. More configurations of RePoNet have been specified in supplementary materials. To validate the effectiveness, we conduct experiments on both REAL275 and Wild6D.
Fully-supervised setting. We share the same architecture and configurations with semi-supervised training but utilize all annotations of CAMERA25 and REAL275 .
2 Ablation Study
3 Comparison with State-of-the-art Methods
Performance on REAL275. We split all existing methods into two groups by whether using full annotations of real data and compare their performance with RePoNet under both fully-supervised setting and semi-supervised setting, denoted as “RePoNet-sup" and “RePoNet-semi" respectively. As listed in Table 5, RePoNet-sup can achieve competitive results among all existing approaches under the fully-supervised setting. Note our method is complimentary to the techniques proposed in SGPA, and our key contribution and focus lies on the semi-supervised counterpart for generalization. As shown in the bottom part of Table 5, the RePoNet-semi outperforms CPS++ by a large margin under the semi-supervised setting In Table 5, CPS++ only reports the performance under 10 degree, 10 cm.. Moreover, compared with the existing fully-supervised methods, RePoNet-semi achieves a comparable or even better performance without any 6D pose annotations. Additionally, our approach can perform shape reconstruction, but it’s not our goal and we report its performance in the supplementary material.
Performance on Wild6D. Finally, we evaluate the performance of RePoNet on the Wild6D testing set as reported in Table 6. For some existing works, we directly test their pre-trained models trained on CAMERA75 along with REAL275. Since FS-Net does not release the model, we cannot experiment on it. It can be observed that the pre-trained models cannot generalize to Wild6D due to the limited diversity of real data during training. For example, the performance of Shape-Prior on three pose estimation metrics are all lower than 15%. On the other hand, the RePoNet with semi-supervised learning on Wild6D achieves 29.5% and 34.4% on 5 degree, 2cm and 5 degree, 5cm, which is better than Shape-Prior . This significant improvement shows the better generalization ability of RePoNet. Also, the superior performance of RePoNet-semi over RePoNet-syn shows our proposed semi-supervised training can effectively leverage the in-the-wild data.
Discussion
Conclusion. In this work, we consider the problem of generalizing 6D pose estimation to in-the-wild objects. Most of existing work are restricted in constrained environments due to limited number of annotated data. In an effort to resolve this limitation, we collect Wild6D, a new RGBD dataset with diverse instances and backgrounds. Instead of annotating them, we also propose RePoNet, a network that can leverage both the synthetic data and real-world RGBD data. Without using any 3D annotations on real-world data, RePoNet outperforms state-of-the-art methods on existing datasets and Wild6D test set by a large margin.
Limitations. Our proposed RePoNet may be hard to generalize to unseen categories. This is because RePoNet depends on the categorical mesh prior which is pre-defined for each category leading to the learnt deformation model fixed to a specific category. Alternative way to address this problem is to represent object shape via a deformation of a unit sphere to get rid of the category-level prior.
References
Appendix
Appendix A Discussion on Wild6D.
Statistics Analysis. We create a new large-scale benchmark for category-level object pose estimation called Wild6D, which consists of 5,166 videos (1.1 million images) over 1722 different object instances and 5 categories, i.e., bottle, bowl, camera, laptop, and mug. The distribution of the over objects per category in Wild6D is illustrated in Table 7.
Mask Segmentation Quality. Although the mask segmentation quality may influence the training process of RePoNet, i.e., silhouette matching loss, we found in most cases the mask segmentations are satisfactory and good enough to conduct the silhouette matching loss. The mask segmentation results of several samples are shown in Fig.4. We provide the mask segmentation along with the raw RGB image and depth image for each video frame.
Baselines on Wild6D. In original paper, we evaluate several existing work on Wild6D testing set. More specifically, we utilize the official released model trained on NOCS CAMERA25 and REAL275 training set in a fully-supervised manner to estimate the 6D pose for each unique object in Wild6D. Comparing with most existing work, RePoNet not only has a better generalization ability, but also leverages the in-the-wild data effectively.
Appendix B Architecture Details
As describe in our paper, the RePoNet has two parallel branches: Pose Network and Shape Network. The detailed architecture of each network are illustrated in Fig. 5.
Semi-supervised setting. We use the training data of CAMERA25 along with the corresponding annotations and jointly train the model with images of REAL275 or Wild6D without any 6D pose annotations. After cropping the object from the RGBD image, we first resize it to and then randomly sample 1,024 points from both color image and depth map. To obtain the categorical shape prior, we choose a CAD model per category from the CAMERA25 training set manually and reduce its number of vertices to 1,024 as well. More configurations of RePoNet have been specified in supplementary materials. We adopt Adam to optimize our model with the initial learning rate of 0.0001. The learning rate is halved every 10 epochs until convergence. We empirically set the balance parameters and to and , respectively.
Appendix C Implementation Details
Appendix D More experiments
We analyze the effect of using different fractions of unlabeled real data used during semi-supervised learning. We uniformly sample every 10% fraction of collected Wild6D training data for semi-supervised learning and evaluate the performance on REAL275 and Wild6D testing sets. As shown in Fig. 6, with more real data used during training, the object pose estimation performance is getting better.
D.2 Shape reconstruction
To evaluate the shape reconstruction performance, we computed the Chamfer Distance of the reconstructed object mesh with the ground truth one and compared it with other methods. Since the Wild6D does not provide the CAD models, we just conduct this experiment on REAL275 . From Table 8, we observe that the average distance over six categories of our proposed method is much lower than Shape-Prior and SGPA under both fully-supervised setting and semi-supervised setting. However, it is worse than CASS , especially on camera category. We believe this is because our method deforms the object shape from the predefined mesh and the shape variance across different cameras is large which may degenerate the performance.
D.3 Semi-supervised on CO3D
As discussed in Related Work, the recent proposed CO3D , although with diverse instances in different categories, is difficult to be used for 6D pose estimation due to the error of depth maps predicted via COLMAP . To validate this point, we compare the RePoNet model trained with CAMERA25 and CO3D with the model trained with CAMERA25 data and Wild6D in Table 9. There is no large performance improvement observed by involving the CO3D data, although it also provides hundreds of object-centric real videos. On the other hand, by using the Wild6D data, the pose estimation performance improves a lot. For instance, the improvement on degree, cm is versus . Hence, our collected Wild6D data with real RGBD images is much more suitable and feasible for 6D pose estimation than existing datasets.
D.4 More visualizations
We show more qualitative results of our proposed RePoNet model on the Wild6D in Fig. 7 and Fig. 8. It can be observed that our proposed RePoNet can estimate object pose and size accurately across diverse instances and under d different background scenes.