FS6D: Few-Shot 6D Pose Estimation of Novel Objects

Yisheng He, Yao Wang, Haoqiang Fan, Jian Sun, Qifeng Chen

Introduction

6D object pose estimation aims to predict a rigid transformation from the object coordinate system to the camera coordinate system, which benefits various applications, including robotic manipulation, augmented reality, autonomous driving, etc. The explosive development of deep learning has brought significant improvement to this problem. With recent works reaching nearly 99% recall accuracy on existing benchmarks , one may get the impression that the 6D object pose problem has been solved, which is not the case. We argue that the current problem has been simplified with strict restrictions. They are under the close-set assumption that the training and testing data are drawn from the same object space, which, however, does not adhere to the real dynamic worlds. Moreover, extravagant high-fidelity CAD models and large-scale datasets are required for training to obtain good performance on new objects under the current instance-level pose estimation setting.

The recently proposed category-level pose estimation task loosens the restriction with generalizability to novel objects within the same categories. However, it is still limited in the close-set assumption of predefined categories. Instead, in this work, we study a new open-set problem, the few-shot 6D object pose estimation: estimating 6D pose of unknown objects by only a few views of the objects without extra training. As shown in Figure 1, in our setting, only a few labeled RGBD images of novel objects are provided, and no high-fidelity CAD models are required. The goal of the problem is to bridge the capability gap between machine learning algorithms and flexible human visual systems that can locate and estimate the pose of a novel object given only several views of it. Besides, it has a wide range of real-world applications in robotic vision systems, i.e., fast registration of novel objects for robotic manipulation and home robots.

Under the observation that human beings utilize both appearance and geometric information to match and locate a new object, we propose a dense RGBD prototypes matching framework to tackle the problem. Specifically, transformers are utilized to fully explore the semantic and geometric relationship between the query scene patch and the support views of novel objects. Moreover, we point out that large-scale datasets’ diverse shape and appearance priors are essential to empower networks to generalize on novel objects. Therefore, we introduce a large-scale photorealistic dataset (ShapeNet6D) with diverse shapes and appearances for prior learning. To our knowledge, ours (800K images of 12K objects) is the largest and most diverse dataset for 6D pose algorithms. To bridge the domain gap between rendered RGB images and real-world scenes, we introduce a simple and effective online texture blending augmentation, which further enriches the appearance diversity and facilitates network performance at a low cost.

To summarize, the contributions of this work are:

We introduce a challenging open-set problem, the few-shot 6D object pose estimation, and establish a benchmark to study it.

We formulate the problem by dense RGBD prototypes matching and introduce FS6D-DPM, which fully leverage appearance and geometric information to tackle the problem.

Datasets: We introduce ShapeNet6D, a large-scale photorealistic dataset with diverse shapes and appearances for prior learning of few-shot 6D pose estimation algorithms. We also introduce an online texture blending augmentation to obtain scenes of texture-rich objects without domain gaps at a low cost.

Related Work

Instance-level pose estimation retrieves pose parameters of known object instances. Matching based approaches requires precise CAD models to render thousands of templates and establish hand-craft or learned codebook for matching. Learning-based approaches includes direct pose regression , dense correspondence exploration and recent keypoint-based approaches , which improve the performance by large margins. Despite compelling results, these approaches can only deal with scenarios of known objects with high-fidelity CAD models. Instead, the recent category-level pose estimation improves the generalizability by estimating unseen object instances within the known categories. Normalized Object Coordinate Space (NOCS) or shape deformation based approaches are proposed. However, both traditional instance- and category-level pose estimation problems are under the close-set setting, assuming that the training and testing data are within the same predefined instance or category spaces. While such close-set setting does not adhere to the real dynamic world, we instead define a new open-set problem, the few-shot 6D pose estimation. Algorithms developed in our open-set setting can be flexibly applied to unknown objects without extra training with only a few labeled RGBD images, no matter they are within the trained categories or not.

2 Possible Few-Shot Pose Estimation Solutions

Local Image Feature Matching. Local feature matching can establish the correspondence between two images for the few-shot pose estimation problem. Existing methods can be categorized into detector-based and detector-free . While these algorithms only leverage the grey-scale images, the performance drops on texture-less objects. Instead, we fully leverage both the appearance and the geometric information and generalize well in more scenarios.

Point Cloud Registration. One line of point cloud registration algorithms solve the problem by detecting 3D keypoints , extracting feature descriptors and estimating the relative transformation. Several end-to-end approaches are also proposed. However, these algorithms heavily rely on fine point clouds and fail on objects that are not captured by depth sensors, i.e., reflective ones. Instead, we fully leverage the complementary information in RGBD images for dense prototypes extraction and matching to retrieve better object pose parameters.

3 Metric learning on few-shot learning problems

Metric learning techniques have been applied to several few-shot learning problems, including classification and segmentation . The representative prototypical network for classification map the support and query images into a global embedding space and then retrieve the class label of query image based on the support embedding, named prototype. The recent metric learning-based approaches in more challenging segmentation areas utilize similar technique but output per-pixel prediction on the query images by matching per-pixel query features with global average prototypes or part-level prototypes . While sparse support prototypes are enough to solve the above problems, few-shot pose estimation requires more dense correspondence exploration on pixel-level support prototypes and query features, which is more challenging.

Proposed Method

We introduce the problem setting of the few-shot 6D object pose estimation and the derived domain generalization problem.

The few-shot 6D object pose estimation. We formulate the new open-set task, the few-shot 6D pose estimation as follows. Given kk support RGBD patches P={p1,p2,...,pk}P=\{p_{1},p_{2},...,p_{k}\} of a novel object with pose parameters as support frames, the inference task is to retrieve the 6D pose parameters of that novel object in the query novel scene image II. Compared to current close-set setting, the proposed open-set one eliminates the reliance on precise CAD models and focuses on the generalizability of trained models on unseen objects. Specifically, once the model is trained, we expect to apply it on novel scenes of novel objects by a few views without extra training. It bridges the gap between machine learning algorithms and flexible human visual systems. Moreover, it enables real-world applications, i.e., fast registration of new objects for robotic manipulation and service home robots.

The generalization requirement of the open-set problem also derive another interesting research question to study:

Domain generalization. The domain generalization targets to reduce the domain gaps between models trained on the synthesis and real-world data. It has been introduced to the 6D pose estimation field to deal with the lack of data . However, this field is less explored as existing real-world benchmarking datasets for the close-set problem has been well established: real-world training data for the object to be estimated is available. While existing datasets are small with limited objects, in our few-shot open-set setting, diversity of shape and appearance are crucial to the generalizability of few-shot 6D object pose estimation algorithms. However, capturing and labeling such a large-scale real-world dataset is not practical due to the high cost (money and time). It is crucial to fully leverage the geometry and appearance diversity in our large-scale photorealistic datasets and generalize to the real world. The domain generalization problem is thus an important problem to study for the few-shot 6D pose estimation.

2 Datasets

The prior learned from large-scale datasets is crucial to the performance and generalizability of few-shot learning algorithms. ImageNet , for example, has been widely used for network pre-training in several few-shot learning tasks, i.e., object detection and segmentation. While 2D vision tasks rely more on the semantic prior in RGB images, for the few-shot 6D object pose estimation, both shape and semantic prior are crucial for the generalizability of the network. However, existing datasets for 6D object pose estimation are small and lack diversity in shape and appearance to provide enough prior for the generalization capability. Therefore, we keep their role as real-world benchmark datasets and propose a new large-scale dataset, ShapeNet6D, with diverse shapes and appearances for prior learning.

The proposed ShapeNet6D is a large-scale photorealistic dataset containing RGBD scene images of more than 12K object instances from the ShapeNet repository. Each scene image is labeled with ground truth information for the 6D pose estimation problem, including instance semantic segmentation and pose parameters of each object. As we demonstrate empirically, the diversity of shape and appearance is crucial for the network to generalize. While it is not practical to collect and label such a large-scale, diverse dataset in the real world due to the high cost (time and money), we instead generate photorealistic images by physically-based rendering. Our approach is inspired by the successful application of photorealistic datasets in while improving the diversity of object shape and appearance. Specifically, we utilize the physically-based rendering engine, Blenderhttps://www.blender.org that simulates the flow of light energy by ray tracing to render realistic scene images. To arrange a scene to render, we first randomly select several objects from ShapeNet, apply random material and texture, and drop them into a box with the PyBullet physics engine integrated into Blender. To enrich the variety of the background, we randomly selected physically-based rendering material from the HDRI Havenhttps://hdrihaven.com/hdris and applied them to the wall of the box. Random environment lights are also added to generate diverse lighting conditions. Finally, the RGBD scene image is rendered from a random camera pose, and the ground truth instance semantic segmentation labels and pose parameters of each object are also obtained. Statistics about ShapeNet6D compared to existing 6D pose benchmark datasets are shown in Table 1. ShapeNet6D is on a larger scale and is more diverse in shape and appearance, which provides better prior to the few-shot pose estimation problem as we showed empirically.

2.2 Online texture blending

As one of the crucial clues to solve the few-shot 6D pose estimation problem, the texture field is also essential to the performance of the few-shot 6D object pose estimation. However, it is labor-intensive and time-consuming to generate textures and materials for objects that can be rendered to be photorealistic. The rendered RGB images tend to have more significant domain gaps between the real world as well. Moreover, to produce photorealistic images, time- and computation-consuming techniques like ray tracing are required. Therefore, the images should be pre-processed offline and stored before network training, which costs a lot of storage space for a large-scale dataset. On the other hand, real-world RGB images captured from various cameras are easy to access, i.e., ImageNet , and MS-COCO . It motivates us to leverage efficient texture wrapping techniques to generate scenes of objects with rich real-world texture to serve as online data argumentation. Specifically, the mesh is first unwrapped to obtain a UV map. For each triangle, we get the UV coordinate of each vertex and then utilize it to determine the UV coordinate of each pixel by linear interpolation during rasterization. The UV coordinate is then applied to lookup the color value from a texture map randomly sampled from the real-world ImageNet , and MS-COCO . Previous works render images with artificial simulations, i.e., Beckmann model , which change the domain and cause domain gaps. Instead, we applied no simulation, so the composite images are kept in the real domain, i.e., the lighting condition, sensors noise of the real-world images are preserved. Moreover, such a simple blending strategy can be implemented fast to serve online. Moreover, we can combine it with online shape deformation to produce data with rich appearance and shape diversity for training, as shown in Figure 2.

3 FS6D-DPM

Prototypes-based few-shot learning. We first briefly introduce the prototypes-based algorithms for few-shot learning. It has been successfully applied to various few-shot 2D vision tasks, i.e., classification and semantic segmentation. Specifically, a pre-trained Siamese backbone is utilized for feature extraction from the support and the query images. Then, global average pooling is applied on the extracted support feature maps to obtain the support prototypes. This global average prototype is then applied to calculate the similarity between the global features (in classification) or dense pixel-wise features (in semantic segmentation) extracted from the query image for prediction. However, these tasks’ global-to-global or global-to-local correspondence is not enough to recover 6D object pose parameters. This work, instead, proposes a dense prototypes extraction module to establish the local-to-local correspondence between the support RGBD images and the query scene patch for pose estimation.

Transformer . Transformers networks are first introduced in Natural Language Processing and are brought into many vision tasks. The multi-head attention mechanism enables it to capture the long-term dependency even on an unordered set. Specifically, given three vectors as inputs, namely query QQ, key KK, and value VV. The attention mechanism is to retrieve information II from the value s.t. the similarity between QQ and KK, denoted as:

Gifted with the capability of capturing long-term dependency, the Transformer networks have been successfully applied to aggregate contextual information in the local feature matching and point cloud registration field. In this work, we further extend it to dense RGBD prototypes matching for few-shot 6D pose estimation.

3.2 Overview

To build a few-shot pose estimation algorithm that can generalize well to novel objects, it is crucial to fully explore the semantic and geometric relationship between the given support views and the query scene patch, as shown in Figure 4. In this section, we introduce our dense prototypes matching framework to tackle this challenging problem. As shown in Figure 3, our framework consists of three main parts. Firstly, a Siamese RGBD feature extraction backbone is utilized to extract rich semantic and geometric features for each pixel/point. Then, a dense prototypes extraction network based on transformers is applied to extract dense RGBD prototypes from the support view and point-wise local features from the query scene patch for similarity calculation. Finally, after the correspondence between dense prototypes and scene features is established, the Umeyama algorithm is leveraged to estimate the 6D pose parameters.

3.3 Feature Extraction Backbone

The first step is to extract rich semantic and geometric features from the given RGBD images. As a fundamental problem, many works have studied this representation learning task. Recently, FFB6D introduce a full flow bidirectional fusion network for 6D pose estimation and significantly improve the performance of close-set pose estimation. Specifically, bidirectional local feature fusion blocks are added into each encoding and decoding layer to bridge the information gap and improve the quality of extracted semantic and geometric features (see for details). In this work, we leverage FFB6D to build a Siamese network for feature extraction from the support images and the query scenes.

3.4 Dense Prototypes Extraction and Matching

Now we have obtained dense features from the Siamese feature extraction backbone. We then extract dense support prototypes and query features to calculate the similarity and establish the correspondence. To extract descriptive and representative dense RGBD prototypes from the support views and dense query features from query scenes, it is crucial to fully leverage the structural geometric information residing in point clouds and semantic information abiding in RGB images. Besides, contextual information between the support shot and the query patch is also essential to improve the precision of similarity calculation and correspondence exploration.

Considering the power of transformers on long-term dependency capturing, we utilize the optimized Linear Transformers to serve the above two purposes. As shown in the middle part of Figure 3, we first establish self-attention on the extracted feature maps to strengthen the geometric and semantic information residing in the extracted dense prototypes and dense query features. We regard the extracted features as query, key, and value and fed them into the Linear Transformer networks to enhance the semantic and geometric features. Meanwhile, a cross-attention module is also applied to explore the contextual information between the support prototypes and the query scene features. Precisely, to extract contextual information from the support prototypes to the query scene features, we took each scene feature as a query and the dense prototypes as keys and values to the Linear Transformers. Contextual information from query scene features to support prototypes is enhanced similarly. With extracted contextual information, another self-attention modules are applied to enhance the geometric and semantic features further. In this way, we obtain dense support prototypes and query features with rich semantic, geometric and contextual information. Unlike prototype-based few-shot classification and segmentation algorithms that calculate the similarity by cosine distance, we follow local feature matching pipelines to establish the dense correspondence by calculating C(i,j)=⟨P(i),Q(j)⟩C(i,j)=\langle P(i),Q(j)\rangle with P(i)P(i) the ithi_{th} prototype, Q(j)Q(j) the jthj_{th} query feature and ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle the inner product. The Sinkhorn Algorithm is applied for differentiable optimization as well.

3.5 Pose Parameters Estimation

After the correspondence between the dense prototypes and the query scene features is established, we utilize the Umeyama algorithms to recover the pose parameters. Specifically, given a set of matched pairs M={(pi,qi),1≤i≤N}\mathcal{M}=\{(p_{i},q_{i}),1\leq i\leq N\} with pip_{i}, qiq_{i} the 3D coordinate of matched prototypes and queries, the Umeyama algorithms estimate the rotation RR and translation TT by minimizing:

To eliminate the influence of outliers. The RANSAC algorithms are also applied.

Given KK support views of a novel object, we can obtain KK predicted pose parameters along with their losses. We select the one with minimum loss as our final prediction.

Experiments

The LineMOD and the YCB-Video are two popular datasets for 6D object pose estimation. The LineMOD dataset contains 13 videos of 13 low-textured objects, while the YCB-Video dataset consists of 92 RGBD videos of 21 YCB objects. For the few-shot pose estimation problem, we select 16 shots for each object for pose estimation. We also follow the strategy of other well-established few-shot problems, i.e., segmentation, and split the dataset into different groups. Specifically, we split the objects into three groups for each dataset and select one for testing and the remaining two for training each time (see the supplementary material for details).

2 Evaluation Metrics

The average distance metrics ADD and ADDS are widely used for performance evaluation of 6D pose estimation. For an object O\mathcal{O} consists of vertexes vv, the ADD of asymmetric objects with the predicted pose RR, TT and ground truth pose R∗R^{*}, T∗T^{*} is calculated by:

For symmetric objects, the ADDS based on the closest point distance is defined as:

In the YCB-Video dataset, the area under the accuracy-threshold curve obtained by varying the distance threshold (ADDS and ADD AUC) is reported following . In the LineMOD datasets, we report the distance less than 10% objects diameter recall (ADD-0.1d) as in .

3 Baselines

Possible solutions to the few-shot 6D object pose estimation problem include local image feature matching, point cloud registration, and template matching. We select the state-of-the-art solution in each direction as our baseline.

LoFTR is a detector-free deep learning architecture for local image feature matching. It uses the self- and cross-attention layers in Transformers to obtain high-quality matches.

PREDATOR is a neural architecture for pairwise 3D point cloud registration with deep attention to the overlap region. It learns to detect the overlap region between two unregistered scans and focus on that region when sampling feature points.

Template Matching. Template matching approaches discrete pose estimation problem into classification problem. These approaches rely on CAD models to generate thousands of templates and retrieve the closest one to the scene. However, we eliminate the dependency of precise object CAD models in our problem. Besides, capturing, labeling, and storing thousands of support shots are time- and storage-consuming. We assign the view with rotation closest to the ground truth and the center shift as translation to reveal the upper bound of these approaches.

For a fair comparison, all baselines and the proposed one are not equipped with iterative refinement, e.g., ICP .

4 Training and Implementation

We crop object patches with ground-truth bounding boxes for our model and resize them to 255×255255\times 255 as input. The correspondence is optimized by negative log-likelihood loss . For a fair comparison, we pretrained all models on ShapeNet6D with online data augmentation for two epochs and fine-tuned on benchmark datasets for five epochs. We select 16 different views for each object as support images.

5 Benchmark Results

Results on LineMOD and YCB-Video datasets. Quantitative results on the YCB-Video and the LineMOD dataset are shown in Table 2 and Table 3 respectively. Thanks to the joint reasoning of appearance and geometric relationship between the support and query images, our method outperforms the state-of-the-art local image feature matching method and point cloud registration algorithms by large margins. Some qualitative results are shown in Figure 5.

Domain generalization. As is shown in Table 5, our model trained on ShapeNet6D with online data augmentation is 4.1%4.1\% behind the fine-tuned one. Considering the small shape and appearance diversity in the LineMOD dataset, compared with ShapeNet6D, we think the performance drop mainly comes from the domain gap. More future works are expected to bridge this gap to fully explore the power of shape and appearance diversity in ShapeNet6D, e.g., designing domain invariant algorithms.

6 Ablation Study

Effect of pre-training on the large-scale ShapeNet6D. As shown in Table 5, FS6D-DPM trained on ShapeNet6D outperforms the one trained from scratch on the LineMOD dataset by a large margin (+11%+11\%), proving the efficacy of the shape and appearance diversity resides in ShapeNet6D.

Effect of online texture blending. As shown in Table 4, the proposed online texture blending provides diverse texture prior and improves the performance on texture-rich objects in the YCB-Video dataset by large margins.

Discussion and Limitations

In this work, we study a challenging open-set problem, the few-shot 6D object pose estimation. We point out the essence of appearance and geometric information to tackle the problem and propose FS6D-DPM as a solid baseline to solve it. Furthermore, we show that prior from diverse shapes and appearances are crucial to the generalizability of few-shot 6D pose estimation algorithms and introduce a large-scale dataset (ShapeNet6D) for network pre-training. An online texture blending augmentation is proposed to bridge the domain gap as well.

However, there are still some limitations in this work. Firstly, we focus on the pose estimation problem and rely on object detection algorithms to crop out the region of interested objects. Though various off-the-shelf few-shot object detection algorithms are available, a joint framework is more practical. Secondly, despite being diverse in shape and appearance, the proposed large-scale ShapeNet6D is synthesis, and the domain gaps problem is not tackled yet. Future directions include domain invariant pose estimation algorithms or large-scale real-world datasets. Lastly, there is still a significant performance gap between few-shot algorithms and those trained under the close-set setting. We expect more future research, e.g., leveraging 3D keypoint-based techniques to bridge this gap.

Acknowledgements This work is supported by Guangzhou Okay Information Technology with the project GZETDZ18EG05.

Appendix A Appendix

Model running time. We run our model on a single NVIDIA GeForce RTX 2080Ti GPU. For each support view, it takes an average time of 72 ms for neural network forwarding and 113 ms for pose alignment using the Umeyama algorithm with RANSAC .

Pose estimation with ground-truth segmentation. In the main paper, we utilize relaxed ground-truth object bounding boxes to crop out regions of interested objects from the query scene for pose estimation. While LatentFusion utilizes stricter ground-truth segmentation to segment out objects, we report our results on LineMOD dataset following their setting. Specifically, We use our model trained only with ShapeNet6D without fine-tuning on the real LineMOD dataset. As is shown in Table 6, our model without any refinement already surpasses iterative refined LatentFusion. Equipped with post-refinement by ICP, our model obtains further improvement. Moreover, our model (0.34 fps) is 18X faster than LantentFusion (0.018 fps) on a RTX 2080Ti, when both use 16 support views.

Effect of the different number of support views. We ablate the effect of the different number of support views in Table 8. As is shown in the table, our algorithm gets better performances when the number of support views increases. Moreover, it only gains margin performance when we have more than 16 views, which shows that our algorithm does not need too many support views and can get good pose results under the few-shot setting.

Details results on the LineMOD dataset. See Table 7.

Visualization of ShapeNet6D Example images in ShapeNet6D are shown in Fig. 6.

A.2 Implementation Details

Grouping information of benchmark datasets. We split the LineMOD dataset into three groups. Objects in different groups have no intersection. During network fine-tuning, we select two groups for training and one group as novel objects for testing. The group information of the LineMOD dataset is shown in Table 10. We split the YCB-Video dataset into three groups in a similar way. Group information of the YCB-Video dataset is shown in Table 9.

Support views selection. We select support views from the training set since we do not have the real-world objects in the LineMOD and YCB-Video datasets to capture the support views. We select 16 support views using the farthest rotation sampling for each object to ensure that each part of the object is visible. Specifically, we initialize the set of selected views with a random view from the training set for each object. We then add another object view with the farthest rotation distance from views in the selected set. We repeat this procedure until 16 views of the target object are obtained. We define the distance between two rotations as the Euclidean distance between two unit quaternions following . The formula is as:

where ∣∣⋅∣∣||\cdot|| denotes the Euclidean norm and q1,q2q_{1},q_{2} the two unit quaternions.

Given the target object’s mask labels and pose parameters in the selected support views, we crop out the object region and transform the object point cloud back to the object coordinate system to serve as a reference frame to define the 6D object pose.

A.3 Fast Registration of Novel Objects

Given a novel object and an RGBD sensor with known intrinsic parameters, we can quickly obtain support views of the novel object in several ways. We provide some examples as follows:

Select from an RGBD video of the novel object. The most simple way is to select support views from an RGBD video of the target object. Specifically, we first place the target object in the center of a clean plane and then capture a video by slowly moving the camera around the object. We use the first frame to define the object coordinate system. Specifically, we mask out the object region by removing the background plane with a plane detection algorithm or least-square-fitting of a plane on the scene point cloud. We define the object coordinate system based on the object point cloud of the first frame. Then, we calculate the pose between the following frame and the first frame. Since the pose difference between adjacent frames of a video is small and the scene background is a clean plane, we can utilize registration algorithms, i.e., ICP , Go-ICP to calculate the relative pose parameters between adjacent frames and obtain the pose parameters between each frame and the first frame. Finally, we can select support views by the farthest rotation sampling algorithm as in Section A.2.

To further improve the accuracy of relative pose parameters, we can put the object on a marker board (a plane with several markers on it) and utilize markers to obtain more accurate relative poses.

Collect with the robot arm. For robotic manipulation, we have a robot arm with a camera in hand. We first calibrate the robot arm and the camera between an observed region with a marker board. We define several viewing points with known pose parameters. We then place the novel target object to the observed region and utilize the robot arm to move the camera to those predefined viewing points to capture support views of the novel objects. The pose parameters of support views will be more accurate due to the robustness of the robotic manipulation system.

References