BOP: Benchmark for 6D Object Pose Estimation
Tomas Hodan, Frank Michel, Eric Brachmann, Wadim Kehl, Anders Glent Buch, Dirk Kraft, Bertram Drost, Joel Vidal, Stephan Ihrke, Xenophon Zabulis, Caner Sahin, Fabian Manhardt, Federico Tombari, Tae-Kyun Kim, Jiri Matas, Carsten Rother
Introduction
Estimating the 6D pose, i.e. 3D translation and 3D rotation, of a rigid object has become an accessible task with the introduction of consumer-grade RGB-D sensors. An accurate, fast and robust method that solves this task will have a big impact in application fields such as robotics or augmented reality.
Many methods for 6D object pose estimation have been published recently, e.g. , but it is unclear which methods perform well and in which scenarios. The most commonly used dataset for evaluation was created by Hinterstoisser et al. , which was not intended as a general benchmark and has several limitations: the lighting conditions are constant and the objects are easy to distinguish, unoccluded and located around the image center. Since then, some of the limitations have been addressed. Brachmann et al. added ground-truth annotation for occluded objects in the dataset of . Hodaň et al. created a dataset that features industry-relevant objects with symmetries and similarities, and Drost et al. introduced a dataset containing objects with reflective surfaces. However, the datasets have different formats and no standard evaluation methodology has emerged. New methods are usually compared with only a few competitors on a small subset of datasets.
This work makes the following contributions:
Eight datasets in a unified format, including two new datasets focusing on varying lighting conditions, are made available (Fig. 1). The datasets contain: i) texture-mapped 3D models of 89 objects with a wide range of sizes, shapes and reflectance properties, ii) 277K training RGB-D images showing isolated objects from different viewpoints, and iii) 62K test RGB-D images of scenes with graded complexity. High-quality ground-truth 6D poses of the modeled objects are provided for all images.
An evaluation methodology based on that includes the formulation of an industry-relevant task, and a pose-error function which deals well with pose ambiguity of symmetric or partially occluded objects, in contrast to the commonly used function by Hinterstoisser et al. .
A comprehensive evaluation of 15 methods on the benchmark datasets using the proposed evaluation methodology. We provide an analysis of the results, report the state of the art, and identify open problems.
An online evaluation system at bop.felk.cvut.cz that allows for continuous submission of new results and provides up-to-date leaderboards.
The progress of research in computer vision has been strongly influenced by challenges and benchmarks, which enable to evaluate and compare methods and better understand their limitations. The Middlebury benchmark for depth from stereo and optical flow estimation was one of the first that gained large attention. The PASCAL VOC challenge , based on a photo collection from the internet, was the first to standardize the evaluation of object detection and image classification. It was followed by the ImageNet challenge , which has been running for eight years, starting in 2010, and has pushed image classification methods to new levels of accuracy. The key was a large-scale dataset that enabled training of deep neural networks, which then quickly became a game-changer for many other tasks . With increasing maturity of computer vision methods, recent benchmarks moved to real-world scenarios. A great example is the KITTI benchmark focusing on problems related to autonomous driving. It showed that methods ranking high on established benchmarks, such as the Middlebury, perform below average when moved outside the laboratory conditions.
Unlike the PASCAL VOC and ImageNet challenges, the task considered in this work requires a specific set of calibrated modalities that cannot be easily acquired from the internet. In contrast to KITTY, it was not necessary to record large amounts of new data. By combining existing datasets, we have covered many practical scenarios. Additionally, we created two datasets with varying lighting conditions, which is an aspect not covered by the existing datasets.
Evaluation Methodology
The proposed evaluation methodology formulates the 6D object pose estimation task and defines a pose-error function which is compared with the commonly used function by Hinterstoisser et al. .
Methods for 6D object pose estimation report their predictions on the basis of two sources of information. Firstly, at training time, a method is given a training set , where is an object identifier. Training data may have different forms, e.g. a 3D mesh model of the object or a set of RGB-D images showing object instances in known 6D poses. Secondly, at test time, the method is provided with a test target defined by a pair , where is an image showing at least one instance of object . The goal is to estimate the 6D pose of one of the instances of object visible in image .
If multiple instances of the same object model are present, then the pose of an arbitrary instance may be reported. If multiple object models are shown in a test image, and annotated with their ground truth poses, then each object model may define a different test target. For example, if a test image shows three object models, each in two instances, then we define three test targets. For each test target, the pose of one of the two object instances has to be estimated.
This task reflects the industry-relevant bin-picking scenario where a robot needs to grasp a single arbitrary instance of the required object, e.g. a component such as a bolt or nut, and perform some operation with it. It is the simplest variant of the 6D localization task and a common denominator of its other variants, which deal with a single instance of multiple objects, multiple instances of a single object, or multiple instances of multiple objects. It is also the core of the 6D detection task, where no prior information about the object presence in the test image is provided .
2 Measuring Error
To calculate the error of an estimated pose w.r.t. the ground-truth pose in a test image , an object model is first rendered in the two poses. The result of the rendering is two distance maps A distance map stores at a pixel the distance from the camera center to a 3D point that projects to . It can be readily computed from the depth map which stores at the coordinate of and which can be obtained by a Kinect-like sensor. and . As in , the distance maps are compared with the distance map of the test image to obtain the visibility masks and , i.e. the sets of pixels where the model is visible in the image (Fig. 3). Given a misalignment tolerance , the error is calculated as:
The object pose can be ambiguous, i.e. there can be multiple poses that are indistinguishable. This is caused by the existence of multiple fits of the visible part of the object surface to the entire object surface. The visible part is determined by self-occlusion and occlusion by other objects and the multiple surface fits are induced by global or partial object symmetries.
Definition (1) is different from the original definition in where the pixel-wise cost linearly increases to as increases to . The new definition is easier to interpret and does not penalize small distance differences that may be caused by imprecisions of the depth sensor or of the ground-truth pose.
Criterion of Correctness.
Comparison to Hinterstoisser et al.
Datasets
We collected six publicly available datasets, some of which we reduced to remove redundancies Identifiers of the selected images are available on the project website. and re-annotated to ensure a high quality of the ground truth. Additionally, we created two new datasets focusing on varying lighting conditions, since this variation is not present in the existing datasets. An overview of the datasets is in Fig. 1 and a detailed description follows.
The datasets consist of texture-mapped 3D object models and training and test RGB-D images annotated with ground-truth 6D object poses. The 3D object models were created using KinectFusion-like systems for 3D surface reconstruction . All images are of approximately VGA resolution.
For training, a method may use the 3D object models and/or the training images. While 3D models are often available or can be generated at a low cost, capturing and annotating real training images requires a significant effort. The benchmark is therefore focused primarily on the more practical scenario where only the object models, which can be used to render synthetic training images, are available at training time. All datasets contain already synthesized training images. Methods are allowed to synthesize additional training images, but this option was not utilized for the evaluation in this paper. Only T-LESS and TUD-L include real training images of isolated, i.e. non-occluded, objects.
To generate the synthetic training images, objects from the same dataset were rendered from the same range of azimuth/elevation covering the distribution of object poses in the test scenes. The viewpoints were sampled from a sphere, as in , with the sphere radius set to the distance of the closest object instance in the test scenes. The objects were rendered with fixed lighting conditions and a black background.
The test images are real images from a structured-light sensor – Microsoft Kinect v1 or Primesense Carmine 1.09. The test images originate from indoor scenes with varying complexity, ranging from simple scenes with a single isolated object instance to very challenging scenes with multiple instances of several objects and a high amount of clutter and occlusion. Poses of the modeled objects were annotated manually. While LM, IC-MI and RU-APC provide annotation for instances of only one object per image, the other datasets provide ground-truth for all modeled objects. Details of the datasets are in Tab. 1.
2 The Dataset Collection
LM (a.k.a. Linemod) has been the most commonly used dataset for 6D object pose estimation. It contains 15 texture-less household objects with discriminative color, shape and size. Each object is associated with a test image set showing one annotated object instance with significant clutter but only mild occlusion. LM-O (a.k.a. Linemod-Occluded) provides ground-truth annotation for all other instances of the modeled objects in one of the test sets. This introduces challenging test cases with various levels of occlusion.
IC-MI/IC-BIN [34, 7].
IC-MI (a.k.a. Tejani et al.) contains models of two texture-less and four textured household objects. The test images show multiple object instances with clutter and slight occlusion. IC-BIN (a.k.a. Doumanoglou et al., scenario 2) includes test images of two objects from IC-MI, which appear in multiple locations with heavy occlusion in a bin-picking scenario. We have removed test images with low-quality ground-truth annotations from both datasets, and refined the annotations for the remaining images in IC-BIN.
T-LESS [16].
It features 30 industry-relevant objects with no significant texture or discriminative color. The objects exhibit symmetries and mutual similarities in shape and/or size, and a few objects are a composition of other objects. T-LESS includes images from three different sensors and two types of 3D object models. For our evaluation, we only used RGB-D images from the Primesense sensor and the automatically reconstructed 3D object models.
RU-APC [28].
This dataset (a.k.a. Rutgers APC) includes 14 textured products from the Amazon Picking Challenge 2015 , each associated with test images of a cluttered warehouse shelf. The camera was equipped with LED strips to ensure constant lighting. From the original dataset, we omitted ten objects which are non-rigid or poorly captured by the depth sensor, and included only one from the four images captured from the same viewpoint.
TUD-L/TYO-L.
Two new datasets with household objects captured under different settings of ambient and directional light. TUD-L (TU Dresden Light) contains training and test image sequences that show three moving objects under eight lighting conditions. The object poses were annotated by manually aligning the 3D object model with the first frame of the sequence and propagating the initial pose through the sequence using ICP. TYO-L (Toyota Light) contains 21 objects, each captured in multiple poses on a table-top setup, with four different table cloths and five different lighting conditions. To obtain the ground truth poses, manually chosen correspondences were utilized to estimate rough poses which were then refined by ICP. The images in both datasets are labeled by categorized lighting conditions.
Evaluated Methods
The evaluated methods cover the major research directions of the 6D object pose estimation field. This section provides a review of the methods, together with a description of the setting of their key parameters. If not stated otherwise, the image-based methods used the synthetic training images.
For each pixel of an input image, a regression forest predicts the object identity and the location in the coordinate frame of the object model, a so called “object coordinate”. Simple RGB and depth difference features are used for the prediction. Each object coordinate prediction defines a 3D-3D correspondence between the image and the 3D object model. A RANSAC-based optimization schema samples sets of three correspondences to create a pool of pose hypotheses. The final hypothesis is chosen, and iteratively refined, to maximize the alignment of predicted correspondences, as well as the alignment of observed depth with the object model. The main parameters of the method were set as follows: maximum feature offset: , features per tree node: 1000, training patches per object: 1.5M, number of trees: 3, size of the hypothesis pool: 210, refined hypotheses: 25. Real training images were used for TUD-L and T-LESS.
Brachmann-16 [2].
The method of is extended in several ways. Firstly, the random forest is improved using an auto-context algorithm to support pose estimation from RGB-only images. Secondly, the RANSAC-based optimization hypothesizes not only with regard to the object pose but also with regard to the object identity in cases where it is unknown which objects are visible in the input image. Both improvements were disabled for the evaluation since we deal with RGB-D input, and it is known which objects are visible in the image. Thirdly, the random forest predicts for each pixel a full, three-dimensional distribution over object coordinates capturing uncertainty information. The distributions are estimated using mean-shift in each forest leaf, and can therefore be heavily multi-modal. The final hypothesis is chosen, and iteratively refined, to maximize the likelihood under the predicted distributions. The 3D object model is not used for fitting the pose. The parameters were set as: maximum feature offset: , features per tree node: 100, number of trees: 3, number of sampled hypotheses: 256, pixels drawn in each RANSAC iteration: 10K, inlier threshold: .
Tejani-14 [34].
Linemod is adapted into a scale-invariant patch descriptor and integrated into a regression forest with a new template-based split function. This split function is more discriminative than simple pixel tests and accelerated via binary bit-operations. The method is trained on positive samples only, i.e. rendered images of the 3D object model. During the inference, the class distributions at the leaf nodes are iteratively updated, providing occlusion-aware segmentation masks. The object pose is estimated by accumulating pose regression votes from the estimated foreground patches. The baseline evaluated in this paper implements but omits the iterative segmentation/refinement step and does not perform ICP. The features and forest parameters were set as in : number of trees: 10, maximum depth of each tree: 25, number of features in both the color gradient and the surface normal channel: 20, patch size: 1/2 the image, rendered images used to train each forest: 360.
Kehl-16 [22].
Scale-invariant RGB-D patches are extracted from a regular grid attached to the input image, and described by features calculated using a convolutional auto-encoder. At training time, a codebook is constructed from descriptors of patches from the training images, with each codebook entry holding information about the 6D pose. For each patch descriptor from the test image, -nearest neighbors from the codebook are found, and a 6D vote is cast using neighbors whose distance is below a threshold . After the voting stage, the 6D hypothesis space is filtered to remove spurious votes. Modes are identified by mean-shift and refined by ICP. The final hypothesis is verified in color, depth and surface normals to suppress false positives. The main parameters of the method with the used values: patch size: , patch sampling step: , -nearest neighbors: 3, threshold : , number of extracted modes from the pose space: 8. Real training images were used for T-LESS.
2 Template Matching Methods
A template matching method that applies an efficient cascade-style evaluation to each sliding window location. A simple objectness filter is applied first, rapidly rejecting most locations. For each remaining location, a set of candidate templates is identified by a voting procedure based on hashing, which makes the computational complexity largely unaffected by the total number of stored templates. The candidate templates are then verified as in Linemod by matching feature points in different modalities (surface normals, image gradients, depth, color). Finally, object poses associated with the detected templates are refined by particle swarm optimization (PSO). The templates were generated by applying the full circle of in-plane rotations with step to a portion of the synthetic training images, resulting in 11–23K templates per object. Other parameters were set as described in . We present also results without the last refinement step (Hodaň-15-nr).
3 Methods Based on Point-Pair Features
Drost-10-edge.
An extension of which additionally detects 3D edges from the scene and favors poses in which the model contours are aligned with the edges. A multi-modal refinement minimizes the surface distances and the distances of reprojected model contours to the detected edges. The evaluation was performed using the same software and parameters as Drost-10, but with activated parameter train_3d_edges during the model creation.
Vidal-18 [35].
The point cloud is first sub-sampled by clustering points based on the surface normal orientation. Inspired by improvements of , the matching strategy of was improved by mitigating the effect of the feature discretization step. Additionally, an improved non-maximum suppression of the pose candidates from different reference points removes spurious matches. The most voted 500 pose candidates are sorted by a surface fitting score and the 200 best candidates are refined by projective ICP. For the final 10 candidates, the consistency of the object surface and silhouette with the scene is evaluated. The sampling distance for model, scene and features was set to 5% of the object diameter, and 20% of the scene points were used as the reference points.
4 Methods Based on 3D Local Features
A RANSAC-based method that iteratively samples three feature correspondences between the object model and the scene. The correspondences are obtained by matching 3D local shape descriptors and are used to generate a 6D pose candidate, whose quality is measured by the consensus set size. The final pose is refined by ICP. The method achieved the state-of-the-art results on earlier object recognition datasets captured by LIDAR, but suffers from a cubic complexity in the number of correspondences. The number of RANSAC iterations was set to 10000, allowing only for a limited search in cluttered scenes. The method was evaluated with several descriptors: 153d SI , 352d SHOT , 30d ECSAD , and 1536d PPFH . None of the descriptors utilize color.
Buch-17 [4].
Evaluation
The methods reviewed in Sec. 4 were evaluated by their original authors on the datasets described in Sec. 3, using the evaluation methodology from Sec. 2.
The parameters of each method were fixed for all objects and datasets. The distribution of object poses in the test scenes was the only dataset-specific information used by the methods. The distribution determined the range of viewpoints from which the object models were rendered to obtain synthetic training images.
Pose Error.
Performance Score.
The performance is measured by the recall score, i.e. the fraction of test targets for which a correct object pose was estimated. Recall scores per dataset and per object are reported. The overall performance is given by the average of per-dataset recall scores. We thus treat each dataset as a separate challenge and avoid the overall score being dominated by larger datasets.
Subsets Used for the Evaluation.
We reduced the number of test images to remove redundancies and to encourage participation of new, in particular slow, methods. From the total of 62K test images, we sub-sampled 7K, reducing the number of test targets from 110K to 17K (Tab. 1). Full datasets with identifiers of the selected test images are on the project website. TYO-L was not used for the evaluation presented in this paper, but it is a part of the online evaluation.
2 Results
Speed.
Open Problems.
Occlusion is a big challenge for current methods, as shown by scores dropping swiftly already at low levels of occlusion (Fig. 4, right). The big gap between LM and LM-O scores provide further evidence. All methods perform on LM by at least 30% better than on LM-O, which includes the same but partially occluded objects. Inspection of estimated poses on T-LESS test images confirms the weak performance for occluded objects. Scores on TUD-L show that varying lighting conditions present a serious challenge for methods that rely on synthetic training RGB images, which were generated with fixed lighting. Methods relying only on depth information (e.g. Vidal-18, Drost-10) are noticeably more robust under such conditions. Note that Brachmann-16 achieved a high score on TUD-L despite relying on RGB images because it used real training images, which were captured under the same range of lighting conditions as the test images. Methods based on 3D local features and learning-based methods have very low scores on T-LESS, which is likely caused by the object symmetries and similarities. All methods perform poorly on RU-APC, which is likely because of a higher level of noise in the depth images.
Conclusion
We have proposed a benchmark for 6D object pose estimation that includes eight datasets in a unified format, an evaluation methodology, a comprehensive evaluation of 15 recent methods, and an online evaluation system open for continuous submission of new results. With this benchmark, we have captured the status quo in the field and will be able to systematically measure its progress in the future. The evaluation showed that methods based on point-pair features perform best, outperforming template matching methods, learning-based methods and methods based on 3D local features. As open problems, our analysis identified occlusion, varying lighting conditions, and object symmetries and similarities.
Acknowledgements
We gratefully acknowledge Manolis Lourakis, Joachim Staib, Christoph Kick, Juil Sock and Pavel Haluza for their help. This work was supported by CTU student grant SGS17/185/OHK3/3T/13, Technology Agency of the Czech Republic research program TE01020415 (V3C – Visual Computing Competence Center), and the project for GAČR, No. 16-072105: Complex network methods applied to ancient Egyptian data in the Old Kingdom (2700–2180 BC).