BOP Challenge 2022 on Detection, Segmentation and Pose Estimation of Specific Rigid Objects

Martin Sundermeyer, Tomas Hodan, Yann Labbe, Gu Wang, Eric Brachmann, Bertram Drost, Carsten Rother, Jiri Matas

Introduction

Estimating the 6D pose, i.e., the 3D translation and 3D rotation, of specific rigid objects from a single image is an important task for application fields such as robotic manipulation, augmented reality, or autonomous driving. The BOP Challenge 2022 is the fourth in a series of public challenges that are part of the BOPBOP stands for Benchmark for 6D Object Pose Estimation . project aiming to continuously report the state of the art in 6D object pose estimation. The first challenge was organized in 2017 and the results were published in . Results of the second challenge from 2019 , the third from 2020 , and the fourth from 2022 are included and discussed in this paper.

Participants of the 2022 challenge were competing on three tasks: 6D object localization, 2D object detection, and 2D object segmentation. The 6D object localization task has the same evaluation methodology and leaderboard since 2019, while the latter two tasks were introduced in 2022.

In the 6D object localization task, methods report their predictions on the basis of two sources of information. Firstly, at training time, a method is given 3D object models and training images showing the objects in known 6D poses. Secondly, at test time, the method is provided with a test image and a list of object instances visible in the image, and the goal is to estimate 6D poses of the listed instances. The images consist of RGB-D (aligned color and depth) channels and intrinsic camera parameters are known.

The 2D object detection and segmentation tasks were introduced to address the design of the majority of recent object pose estimation methods, which start by detecting/segmenting objects and then estimate their poses from the predicted image regions. Evaluating the detection/segmentation and pose estimation stages separately enables a better understanding of advances in the two stages. To create an opportunity for detector-agnostic comparison of pose estimation methods and to allow participants to focus only on the pose estimation stage, we also provided default detections and segmentations from Mask R-CNN trained for CosyPose , the winning method in 2020.

The challenge primarily focuses on the practical scenario where no real images are available at training time, only the 3D object models and images synthesized using the models. While capturing real images of objects under various conditions and annotating the images with 6D object poses requires a significant human effort , the 3D models are either available before the physical objects, which is often the case for manufactured objects, or can be reconstructed at an admissible cost. Approaches for reconstructing 3D models of opaque, matte and moderately specular objects are established and promising approaches for transparent and highly specular objects are emerging .

In the 2019 challenge, methods using the depth image channel were mostly based on point pair features (PPF’s) and clearly outperformed methods relying only on the RGB channels, all of which were based on deep neural networks (DNN’s). DNN-based methods need large amounts of annotated training images, which had been typically obtained by OpenGL rendering of the 3D object models on random backgrounds . However, as suggested in , the evident domain gap between these “render & paste” training images and real test images limits the potential of the DNN-based methods. To reduce the gap between the synthetic and real domains and thus to bring fresh air to the DNN world, we joined the development of BlenderProcgithub.com/DLR-RM/BlenderProc , an open-source, physically-based renderer (PBR). For the 2020 challenge, we then provided participants with 350K PBR training images (see for examples), which helped the DNN-based methods to achieve noticeably higher accuracy and to finally catch up with the PPF-based methods.

In the 2022 challenge, DNN-based methods for 6D object localization clearly outperformed PPF-based methods in both accuracy and speed, with the performance gains coming mostly from advances in network architectures and training schemes. The largest improvements were achieved on challenging industry-relevant datasets ITODD and T-LESS , and on the HB dataset which includes diverse objects captured under various levels of occlusion. Remarkably, RGB methods from 2022 surpassed RGB-D methods from 2020, the performance gap between methods trained only on PBR images and methods trained also on real images noticeably shrinked, and some methods started training on the depth image channel in addition to the RGB channels. On the new 2D object detection and segmentation tasks, large gains were achieved w.r.t. a baseline from 2020.

Sec. 2 of this paper defines the evaluation methodology, Sec. 3 introduces datasets, Sec. 4 describes the experimental setup and analyzes the results, Sec. 5 presents the awards of the BOP Challenge 2022, and Sec. 6 concludes the paper.

Evaluation Methodology

Methods are evaluated on the task of 6D object localization, as in 2019 and 2020 , and additionally on the tasks of 2D object detection and 2D object segmentation. The tasks are defined below together with accuracy scores that are used to compare methods. Participants could submit their results to any of the three tasks. Note that although all BOP datasets currently include RGB-D images (Sec. 3), a method may have used any of the image channels.

Training input: At training time, a detection/segmentation method is provided a set of training images showing objects annotated with ground-truth 2D bounding boxes (for the detection task) and binary masks (for the segmentation task). The boxes are amodal (covering the whole object silhouette, including the occluded parts) while the masks are modal (covering only the visible object part). The method can also use 3D mesh models that are available for the objects (e.g., to synthesize extra training images).

Test input: At test time, the method is given an image showing an arbitrary number of instances of an arbitrary number of objects from a considered dataset. No prior information about the visible object instances is provided.

Test output: The method produces a list of amodal 2D bounding boxes (for detection) and modal binary masks (for segmentation) with confidences.

Metrics: Following the the evaluation methodology from the COCO 2020 Object Detection Challenge , the detection/segmentation accuracy is measured by the Average Precision (AP). Specifically, a per-object APO\text{AP}_{O} score is calculated by averaging the precision at multiple Intersection over Union (IoU) thresholds: [0.5,0.55,…,0.95][0.5,0.55,\dots,0.95]. The accuracy of a method on a dataset DD is measured by APD\text{AP}_{D} calculated by averaging per-object APO\text{AP}_{O} scores, and the overall accuracy on the core datasets (Sec. 3) is measured by APC\text{AP}_{C} defined as the average of the per-dataset APD\text{AP}_{D} scores.

Analagous to the 6D localization task, only object instances for which at least 10%10\% of the projected surface area is visible need to be detected/segmented. Correct predictions for objects that are visible from less than 10%10\% are filtered out and not counted as false positives. Up to 100100 predictions with the highest scores per image are considered.

2 6D Object Localization Task

As in the 2019 and 2020 editions of the challenge, methods are evaluated on the task of 6D localization of a varying number of instances of a varying number of objects from a single image. This variant of the 6D object localization task is referred to as ViVo and defined as follows.See Sec. A.1 in for a discussion on why the methods are evaluated on 6D object localization instead of 6D object detection, where no prior information about the visible object instances is provided .

Training input: A method is provided a set of training images showing objects annotated with 6D poses, and 3D mesh models of the objects (typically with a color texture). A 6D pose is defined by a matrix P=[R ∣ t]\textbf{P}=[\mathbf{R}\,|\,\mathbf{t}], where R\mathbf{R} is a 3D rotation matrix, and t\mathbf{t} is a 3D translation vector. The matrix P defines a rigid transformation from the 3D space of the object model to the 3D space of the camera.

Test input: The method is given an image unseen during training and a list L=[(o1,n1),L=[(o_{1},n_{1}), …,\dots, (om,nm)](o_{m},n_{m})], where nin_{i} is the number of instances of object oio_{i} visible in the image.

Test output: The method outputs a list E=[E1,<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>…</mo><moseparator="true">,</mo></mrow><annotationencoding="application/x−tex">…,</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3174em;vertical−align:−0.1944em;"></span><spanclass="minner">…</span><spanclass="mspace"style="margin−right:0.1667em;"></span><spanclass="mpunct">,</span></span></span></span></span>Em]E=[E_{1},<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>…</mo><mo separator="true">,</mo></mrow><annotation encoding="application/x-tex">\dots,</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3174em;vertical-align:-0.1944em;"></span><span class="minner">…</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mpunct">,</span></span></span></span></span>E_{m}], where EiE_{i} is a list of nin_{i} pose estimates with confidences for instances of object oio_{i}.

Metrics: The 6D object localization task is evaluated as in the 2020 challenge . In short, the error of an estimated pose w.r.t. the ground-truth pose is calculated by three pose-error functions: Visible Surface Discrepancy (VSD) which treats indistinguishable poses as equivalent by considering only the visible object part, Maximum Symmetry-Aware Surface Distance (MSSD) which considers a set of pre-identified global object symmetries and measures the surface deviation in 3D, and Maximum Symmetry-Aware Projection Distance (MSPD) which considers the object symmetries and measures the perceivable deviation. An estimated pose is considered correct w.r.t. a pose-error function ee, if e<θee<\theta_{e}, where e∈{VSD,MSSD,MSPD}e\in\{\text{VSD},\text{MSSD},\text{MSPD}\} and θe\theta_{e} is the threshold of correctness. The fraction of annotated object instances for which a correct pose is estimated is referred to as Recall. The Average Recall w.r.t. a function ee, denoted as ARe\text{AR}_{e}, is defined as the average of the Recall rates calculated for multiple settings of the threshold θe\theta_{e} and also for multiple settings of a misalignment tolerance τ\tau in the case of VSD. The accuracy of a method on a dataset DD is measured by: ARD=(ARVSD+ARMSSD+ARMSPD) / 3\text{AR}_{D}=(\text{AR}_{\text{VSD}}+\text{AR}_{\text{MSSD}}+\text{AR}_{\text{MSPD}})\,/\,3, which is calculated over estimated poses of all objects from DD. The overall accuracy on the core datasets is measured by ARC\text{AR}_{C} defined as the average of the per-dataset ARD\text{AR}_{D} scores.When calculating ARC, scores are not averaged over objects before averaging over datasets, which is done when calculating APC\text{AP}_{C} (Sec. 2.1) to comply with the original COCO evaluation methodology .

Datasets

BOP currently includes twelve datasets in a unified format – sample test images are in Fig. 2 and dataset parameters in Tab. 1. Seven from the twelve were selected as core datasets: LM-O, T-LESS, ITODD, HB, YCB-V, TUD-L, IC-BIN. A method had to be evaluated on all core datasets to be considered for the main challenge awards (Sec. 5).

Each dataset includes 3D object models and training and test RGB-D images annotated with ground-truth 6D object poses. The object models are provided in the form of 3D meshes (in most cases with a color texture) which were created manually or using KinectFusion-like systems for 3D reconstruction . While all test images are real, training images may be real and/or synthetic. The seven core datasets include a total of 350K photorealistic PBR (physically-based rendered) training images generated and automatically annotated using BlenderProc . Example images are shown in and a detailed description of the generation process and an analysis of the importance of PBR training images is provided in Sec. 3.2 and 4.3 of the 2020 challenge paper . Datasets T-LESS, TUD-L and YCB-V include also real training images, and most datasets additionally include training images obtained by OpenGL rendering of the 3D object models on a black background. Test images were captured in scenes with graded complexity, often with clutter and occlusion. The HB and ITODD datasets include also real validation images – in this case, the ground-truth poses are publicly available only for the validation and not for the test images. The datasets can be downloaded from the BOP websitebop.felk.cvut.cz/datasets and more details about the datasets can be found in Chapter 7 of .

Results and Discussion

This section presents results of the BOP Challenge 2022, compares them with results from 2019 and 2020 challenge editions, and summarizes the main messages for our field.

In total, 49 methods were evaluated on the ViVo variant of the 6D object localization task on all seven core datasets – 11 methods in 2019, 15 in 2020, and 23 in 2022. Additionally, 8 methods were evaluated on the new detection task and 8 methods on the new segmentation task.

Participants of the BOP Challenge 2022 were submitting results of their methods to the online evaluation system at bop.felk.cvut.cz from May 1, 2022 until the deadline on October 16, 2022. The methods were evaluated on the ViVo variant of the 6D object localization task as described in Sec. 2.2 and on the 2D object detection and segmentation tasks as described in Sec. 2.1. The evaluation scripts are publicly available in the BOP toolkit.github.com/thodan/bop_toolkit

A method had to use a fixed set of hyper-parameters across all objects and datasets. For training, a method may have used the provided object models and training images, and rendered extra training images using the object models. However, not a single pixel of test images may have been used for training, nor the individual ground-truth poses or object masks provided for the test images. Ranges of the azimuth and elevation camera angles, and a range of the camera-object distances determined by the ground-truth poses from test images is the only information about the test set that may have been used during training.

Only subsets of test images were used to remove redundancies and speed up the evaluation, and only object instances for which at least 10%10\% of the projected surface area is visible were considered in the evaluation.

2 6D Object Localization Results

An overview of the 6D object localization results is in Tab. 2 and properties of the evaluated methods in Tab. 3. In 2022, all 23 of the new submissions rely on DNN’s in their pipelines and 18 of them outperform CosyPose , the top-performing method from the 2020 challenge. The best method from 2022, GDRNPP , is purely learning-based and achieves 83.7 ARC, outperforming CosyPose by substantial 13.9 points in ARC (#1−-#19 in Tab. 2). Gains in accuracy are most notable on the industrial ITODD dataset where GDRNPP reaches 67.9 ARC (+36.6 ARC w.r.t. CosyPose). This result is significant as ITODD reflects a challenging industrial scenario and was previously dominated by PPF-based approaches, the best of which, KoenigHybrid (#24), achieved 48.3 ARC.

GDRNPP dominates in 2022: The GDRNPP method was evaluated in seven variants, four of which are on top of the leaderboard. The variants were tailored towards different BOP 2022 awards (Sec. 5) by relying on different data domains and modalities and on different detection and pose refinement methods. Having results of these variants enables to understand the importance of individual aspects of the pipeline. The common ground is the Geometrically-Guided Direct Regression Network (GDR-Net) , which takes an RGB object crop as input and densely predicts 2D-3D correspondences, identities of surface fragments , and a mask of the visible object part. Then, instead of applying PnP-RANSAC , the predictions are concatenated and fed into a small CNN with a fully connected head that regresses a scale-invariant translation and a 3D rotation using the allocentric 6D representation . The 3D rotation loss takes into account object symmetries that are provided in the BOP datasets. For BOP 2022, GDR-Net was modified by exchanging the ResNet34 backbone with ConvNext , predicting both modal and amodal masks as intermediate representations, and applying stronger domain randomization. The winning GDRNPP variant trains YOLOX for object detection and GDR-Net for pose estimation on the provided PBR and real RGB images, and refines the poses by a multi-hypotheses refinement method inspired by Coupled Iterative Refinement (CIR) , which is trained on PBR and real RGB-D images.

Training on depth: Methods RCVPose3D (#7) and CIR (#12; a variant is also used in #1, 2, 4), started benefiting from learning on the depth channel in addition to the RGB channels (only PointVoteNet2 applied a neural network to the depth channel in 2020). On the flip side, the multi-hypotheses refinement methods can be time-intensive – the CIR-based approach increases the inference time of GDRNPP by 6.03s per image on average (#1−-#3).

Increased accuracy & speed: The third GDRNPP entry replaces the CIR-based refinement , which is used in the top two entries, by a fast and simple depth-based adjustment of the 3D translation and still achieves impressive 80.5 ARC in just 0.23s per image. In comparison, the best method in 2020 that took less than 1s per image is KoenigHybrid (#24) with 63.9 ARC and 0.63s per image.

RGB-only from 2022 beats RGB-D from 2020: The best method that relies only on RGB image channels at both training and test time is a variant of GDRNPP (#13). Without any pose refinement, this method achieves 72.8 ARC which is +9.1 w.r.t. CosyPose that applies RGB-based pose refinement (#25) and +3.0 w.r.t. to the overall best method from 2020, i.e., CosyPose with a depth-based ICP (#19).

Synthetic-to-real gap shrinks further: Another important result was achieved by the GDRNPP variant that is trained only on the provided synthetic PBR images rendered with BlenderProc . With 82.7 ARC, this variant achieves the second highest accuracy. On datasets with real training images (T-LESS, YCB-V, TUD-L), the synthetically trained variant is only -2.5 ARC on average behind the winning method that was trained on both PBR and real training images. In the RGB-only setting, the synthetic-to-real gap has been reduced on the three datasets from Δ\Delta15.8 ARC (observed on CosyPose in 2020; #25−-#29) to Δ\Delta6.2 ARC (observed on GDRNPP in 2022; #13−-#18). The BOP 2020 results demonstrated the importance of training on PBR images over training on rasterized images with random backgrounds. The BOP 2022 results confirm this observation and also suggest that the synthetic-to-real gap monotonically shrinks as the accuracy of methods increases (see, e.g., #25−-#29, #14−-#20, #5−-#9, #1−-#2 in Tab. 2).

Scalability in the number of objects: The advancement in the synthetic-to-real transfer is crucial for increasing the scope of applications. In addition, real world applications require methods whose computational and memory resources scale gracefully with the amount of target objects. The top four GDRNPP variants are all trained with at least one pose network per object. This means that the training and inference time complexity and the inference memory increase linearly with the number of target objects. When GDRNPP is trained with one pose network per BOP dataset containing 2–33 objects (Tab. 1), it achieves only 74.8 ARC (#11) and is outperformed by, e.g., Extended_FCOS+PFA (#5) that reaches 78.7 ARC with one pose network per dataset. This raises the question how the results would change if was trained per object.

2D detection followed by 6D pose estimation: Almost all 6D object localization methods evaluated in 2022 start by detecting the object instances in RGB images by predicting their 2D bounding boxes. Some methods also predict 2D object masks in the detected regions at training time for loss calculation or extra supervision , and some predict 2D masks at both training and inference time and use them to establish correspondences . The only exception is RCVPose3D , which does not start by detecting object instance in the RGB image channels and instead segments the object instances in 3D point clouds calculated from the depth image channel.

Detector-agnostic results: Eleven methods use the default 2D object detections (Default in column Det./seg. in Tab. 3), which were provided to participants of the 2022 challenge and produced by Mask R-CNN trained for the first stage of CosyPose in 2020. Three of these methods use detections from Mask R-CNN trained only on PBR images, and eight use detections from Mask R-CNN trained on synthetic and real images (where the synthetic include PBR and additional images synthesized by the authors of ). Among the eight methods, GDRNPP is once again at the top with 79.8 ARC (#4). We can therefore conclude that the pose estimation performance of the GDRNPP pipeline is performing best independent of the used detection method. However, the accuracy gap to other methods decreases with the default detections, e.g., from +7.2 ARC (#1−-#8) to +3.3 ARC (#4−-#8) w.r.t. ZebraPose .

3 2D Object Detection Results

As shown in Tab. 4, the YOLOX detector from GDRNPP has the top performance of 77.3 APC. This detector employs a ConvNext backbone and was trained with the Ranger optimizer and strong data augmentation. Mask R-CNN from CosyPose only achieves 60.5 APC (-16.8 APC), which explains the +3.9 ARC gain in the pose accuracy (#1−-#4 in Tab. 2). YOLOX is relatively insensitive to the image domain, improving only +3.5 APC (#1−-#2 in Tab. 4) when trained also on real images. Mask R-CNN yields +4.8 APC (#6−-#7) and FCOS yields +5.4 APC (#3−-#4) in such a comparison.

Although all 2D object detection methods rely only on RGB and ignore the depth channel, they work remarkably well even on the texture-less objects from T-LESS (see the BOP website for per-dataset scores). However, detections from YOLOX on YCB-V in Fig. 1 reveal a limitation of the RGB-only detection that fails to distinguish the two differently sized clamps. This detection failure can cause wrong pose estimates even though the rendered scene seems perfectly plausible. Depth data could help to disambiguate the object scale in such cases.

4 2D Object Segmentation Results

We see an improvement from 40.5 APC achieved by the default masks from Mask R-CNN to 58.7 APC achieved by masks from ZebraPoseSAT (+18.2 APC; #1−-#7 in Tab. 5). Interestingly, ZebraPoseSAT predicts the high-quality masks in regions determined by the default detections from Mask R-CNN (#6 in Tab. 4) and would likely achieve even higher segmentation accuracy if relying on detections from YOLOX trained for GDRNPP. As mentioned in Sec. 4.2, most 6D object localization methods evaluated in 2022 start by 2D object detection. Leveraging 2D object segmentation instead could improve results on objects with irregular shapes which are included, e.g., in the industrial ITODD dataset .

Awards

The following BOP Challenge 2022 awards were presented at the 7th Workshop on Recovering 6D Object Pose cmp.felk.cvut.cz/sixd/workshop_2022 organized at the ECCV 2022 conference. The awards are based on the 6D object localization results in Tab. 2, method properties in Tab. 3, the 2D object detection results in Tab. 4, and the 2D object segmentation results in Tab. 5.

The GDRNPP submissions were prepared by Xingyu Liu, Ruida Zhang, Chenyangguang Zhang, Bowen Fu, Jiwen Tang, Xiquan Liang, Jingyi Tang, Xiaotian Cheng, Yukang Zhang, Gu Wang, Xiangyang Ji; Extended_FCOS+PFA by Yang Hai, Rui Song, Zhiqiang Liu, Jiaojiao Li, Mathieu Salzmann, Pascal Fua, Yinlin Hu; ZebraPoseSAT by Yongzhi Su, Praveen Nathan, Torben Fetzer, Jason Rambach, Didier Stricker, Mahdi Saleh, Yan Di, Nassir Navab, Benjamin Busam, Federico Tombari, Yongliang Lin, Yu Zhang, Coupled Iterative Refinement by Lahav Lipson, Zachary Teed, Ankit Goyal, and Jia Deng; and RCVPose3D by Yangzheng Wu, Alireza Javaheri, Mohsen Zand, Michael Greenspan.

Awards for 6D object localization methods:

The Overall Best Method: GDRNPP-PBRReal-RGBD-MModel

The Best RGB-Only Method: GDRNPP-PBRReal-RGB-MModel

The Best Fast Method (less than 1s per image): GDRNPP-PBRReal-RGBD-MModel-Fast

The Best BlenderProc-Trained Method: GDRNPP-PBR-RGBD-MModel

The Best Single-Model Method (trained per dataset): Extended_FCOS+PFA-MixPBR-RGBD

The Best Open-Source Method: GDRNPP-PBRReal-RGBD-MModel

The Best Method On Default Detections/Segment.: GDRNPP-PBRReal-RGBD-MModel-OfficialDet

The Best Method on T-LESS, ITODD, YCB-V, HB: GDRNPP-PBRReal-RGBD-MModel

The Best Method on LM-O: Extended_FCOS+PFA-MixPBR-RGBD

The Best Method on TUD-L: Coupled Iterative Refinement (CIR)

The Best Method on IC-BIN: RCVPose3D_SingleModel_VIVO_PBR

Awards for 2D object detection/segmentation methods:

The Overall Best Detection Method: GDRNPPDet_PBRReal

The Best BlenderProc-Trained Detection Method: GDRNPPDet_PBR

The Overall Best Segmentation Method: ZebraPoseSAT-EffnetB4 (DefaultDetection)

The Best BlenderProc-Trained Segment. Method: ZebraPoseSAT-EffnetB4 (DefaultDet+PBR_Only)

Conclusions

In the BOP Challenge 2022, we witnessed another breakthrough in the 6D pose estimation accuracy, efficiency and synthetic-to-real transfer. Methods based on deep neural networks now clearly surpass the traditional methods based on point pair features in both accuracy and speed. Variations of the winning GDRNPP method allowed us to analyze the importance of different aspects related to training domains, modalities and run-time efficiency. Besides, we individually measured 2D detection and segmentation performance and could thereby determine sources of gains in the multi-stage pose estimation pipelines. Despite the progress, accuracy scores have not been saturated on most BOP datasets and we are already looking forward to insights from the next challenge. The online evaluation system at bop.felk.cvut.cz stays open and raw results of all methods will be made publicly available.

References