MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

Yann Labbé, Lucas Manuelli, Arsalan Mousavian, Stephen Tyree, Stan Birchfield, Jonathan Tremblay, Justin Carpentier, Mathieu Aubry, Dieter Fox, Josef Sivic

Introduction

Accurate 6D object pose estimation is essential for many robotic and augmented reality applications. Current state-of-the-art methods are learning-based and require 3D models of the objects of interest at both training and test time. These methods require hours (or days) to generate synthetic data for each object and train the pose estimation model. They thus cannot be used in the context of robotic applications where the objects are only known during inference (e.g. CAD models are provided by a manufacturer or reconstructed ), and where rapid deployment to novel scenes and objects is key.

The goal of this work is to estimate the 6D pose of novel objects, i.e., objects that are only available at inference time and are not known in advance during training. This problem presents the challenge of generalizing to the large variability in shape, texture, lighting conditions, and severe occlusions that can be encountered in real-world applications. Some prior works have considered category-level pose estimation to partially address the challenge of novel objects by developing methods that can generalize to novel object instances of a known class (e.g. mugs or shoes). These methods however do not generalize to object instances outside of training categories. Other methods aim at generalizing to any novel instances regardless of their category . These works present important technical limitations. They rely on non-learning based components for generating pose hypotheses (e.g. PPF ), for pose refinement (e.g. PnP and ICP ), for computing photometric errors in pixel space , or for estimating the object depth (e.g. using only the size of a 2D detection ). These components however inherently cannot benefit from being trained on large amount of data to gain robustness with respect to noise, occlusions, or object variability. Learning-based methods also have the potential to improve as the quality and size of the datasets improve.

Pipelines for 6D pose estimation of known (not novel) objects that consist of multiple learned stages have shown excellent performance on several benchmarks with various illumination conditions, textureless objects, cluttered scenes and high levels of occlusions. We take inspiration from which split the problem into three parts: (i) 2D object detection, (ii) coarse pose estimation, and (iii) iterative refinement via render & compare. We aim at extending this approach to novel objects unseen at training time. The detection of novel objects has been addressed by prior works and is outside the scope of this paper. In this work, we focus on the coarse and refinement networks for 6D pose estimation. Extending the paradigm from presents three major challenges. First, the pose of an object depends heavily on both its visual appearance and choice of coordinate system (defined in the CAD model of the object). In existing refinement networks based on render & compare , this information is encoded in the network weights during training, leading to poor generalization results when tested on novel objects. Second, direct regression methods for coarse pose estimation are trained with specific losses for symmetric objects , requiring that object symmetries be known in advance. Finally, the diversity of shape and visual properties of the objects that can be encountered in real-world applications is immense. Generalizing to novel objects requires robustness to properties such as object symmetries, variability of object shape, and object textures (or absence of).

Contributions. We address these challenges and propose a method for estimating the pose of any novel object in a single RGB or RGB-D image, as illustrated in Figure 1. First, we propose a novel approach for 6D pose refinement based on render & compare which enables generalization to novel objects. The shape and coordinate system of the novel object are provided as inputs to the network by rendering multiple synthetic views of the object’s CAD model. Second, we propose a novel method for coarse pose estimation which does not require knowledge of the object symmetries during training. The coarse pose estimation is formulated as a classification problem where we compare renderings of random pose hypotheses with an observed image, and predict whether the pose can be corrected by the refiner. Finally, we leverage the availability of large-scale 3D model datasets to generate a highly diverse synthetic dataset consisting of 2 million photorealistic images depicting over 20K models in physically plausible configurations. The code, dataset and trained models are available on the project page .

We show that our novel-object pose estimation method trained on our large-scale synthetic dataset achieves state-of-the-art performance on ModelNet . We also perform an extensive evaluation of the approach on hundreds of novel objects from all 7 core datasets of the BOP challenge and demonstrate that our approach achieves performance competitive with existing approaches that require access to the target objects during training.

Related work

In this section, we first review the literature on 6D pose estimation of known rigid objects. We then focus on the practical scenario similar to ours where the objects are not known prior to training.

6D pose estimation of known objects. Estimating the 6D pose of rigid objects is a fundamental computer vision problem that was first addressed using correspondences established with locally invariant features or template matching . These have been replaced by learning-based methods with convolutional neural networks that directly regress sets of sparse or dense features. All these approaches use non-learning stages relying on PnP+Ransac to recover the pose from correspondences in RGB images, or variants of the iterative closest point algorithm, ICP , when depth is available. The best performing methods rely on trainable refinement networks based on render & compare . These methods render a single image of the object, which is not sufficient to provide complete information on the shape and coordinate system of a 3D model to the network. This information is thus encoded in the networks weights when training, which leads to poor generalization when tested on novel objects unseen at training. Our approach renders multiple views of an object to provide this 3D information, making the trained network independent of these object-specific properties.

6D pose estimation of novel objects. Other works consider a practical scenario where the objects are not known in advance. Category-level 6D pose estimation is a popular problem in which CAD models of test objects are not known, but the objects are assumed to belong to a known category. These methods rely on object properties that are common within categories (e.g. handle of a mug) to define and estimate the object pose, and thus cannot generalize to novel categories. Our method requires the 3D model of the novel object instance to be known during inference, but does not rely on any category-level information. Other works address a scenario similar to ours. only estimate the 3D orientation of novel objects by comparing rendered pose hypotheses with the observed image using features extracted by a network. They rely on handcrafted or learning-based DeepIM refiners to recover accurate 6D poses. We instead propose a method that estimates the full 6D pose of the object and show our refinement network significantly outperforms DeepIM when tested on novel object instances. The closest works to ours are OSOP and ZePHyR . OSOP focuses on the coarse estimation by explicitly predicting 2D-2D correspondences between a single rendered view of the object and the observed image, and solves for the pose using PnP or Kabsch which makes inference slower and less robust compared to directly predicting refinement transforms with a network as done in our solution. ZePHyR strongly relies on the depth modality, whereas our approach can also be used in RGB-only images. Finally, investigate using a set of real reference views of the novel object instead of using a CAD model. These approaches have only reported results on datasets with limited or no occlusions. Our use of a deep render & compare network trained on a large-scale synthetic dataset displaying highly occluded object instances enables us to handle highly cluttered scenes with high occlusions like in the LineMOD Occlusion, HomebrewedDB or T-LESS datasets.

Method

Coarse pose estimation. Given an object detection, shown in Figure 1(b), the goal of the coarse pose estimator is to provide an initial pose estimate TCO,coarse\mathcal{T}_{\textit{CO},\text{coarse}} which is sufficiently accurate that it can then be further improved by the refiner. In order to generalize to novel-objects we propose a novel classification based approach that compares observed and rendered images of the object in a variety of poses and selects the rendered image whose object pose best matches the observed object pose.

Figure 2(a) gives an overview of the coarse model. At inference time the network consumes the observed image IoI_{o} along with rendered images {Ir(TCOj)}j=1M\{I_{r}(\mathcal{T}_{\textit{CO}}^{j})\}_{j=1}^{M} of the object in many different poses {TCOj}j=1M\{\mathcal{T}_{\textit{CO}}^{j}\}_{j=1}^{M}. For each pose TCOj\mathcal{T}_{\textit{CO}}^{j} the model predicts a score (Io,Ir(TCOj))→ξj(I_{o},I_{r}(\mathcal{T}_{\textit{CO}}^{j}))\rightarrow\xi_{j} that classifies whether the pose hypothesis is within the basin of attraction of the refiner. The highest scoring pose TCOj∗,j∗=argmax⁡jξj\mathcal{T}_{\textit{CO}}^{j^{*}},j^{*}=\operatorname{argmax}_{j}\xi_{j} is used as the initial pose for the refinement step. Since we are performing classification, our method can implicitly handle object symmetries, as multiple poses can be classified as correct.

Pose refinement model. Given an input image and an estimated pose, the refiner predicts an updated pose estimate. Starting from a coarse initial pose estimate TCO,coarse\mathcal{T}_{\textit{CO},\text{coarse}} we can iteratively apply the refiner to produce an improved pose estimate. Similar to our refiner takes as input observed IoI_{o} and rendered images Ir(TCOk)I_{r}(\mathcal{T}_{\textit{CO}}^{k}) and predicts an updated pose estimate TCOk+1\mathcal{T}_{\textit{CO}}^{k+1}, see Figure 2 (b), where kk refers to the kthk^{th} iteration of the refiner. Our pose update uses the same parameterization as DeepIM and CosyPose which disentangles rotation and translation prediction. Crucially this pose update ΔT\Delta\mathcal{T} depends on the choice of an anchor point O\mathcal{O}, see Appendix for more details. color=orange!30,size=,tickmarkheight=4pt]Lucas: Add proof in appendix In prior work which trains and tests on the same set of objects, the network can effectively learn the position of the anchor point O\mathcal{O} for each object. However in order to generalize to novel objects we must enable the network to infer the anchor point O\mathcal{O} at inference time.

In order to provide information about the anchor point to the network we always render images Ir(TCOk)I_{r}(\mathcal{T}_{\textit{CO}}^{k}) such that the anchor point O\mathcal{O} projects to the image center. Using rendered images from multiple distinct viewpoints {TCO,i}i=1N\{\mathcal{T}_{\textit{CO},i}\}_{i=1}^{N} the network can infer the location of the anchor point O\mathcal{O} as the intersection point of camera rays that pass through the image center, see Figure 2(b).

Additional information about object shape and geometry can be provided to the network by rendering depth and surface normal channels in the rendered image IrI_{r}. We normalize both input depth (if available) and rendered depth images using the currently estimated pose TCOk\mathcal{T}_{\textit{CO}}^{k} to assist the network in generalizing across object scales, see Appendix for more details.

Network architecture. Both the coarse and refiner networks consists of a ResNet-34 backbone followed by spatial average pooling. The coarse model has a single fully-connected layer that consumes the backbone feature and outputs a classification logit. The refiner network has a single fully-connected layer that consumes the backbone feature and outputs 9 values that specify the translation and rotation for the pose update.

2 Training Procedure

Training data. For training, both the coarse and refiner models require RGB(-D)Our method can consume either RGB or RGB-D images depending on the input modalities that are available. images with ground-truth 6D object pose annotations, along with 3D models for these objects. In order for our approach to generalize to novel-objects we require a large dataset containing diverse objects. All of of our methods are trained purely on synthetic data generated using BlenderProc . We generate a dataset of 2 million images using a combination of ShapeNet (abbreviated as SN) and Google-Scanned-Objects (abbreviated as GSO) . Similar to the BOP synthetic data, we randomly sampled objects from our dataset and dropped them on a plane using a physics simulator. Materials, background textures, lighting and camera positions are randomized. Example images can be seen in Figure 1(a) and in the Appendix. Some of our ablations also use the synthetic training datasets provided by the BOP challenge . We add data augmentation similar to CosyPose to the RGB images which was shown to be a key to successful sim-to-real transfer. We also apply data augmentation to the depth images as explained in the appendix.

Refiner model. The refiner model is trained similarly to . Given an image with an object M\mathcal{M} at ground-truth pose TCO∗\mathcal{T}_{\textit{CO}}^{*} we generate a perturbed pose TCO′\mathcal{T}_{\textit{CO}}^{\prime} by applying a random translation and rotation to TCO∗\mathcal{T}_{\textit{CO}}^{*}. Translation is sampled from a normal distribution with a standard deviations of (0.02,0.02,0.05)(0.02,0.02,0.05) centimeters and rotation is sampled as random Euler angles with a standard deviation of 1515 degrees in each axis. The network is trained to predict the relative transformation between the initial and target pose. Following we use a loss that disentangles the prediction of depth, xx-yy translation, and rotation. See the appendix for more details.

Coarse model. Given an input image IoI_{o} of an object M\mathcal{M} and a pose TCO′\mathcal{T}_{\textit{CO}}^{\prime} the coarse model is trained to classify whether pose TCO′\mathcal{T}_{\textit{CO}}^{\prime} is within the basin of attraction of the refiner. In other words, if the refiner were started with the initial pose estimate TCO′\mathcal{T}_{\textit{CO}}^{\prime} would it be able to estimate the ground-truth pose via iterative refinement? Given a ground-truth pose-annotation TCO∗\mathcal{T}_{\textit{CO}}^{*} we randomly sample poses TCO′\mathcal{T}_{\textit{CO}}^{\prime} by adding random translation and rotation to TCO∗\mathcal{T}_{\textit{CO}}^{*}. The positives are sampled from the same distribution used to generate the perturbed poses the refiner network is trained to correct (see above), and other poses sufficiently distinct to this one (see the appendix for more details) are marked as negatives. The model is then trained with binary cross entropy loss.

Experiments

We evaluate our method for 6D pose estimation of novel objects using the seven challenging datasets of the BOP 6D pose estimation benchmark, and the ModelNet dataset. The dataset and the standard 6D pose estimation metrics we use are detailed in Section 4.1. In all our experiments, the objects are considered novel, i.e. they are only available during inference on a new image and they are not used during training. In Section 1, we evaluate the performance of our approach composed of coarse and refinement networks. Notably, we show that (i) our method is competitive with others that require the object models to be known in advance, and (iii) our refiner outperforms current state-of-the-art on the ModelNet and YCB-V datasets. Section 4.3 validates our technical contributions and shows the crucial importance of the training data in the success of our method. Finally, we discuss the limitations in Section 4.4.

We consider the seven core datasets of the BOP challenge : LineMod Occlusion (LM-O) , T-LESS , TUD-L , IC-BIN , ITODD , HomebrewedDB (HB) and YCB-Video (YCB-V) . These datasets exhibit 132 different objects in cluttered scenes with occlusions. These objects present many factors of variation: textured or untextured, symmetric or asymmetric, household or industrial (e.g. watcher pitcher, stapler, bowls, multi-socket plug adaptor) which makes them representative of objects that are typically encountered in robotic scenarios. The ModelNet dataset depicts individual instances of objects from 7 classes of the ModelNet dataset (bathtub, bookshelf, guitar, range hood, sofa, tv stand and wardrobe). We use initial poses provided by adding noise to the ground truth, similar to previous works . The focus is on refining these inital poses. We follow the evaluation protocol of for BOP datasets, and of DeepIM for ModelNet.

2 6D pose estimation of novel objects

Performance of coarse+refiner. Table 1 reports results of our novel-object pose estimation method on the BOP datasets. We first use the detections and pose hypotheses provided by a combination of PPF and SIFT, similar to the state-of-the-art method Zephyr . For each object detection, these algorithms provide multiple pose hypotheses. We find the best hypothesis using the score of our coarse network, and apply 5 iterations of our refiner. Results are reported in row 7. On YCB-V, our method achieves a +10.7 AR score improvement compared to Zephyr (row 6). Averaged across the YCB-V and LM-O datasets, the AR score of our approach is 59.7 compared to 55.7 for Zephyr (row 6). Next, we provide a complete set of results using the detections from Mask-RCNN networks. Please note that since detection of novel objects is outside the scope of this paper, we use the networks trained on the synthetic PBR data of the target objects which are publicly available for each dataset. We report the results of our coarse estimation strategy (Table 1, row 11), and after running the refiner network, on RGB (row 12) or RGB-D (row 13) images. We observe that (i) our refinement network significantly improves the coarse estimates (+41.0 mean AR score for our RGBD refiner) and (ii) the performance of both models is competitive with the learning-based refiner of CosyPose (row 1) while not requiring to be trained on the test objects. The recent SurfEmb performs better than our approach, but heavily relies on the knowledge of the objects for training and cannot generalize to novel objects.

Performance of the refiner. We now focus on the evaluation of our refiner which can be used to refine arbitrary initial poses. Our refiner is the only learning-based approach in Table 1 which can be applied to novel objects. In rows 9 and 10, we apply our refiner to the coarse estimates of CosyPose (row 11). Again, we observe that our refiner significantly improves the accuracy of these initial pose estimates (+23.7 in average for the RGB-D model). Notably, the RGB-only method (row 9) performs better than the CosyPose refiner (row 1) on average, while not having seen the BOP objects during training. This is thanks to our large-scale training on thousands of various objects, while CosyPose is only trained on tens of objects for each dataset.

One iteration of our refiner takes approximately 50 milliseconds on a RTX 2080 GPU, making it suitable for use in an online tracking application. Five iterations of our refiner are also 5 times faster than the object-specific refiner of SurfEmb which takes around 1 second per image crop. Finally, we evaluate our refiner on ModelNet and compare it with the state-of-the-art methods MP-AAE and LatentFusion . For this experiment, we remove the ShapeNet categories that overlap with the test ones in ModelNet from our training set in order to provide a fair comparison on novel instances and novel categories similar to . Results reported in Table 2 show that our refiner significantly outperforms existing approaches across all metrics.

3 Ablations

In this section we perform ablations of our approach to validate our main contributions. Additional ablations are in the appendix. For these ablations, we consider the RGB-only refiner and re-train several models with different configurations of hyper-parameters and training data.

Encoding the anchor point and object shape. As discussed in Section 3.1 the refiner must have information about the anchor point O\mathcal{O} in order to generalize to novel objects. We accomplish this by using 4 rendered views pointing towards the anchor point, see Figure 2(b). Table 3(a) shows the performance of the refiner increases as we increase the number of views from 11 to 44, validating our design choice. Multiple views may also help the network to understand the object’s appearance from alternate viewpoints, thus potentially helping the refiner to overcome large initial pose errors. We also validate our choice to provide a normal map of the object to the network. This information can help the network use subtle object appearance variations that are only visible under different illumination like the details on the cross of a guitar.

Number of training objects. We now show the crucial role of the training data. We report in Table 3 (b) the results for our refiner trained on an increasing number of CAD models. The performance steadily increases with the number of objects, which validates that training on a large number of object models is important to generalize to novel ones. These results also suggest that the performance of our approach could be improved as more datasets of high-quality CAD models like GSO become available.

Variety in the training objects. Next, we restrict the training to different sets of objects. We observe in the bottom of Table 3(b) that models from the GoogleScannedObjects are more important to the performance of the method on the BOP dataset compared to using both ShapeNet and GSO. We hypothesize this is due to the presence of high-quality textured objects in the GSO dataset. Finally, we train our model on the 132 objects of the BOP dataset. When testing on the same BOP objects, the performance benefits from knowing these objects during training is small compared to using our GSO+ShapeNet or GSO dataset.

4 Limitations

While MegaPose shows promising results in robot experiments (please see the supplementary video) and 6D pose estimation benchmarks, there is still room for improvement. We illustrate the failure modes of our approach in the supplementary material. The most common failure mode is due to inaccurate initial pose estimates from the coarse model. The refiner model is trained to correct poses that are within a constrained range but can fail if the initial error is too large. There exist multiple potential approaches to alleviate this problem. We can increase the number of pose hypotheses MM at the expense of increased inference time, improve the accuracy of the coarse model, and increase the basin of attraction for refinement model. Another limitation is the runtime of our coarse model. We use M=520M=520 pose hypotheses per object which takes around 2.5 seconds to be rendered and evaluated by our coarse model. In a tracking scenario however, the coarse model is run just once at the initial frame and the object can be tracked using the refiner which runs at 20Hz. Additionally, our refiner could also be coupled with alternate coarse estimation approaches such as to achieve improved runtime performance.

Conclusion

We propose MegaPose, a method for 6D pose estimation of novel objects. Megapose can estimate the 6D pose of novel objects given a CAD model of the object available only at test time. We quantitatively evaluated MegaPose on hundreds of different objects depicted in cluttered scenes, and performed ablation studies to validate our network design choices and highlight the importance of the training data. We release our models and large-scale synthetic dataset to stimulate the development of novel methods that are practical to use in the context of robotic manipulation where rapid deployment to new scenes with new objects is crucial. While this work focuses on the coarse estimation and fine refinement of an object pose, detecting any unknown object given only a CAD model is still a difficult problem that remains to be solved for having a complete framework for detection and pose estimation of novel objects. Future work will address zero-shot object detection using our large-scale synthetic dataset.

Acknowledgements

This work was partially supported by the HPC resources from GENCI-IDRIS (Grant 011011181R2), the European Regional Development Fund under the project IMPACT (reg. no. CZ.02.1.01/0.0/0.0/15 003/0000468), EU Horizon Europe Programme under the project AGIMUS (No. 101070165), Louis Vuitton ENS Chair on Artificial Intelligence, and the French government under management of Agence Nationale de la Recherche as part of the ”Investissements d’avenir” program, reference ANR-19-P3IA-0001 (PRAIRIE 3IA Institute).

References

Appendix

This appendix is organized as follows. In section A, we provide the equations of the pose update predicted by the refiner network, and show it depends on the anchor point. In section B, we give details on the loss used to train the refiner network. Section C explains the normalization strategy we apply to the observed and rendered depth images of the RGB-D refiner. Section D details the pose hypotheses used during training and inference of the coarse model. In section E, we provide examples of training images and give details on the data augmentation and training hardware. In section 4, we perform additional ablations to validate (i) the contributions of our coarse network, (ii) the choice of hyper-parameter MM. We also provide details on the robot experiments shown in the supplementary video. In section G, we illustrate qualitatively that our approach is robust to illumination condition variations. Section H illustrates the main failure modes of our approach. Finally, section I investigates the robustness of our approach with respect to an incorrect 3D model.

The supplementary video shows predictions of our approach on real images. We apply our approach in tracking mode on several videos. Tracking consists in running the coarse estimator on the first frame of a video sequence, and then applying one iteration of the refiner on each new image, using the prediction in the previous image as the pose initialization at the input of refinement network. This approach can process 20 images per second. The video notably demonstrates the method is robust to occlusion and can be used to perform visually guided robotic manipulation of novel objects.

Appendix A Pose update and anchor point

Pose update. We use the same pose update as DeepIM and CosyPose . The network predicts 9 values corresponding to one 3-vector [vx,vy,vz][v_{x},v_{y},v_{z}] to predict an update of the translation of a 3D anchor point, and two 3-vectors e1,e2e_{1},e_{2} that define a rotation update explained below. The pose update consists in updating (i) the position of a 3D reference point O\mathcal{O} attached on the object, and (ii) the rotation matrix RCOR_{CO} of the object frame expressed in the camera frame (please note the different notations for the anchor point O\mathcal{O} and the object frame OO):

where [xOk,yOk,zOk][x^{k}_{\mathcal{O}},y^{k}_{\mathcal{O}},z^{k}_{\mathcal{O}}] is the 3D position of the anchor point expressed in camera frame at iteration kk, RCOkR^{k}_{CO} a rotation matrix describing the objects orientation expressed in camera frame, fxCf_{x}^{C} and fyCf_{y}^{C} are the (known) focal lengths that correspond to the (virtual) camera associated with the cropped observed image, and R(e1,e2)R(e_{1},e_{2}) is a rotation matrix describing the rotation update recovered from e1,e2e_{1},e_{2} using by orthogonalizing the basis defined by the two predicted rotation vectors e1,e2e_{1},e_{2} similar to . Finally, [xOk+1,yOk+1,zOk+1][x^{k+1}_{\mathcal{O}},y^{k+1}_{\mathcal{O}},z^{k+1}_{\mathcal{O}}] and RCOk+1R^{k+1}_{CO} are, respectively, the translation and rotation after applying the pose update. The 3D translation of the anchor point and the rotation matrix RCOR_{CO} are used to define the pose the object.

Dependency to the anchor point. We now show that the predictions the network must make to correct a pose error between an initial pose TCOk\mathcal{T}_{CO}^{k} and a target pose TCOk+1\mathcal{T}_{CO}^{k+1} is independent of the choice of the orientation of the objects coordinate frame OO but depends on the choice anchor point O\mathcal{O}. Let us denote O1,O2\mathcal{O}^{1},\mathcal{O}^{2} two different anchor points, and RCO1,RCO2R_{CO^{1}},R_{CO^{2}} the rotation matrices of the object (expressed in the fixed camera frame) for two different choices of object coordinate frames O1O^{1} and O2O^{2}. We note tO1O2=O2−O1=[x12,y12,z12]t_{\mathcal{O}^{1}\mathcal{O}^{2}}=\mathcal{O}^{2}-\mathcal{O}^{1}=[x_{12},y_{12},z_{12}] the 3D translation vector between O2\mathcal{O}^{2} and O1\mathcal{O}^{1}; and RO1O2=RCO1TRCO2R_{O^{1}O^{2}}=R_{CO^{1}}^{T}R_{CO^{2}} the rotation of coordinate frame O2O^{2} expressed in O1O^{1}. For one choice of anchor point and object frame, e.g. O1\mathcal{O}_{1} and RCO1R_{CO^{1}}, we derive the predictions the network has to make to correct the error using equations (1),(2),(3),(4):

From eq. (12), we have R1(R2)T=IdR^{1}\left(R^{2}\right)^{T}=Id. In other words, the rotation matrices that the network must predict to correct the errors in scenarios 11 and 22 are the same. The network predictions for the rotation components thus do not depend on the choice of the choice of object coordinate system. However the other components of the translation cannot be simplified further. For example, derivations of eq. (11) leads to vz1−vz2=z12(zO1k+1−zO1k)z1k(z1k+z12)v_{z}^{1}-v_{z}^{2}=\frac{z_{12}\left(z_{\mathcal{O}^{1}}^{k+1}-z_{\mathcal{O}^{1}}^{k}\right)}{z_{1}^{k}(z_{1}^{k}+z_{12})} which is non-zero in the general case where O1O^{1} and O2O^{2} are different and there is an error between the initial and target poses. This proves that different choices of anchor point leads to different predictions. For the network to generalize to a novel object, the network be able to infer the 3D position of the anchor point on this object. We achieve this by rendering multiple views of the objects in which the anchor point reprojects to the center of each image as explained in Section 3.1 of the main paper.

Appendix B Refiner loss

Our refiner network is trained using the same loss as in CosyPose , but without using symmetry information on the objects because it is not typically not available for large-scale datasets of CAD models like ShapeNet or GoogleScannedObjects. We first define the distance DO(T1,T2)D_{O}(\mathcal{T}_{1},\mathcal{T}_{2}) to measure the distance between two 6D poses represented by transformations T1\mathcal{T}_{1} and T2\mathcal{T}_{2} using the 3D points XO\mathcal{X}_{O} of an object O:

where ∣⋅∣|\cdot| is the L1L_{1} norm. In practice, we uniformly sample 20002000 points on the surface of an object’s CAD model to compute this distance. We also define the pose update function FF which takes as input the initial estimate of the pose TCOk\mathcal{T}_{CO}^{k}, the predictions of the neural network [vx,vy,vz][v_{x},v_{y},v_{z}] and RR, and outputs the updated pose:

where the closed form of function FF is expressed in equations (1) (2) (3) (4). We also write [vx⋆,vy⋆,vz⋆][v_{x}^{\star},v_{y}^{\star},v_{z}^{\star}] and R⋆R^{\star} as the target predictions, i.e. the predictions such that TCO⋆=F(TCOk,[vx⋆,vy⋆,vz⋆],R⋆)\mathcal{T}_{CO}^{\star}=F(\mathcal{T}_{CO}^{k},[v_{x}^{\star},v_{y}^{\star},v_{z}^{\star}],R^{\star}), where TCO⋆\mathcal{T}_{CO}^{\star} is the ground truth camera-object pose. The loss used to train the refiner is the following:

where DOD_{O} is the distance defined in eq. (13) and KK is the number of training iterations. The different terms of this loss separate the influence of: xyxy translation (15), relative depth (16) and rotation (17). We sum the loss over K=3K=3 refinement iterations to imitate how the refinement algorithm is applied at test time but the error gradients are not backpropagated through rendering and iterations. For simplicity, we write the loss for a single training sample (i.e. a single object in an image), but we sum it over all the samples in the training set.

Appendix C Depth normalization

When depth measurements are available, the observed depth image and depth images of the renderings are concatenated with the images, as mentioned in Section 3.1 of the main paper. At test time, the objects may be observed at different depth outside of the training distribution. In order for the network to become invariant to the absolute depth values of the inputs, we normalize both observed and rendered depth. Let us denote DD a depth image (rendered or observed are treated similarly). We apply the following operations to DD. (i) Clipping of the metric depth values of DD to lie between and zOk+1z^{k}_{\mathcal{O}}+1, where zOkz^{k}_{\mathcal{O}} is the depth of the anchor point on the object in the input pose at iteration kk:

Appendix D Pose hypotheses in the coarse model

Training hypotheses. Given the ground truth object pose TCO⋆\mathcal{T}_{CO}^{\star}, we generate a perturbed pose TCO′\mathcal{T}_{CO}^{\prime} by applying random translation and rotation to TCO′\mathcal{T}_{CO}^{\prime}. The parameters of this (small) perturbation are sampled from the same distribution as the distribution used to sample the perturbed poses the refiner network is trained to correct. The translation is sampled from a normal distribution with a standard deviations of (0.02,0.02,0.05)(0.02,0.02,0.05) centimeters and rotation is sampled as random Euler angles with a standard deviation of 1515 degrees in each axis.

We then define several poses that depend on TCO′\mathcal{T}_{CO}^{\prime} and cover a large variety of viewing angles of the object. We define a cube of size 2zO′2z^{\prime}_{\mathcal{O}}, where zO′z^{\prime}_{\mathcal{O}} is the zz component of the 3D translation in the pose TCO′\mathcal{T}_{CO}^{\prime}. The CAD model of the object observed under orientation RCO′R_{CO}^{\prime} is placed at the center of the cube. We then place 26 cameras at the locations of each corner, half-side and face centers of the cube. By construction, one of these cameras, which we denote C0C^{0}, has the same camera-object orientation as TCO′\mathcal{T}_{CO}^{\prime}, and all others {Ci}i=1..25\{C^{i}\}_{i=1..25} correspond to cameras observing the object under viewpoints which are sufficiently far from RC0OR_{C^{0}O} and outside the basin of attraction of the refiner by construction. In addition, we apply inplane rotations of 90∘90^{\circ}, 180∘180^{\circ} and 270∘270^{\circ} to each camera, which leads to a total of 26∗4=10426*4=104 cameras with one positive and 103 negatives.

We mark TC0O\mathcal{T}_{C^{0}O} as a positive for the coarse model because the error between TC0O\mathcal{T}_{C^{0}O} and TCO⋆\mathcal{T}_{CO}^{\star} lies within the basin of attraction of the refiner. All other cameras are marked as negatives. During training, the positives account for around 30%30\% of the total numbers of images in a mini-batch.

where fxf_{x} and fyf_{y} are the (known) focal lengths of the camera. We then update the depth estimate zOpguessz_{\mathcal{O}^{p}}^{guess} using the following simple strategy. We project the points of the object 3D model using RpR^{p} and the initial guess of the 3D position of the anchor point we have just defined. These points define a bounding box with dimensions Δuguess,x=(Δuguess,x,Δuguess,y)\Delta u_{guess,x}=(\Delta u_{guess,x},\Delta u_{guess,y}) and the center remains unchanged uguess=udetu_{guess}=u_{det} by construction. We compute an updated depth of the anchor point such that its width and height approximately match the size of the 2D detection:

and use this new depth to compute xOpx_{\mathcal{O}^{p}} and yOpy_{\mathcal{O}^{p}} using equations (20) and (21) that were used to define xOpguessx_{\mathcal{O}^{p}}^{guess} and yOpguessy_{\mathcal{O}^{p}}^{guess}. The rotation RpR^{p} and 3-vector [xOp,yOp,zOp][x_{\mathcal{O}^{p}},y_{\mathcal{O}^{p}},z_{\mathcal{O}^{p}}] define the pose of hypothesis pp. We then use the same strategy used to define the training hypotheses (described above) in order to define 103103 additional viewpoints depending on pp. We repeat the operation P=5P=5 times, for a total of 5∗104=5205*104=520 pose hypotheses.

Appendix E Training details

We generate 2 million photorealistic images using BlenderProc as explained in Section 3.2 of the main paper. Randomly sampled images from the training set are shown in Figure 4.

Data augmentation.

We apply data augmentation to the synthetic images during training. We use the same data augmentation as CosyPose for the RGB images. It includes Gaussian blur, contrast, brightness, colors and shaprness filters from the Pillow library . For the depth images, we take inspiration from the augmentations used in . Augmentations include blur, ellipse dropout, correlated and uncorrelated noise.

GPU hardware and training time.

Training time is respectively 32 and 48 hours for the coarse and refiner models using 32 V-100 GPUs. This training is performed once, and estimating the pose of novel objects does not require any fine-tuning on the target objects.

Appendix F Additional experiments.

In order to evaluate the validate the contributions of our coarse scoring network, we use a set of pose hypotheses generated for novel objects by the commercial Halcon 20.05 Progress software which implements the PPF algorithm described in . Note that these are the same pose hypotheses used in Zephyr . We then find the best hypotheses using the scores of PPF, the scoring network of Zephyr or our coarse network, and report AR results for the LM-O and YCB-V datasets in the table 4. On both datasets, our coarse network is better than the two baselines (PPF and Zephyr) for selecting the best poses among a given set of hypotheses.

Classification-based coarse network.

To validate our classification-based coarse model, we consider a regression-based alternative. We trained a regression-based network similar to the coarse model of CosyPose which takes as input six views of the objects covering viewpoints at the poles of a sphere centered on the object. The network collapsed during training, leading to large errors that cannot be recovered by the refiner and a performance close to zero on the BOP datasets. We hypothesize this failure is due to the presence of symmetric objects in our training set which leads to ambiguous gradients during training. This failure could also be attributed to other factors, such as the difficulty to interpret the full 3D geometry of an object with a CNN given six views of its 3D model captured under distant viewpoints.

Number of coarse pose hypotheses.

MM is an important parameter of our method, which can be used to choose a trade-off between running time and accuracy. The performance significantly improves from M=104 to M=520 (+11.4 AR on BOP5) while keeping the running-time of the coarse model reasonable (1.6 seconds for M=520 compared to 0.3 seconds for M=120). Above M=520, the performance improvement is marginal, e.g. (+0.9 AR) for 4608 hypotheses. Please note that we are still making improvements to our code and have lower runtimes than reported in the paper (1.6s for M=520 compared to the 4s mentioned in line 276).

Robotic grasping experiments.

We performed a qualitative real-robot grasping experiment. For multiple YCB-V objects, we manually annotated one grasp with respect to the object’s coordinate frame. We then placed the considered object (e.g. the drill in the supplementary video) in a scene among other objects representing visual distractors. The object may be placed on the table or on another object. We then take a single RGB image of the scene using a RealSense D415 camera mounted on the gripper of a Franka Emika Panda robot. We detect the object in 2D using the Mask-RCNN detector from CosyPose , and run our Megapose approach composed of coarse and refiner modules for estimating the 6D pose of the object with respect to the camera. We then express the 6D pose of the object and grasp with respect to the robot using the known camera-to-robot extrinsic calibration. We then use a motion planner to generate a robot motion that reaches the estimated grasp pose with the gripper and lift the object. This experiment shown in the supplementary video shows that the pose estimates are of sufficiently high quality to be useful for a robotic manipulation task.

Appendix G Robustness to illumination conditions

In Figure 5, we show qualitative predictions of our approach for the watering can on the TUD-L dataset. Please notice the high accuracy of our approach despite challenging illumination conditions.

Appendix H Failure modes and performance on specific types of objects

We carry out a per-object analysis of the performance of our approach on the YCB-V dataset. For each of the 21 objects of the dataset, we report the percentage of predictions for which the error with the ground truth is within a threshold of 15∘15^{\circ} in rotation and 5cm5\textrm{cm} in translation. Results are reported in Figure 6.

Next, we illustrate the main failure modes of our approach using a set of objects which have a performance below average on this dataset. Examples of failure cases are presented in Figure 7. We observed three main failure modes to our approach. First, we observe the orientation of a novel object may be incorrectly predicted if the object has a similar visual appearance under different viewpoint. We observed this failure mode in particular for textureless objects such as a red bowl that appears similar whether it is standing upside or it is flipped. Second, we observe that our approach may fail to disambiguate the pose of objects that are asymmetric but for which it is necessary to look at fine details on the objects to disambiguate multiple possible poses. An example is a pair of scissors which have left and right handles with slightly different dimensions. In both of these failure modes, we observed that our refiner gets stuck into a local minimal due to an inaccurate coarse estimate outside of the basin of attraction of our refiner model. Finally, using a CAD model with incorrect scale leads to an incorrect estimation of the depth of the object due to the object scale/depth ambiguity in RGB images. We observe for example that the translation estimates of the wooden block of YCB-V have systematically large error despite the rendering of our prediction correctly matching the contours of the object in the observed image. This is because the scale of the CAD model of the wooden block publicly available does not match the correct dimensions of the real object which was used for annotating the ground truth.

Appendix I 3D model quality

Our approach can be applied even if the 3D model of the object does not exactly matches the real object. In figure 8, we show examples of correctly estimated poses using low-fidelity CAD models with low-quality textures or geometric discrepancies between the real object and its 3D model.