GDRNPP: A Geometry-guided and Fully Learning-based Object Pose Estimator

Xingyu Liu, Ruida Zhang, Chenyangguang Zhang, Gu Wang, Jiwen Tang, Zhigang Li, Xiangyang Ji

Introduction

Estimating the 6D pose, \iethe 3D rotation and 3D translation, of objects with respect to the camera is a fundamental problem in computer vision. It has wide applicability to many real-world tasks such as robotic manipulation , augmented reality and autonomous driving . Most traditional methods rely on depth data for this task , while monocular methods lagged considerably behind . Nonetheless, with the advent of deep learning and especially the rise of Convolutional Neural Networks (CNNs), accuracy and robustness of monocular 6D object pose estimation have been consistently improving, even at times surpassing methods relying on depth data .

Different strategies for predicting 6D pose from monocular data have been proposed. For instance, learning of an embedding space for pose or direct regression of the 3D rotation and translation . While these methods generally perform well, they usually lack in accuracy when compared with approaches that instead rely on establishing 2D-3D correspondences prior to estimating the 6D pose .

Differently, this latter class of methods usually involves solving the 6D pose through a variant of the PnnP/RANSAC algorithm. While such a paradigm provides good estimates, it also suffers from several drawbacks. First, these methods are usually trained with a surrogate objective for correspondence regression, which does not necessarily reflect the actual 6D pose error after optimization. In practice, two sets of correspondences can have the same average error while describing completely different poses. Second, these approaches are not differentiable with respect to the estimated 6D pose, which limits learning. For instance, these methods cannot be coupled with self-supervised learning from unlabeled real data , as they require the computation of the pose to be fully differentiable in order to obtain a signal between data and pose. Finally, the RANSAC iterative process can be very time-consuming when dealing with dense correspondences.

To summarize, while methods grounded on 2D-3D correspondences are dominating the field, they still exhibit downsides due to the decoupling of the problem into two separate steps one of which not differentiable. Consequently, some efforts have been devoted to enabling backpropagation through the PnnP/RANSAC stage. However, this either requires a complex training strategy in order to have good initialization of scene coordinates , or can only handle sparse correspondences of a predefined set of keypoints . Recently, the authors of proposed to leverage PointNet in order to approximate PnnP for sparse correspondences. While this proves to work well, PointNet disregards the fact that the correspondences are organized with respect to the image pixels, which tends to strongly deteriorate the performance as shown in .

In this work, we propose to overcome these limitations by establishing 2D-3D correspondences whilst computing the final 6D pose estimate in a fully differentiable way (Fig. 1). In its core, we propose to learn the PnnP optimization, exploiting the fact that the correspondences are organized in image space, which gives a significant boost in performance, outperforming all prior works. To summarize, we make the following contributions:

We revisit the key ingredients in direct 6D pose regression and observe that by choosing appropriate representations for the pose parameters, methods based on direct regression show competitive performance compared with state-of-the-art correspondence-based indirect methods.

We further propose a simple yet effective Geometry-guided Direct Regression Network (GDR-Net) to boost the performance of direct 6D pose regression via leveraging the geometric guidance from dense correspondence-based intermediate representations.

Extensive experiments on LM , LM-O , and YCB-V datasets show that our unified GDR-Net approach achieves accurate, yet real-time and robust, monocular 6D object pose estimation.

Related Work

While methods based on depth data used to dominate the field of 6D pose estimation , Deep Learning-based methods have recently demonstrated promising results for the task at hand . In this section, we review several commonly employed strategies for monocular 6D pose estimation.

Indirect Methods. The most popular approach is to establish 2D-3D correspondences, which are then leveraged to solve for the 6D pose using a variant of the RANSAC-based PnnP algorithm. For instance, and compute the 2D projections of a set of fixed control points (\egthe 3D corners of the encapsulating bounding box). To enhance the robustness, and additionally conduct segmentation coupled with voting for each correspondence. However, the recent trend goes towards predicting dense rather than sparse correspondences . Moreover, while leverages a GAN on top of dense correspondences to increase stability, makes use of fragments in order to account for ambiguities in pose.

Another orthogonal line of works aims at learning a latent embedding of pose which can be utilized for retrieval during inference. These embeddings are commonly either grounded on metric learning employing a triplet loss , or via training of an Auto-Encoder .

Direct Methods. Although indirect methods, leveraging 2D-3D correspondences, are currently performing better, they cannot be directly employed in many tasks, which require the pose estimation to be differentiable . Hence, some methods directly regress the 6D pose, either leveraging a point matching loss or employing separate loss terms for each component . Other methods discretize the pose space and conduct classification rather than regression . A few methods also try to solve a proxy task during optimization. Thereby, proposes to employ an edge-alignment loss using the distance transform, while harnesses differentiable rendering to allow training on unlabeled samples.

Differentiable Indirect Methods. Recently, a few works attempt to make PnnP/RANSAC differentiable. In the authors introduce a novel differentiable way to apply RANSAC via sharing of hypotheses based on the predicted distribution. Nonetheless, these approaches require a complex training strategy, as they expect a good initialization for the scene coordinates. As for PnnP, employs the Implicit Function Theorem to enable the computation of analytical gradients \wrtthe pose loss. Yet, it is computationally expensive especially given too many correspondences since PnnP/RANSAC is still needed for both training and inference. Instead, attempts to learn the PnnP stage with a PointNet-based architecture which learns to infer the 6D pose from a fixed set of sparse 2D-3D correspondences.

Beyond Instance-level 6D Pose Estimation. Noteworthy, a few methods are recently trying to go beyond the instance-level scenario, even estimating the pose , sometimes paired with shape , for previously unseen objects.

Method

Given an RGB image II and a set of NN objects O={ Oi∣i=1,⋯ ,N }\mathcal{O}=\{\,\mathcal{O}_{i}\mid i=1,\cdots,N\,\} together with their corresponding 3D CAD models M={ Mi∣i=1,⋯ ,N }\mathcal{M}=\{\,\mathcal{M}_{i}\mid i=1,\cdots,N\,\}, our goal is to estimate the 6D object pose P=[R∣t]\mathbf{P}=[\mathbf{R}|\mathbf{t}] \wrtthe camera for each object O\mathcal{O} present in II. Notice that R\mathbf{R} describes the 3D rotation and t\mathbf{t} denotes the 3D translation of the detected object.

Fig. 2 presents a schematic overview of the proposed methodology. In the core, we first detect all objects of interest using an off-the-shelf object detector, such as . For each detection, we then zoom in to the corresponding Region of Interest (RoI) and feed it to our network to predict several intermediate geometric feature maps. Finally, we directly regress the associated 6D object pose from the dense correspondence-based intermediate geometric features.

In the following, we first (Sec. 3.1) revisit the key ingredients of direct 6D object pose estimation methods. Afterwards (Sec. 3.2), we illustrate a simple yet effective Geometry-Guided Direct Regression Network (GDR-Net) which unifies regression-based direct methods and geometry-based indirect methods, thus harnessing the best of both worlds.

Direct 6D pose estimation methods usually differ in one or more of the following components. Firstly, the parameterization of the rotation R\mathbf{R} and translation t\mathbf{t}, and secondly, the employed loss for pose. In this section, we investigate different commonly used parameterizations and demonstrate that appropriate choices have significant impact on the 6D pose estimates.

Parameterization of 3D Rotation. Several different parameterization can be employed to describe 3D rotations. Since many representations exhibit ambiguities, \ieRi\mathbf{R}_{i} and Rj\mathbf{R}_{j} describe the same rotation with Ri≠Rj\mathbf{R}_{i}\neq\mathbf{R}_{j}, most works rely on parametrizations that are unique to help training. Therefore, common choices are unit quaternions , log quaternions , or Lie algebra-based vectors .

Nevertheless, it is well-known that all representations with four or fewer dimensions for 3D rotation have discontinuities in the Euclidean space. When regressing a rotation, this introduces an error close to the discontinuities which becomes often significantly large. To overcome this limitation, proposed a novel continuous 6-dimensional representation for R\mathbf{R} in SO(3)SO(3), which has proven promising . Specifically, the 6-dimensional representation R6d\mathbf{R}_{\text{6d}} is defined as the first two columns of R\mathbf{R}

Given a 6-dimensional vector R6d=[r1∣r2]\mathbf{R}_{\text{6d}}=[\mathbf{r}_{1}|\mathbf{r}_{2}], the rotation matrix R=[R⋅1∣R⋅2∣R⋅3]\mathbf{R}=[\mathbf{R}_{\boldsymbol{\cdot}1}|\mathbf{R}_{\boldsymbol{\cdot}2}|\mathbf{R}_{\boldsymbol{\cdot}3}] can be computed according to

where ϕ(∙)\phi(\bullet) denotes the vector normalization operation.

Given the advantages of this representation, in this work we employ R6d\mathbf{R}_{\text{6d}} to parameterize the 3D rotation. Nevertheless, in contrast to , we propose to let the network predict the allocentric representation of rotation Ra6d\mathbf{R}_{\text{a6d}}. This representation is favored as it is viewpoint-invariant under 3D translations of the object. Hence, it is more suitable to deal with zoomed-in RoIs. Note that the egocentric rotation can be easily converted from allocentric rotation given 3D translation and camera intrinsics KK following .

Exemplary, approximate (ox,oy)(o_{x},o_{y}) as the bounding box center (cx,cy)(c_{x},c_{y}) and estimate tzt_{z} using a reference camera distance. PoseCNN directly regresses (ox,oy)(o_{x},o_{y}) and tzt_{z}. Nonetheless, this is not suitable for dealing with zoomed-in RoIs, since it is essential for the network to estimate position and scale invariant parameters.

Therefore, in our work we utilize a Scale-Invariant representation for Translation Estimation (SITE) . Concretely, given the size so=max⁡(w,h)s_{o}=\max(w,h) and center (cx,cy)(c_{x},c_{y}) of the detected bounding box and the ratio r=szoom/sor=s_{\text{zoom}}/s_{o} \wrtthe zoom-in size szooms_{\text{zoom}}, the network regresses the scale-invariant translation parameters tSITE=[δx,δy,δz]T\mathbf{t}_{\text{SITE}}=[\delta_{x},\delta_{y},\delta_{z}]^{T}, where

Finally, the 3D translation can be solved according to Eq. 3.

Disentangled 6D Pose Loss. Apart from the parameterization of rotation and translation, the choice of loss function is also crucial for 6D pose optimization. Instead of directly utilizing distances based on rotation and translation (\eg, angular distance, L1L_{1} or L2L_{2} distances), most works employ a variant of Point-Matching loss based on the ADD(-S) metric in an effort to couple the estimation of rotation and translation.

Inspired by , we employ a novel variant of disentangled 6D pose loss via individually supervising the rotation R\mathbf{R}, the scale-invariant 2D object center (δx,δy)(\delta_{x},\delta_{y}), and the distance δz\delta_{z}.

where ∙^\hat{\bullet} and ∙ˉ\bar{\bullet} denote prediction and ground truth, respectively. To account for symmetric objects, given Rˉ\mathcal{\bar{\mathcal{R}}}, the set of all possible ground-truth rotations under symmetry, we further extend our loss to a symmetry-aware formulation LR,sym=min⁡Rˉ∈RˉLR(R^,Rˉ)\mathcal{L}_{\mathbf{R},\text{sym}}=\underset{\mathbf{\bar{\mathbf{R}}}\in\mathcal{\bar{\mathcal{R}}}}{\min}\mathcal{L}_{\mathbf{R}}(\hat{\mathbf{R}},\bar{\mathbf{R}}).

2 Geometry-guided Direct Regression Network

In this section, we present our Geometry-guided Direct Regression Network, which we dub GDR-Net. Harnessing dense correspondence-based geometric features, we directly regress 6D object pose. Thereby, GDR-Net unifies approaches based on dense correspondences and direct regression.

Network Architecture. As shown in Fig. 2, we feed the GDR-Net with a zoomed-in RoI of size 256×256256\times 256 and predict three intermediate geometric feature maps with spatial size of 64×6464\times 64, which are composed of the Dense Correspondences Map (M2D-3DM_{\text{2D-3D}}), the Surface Region Attention Map (MSRAM_{\text{SRA}}) and the Visible Object Mask (MvisM_{\text{vis}}).

Our network is inspired by CDPN , a state-of-the-art dense correspondence-based method for indirect pose estimation. In essence, we keep the layers for regressing MXYZM_{\text{XYZ}} and MvisM_{\text{vis}}, while removing the disentangled translation head. Additionally, we append the channels required by MSRAM_{\text{SRA}} to the output layer. Since these intermediate geometric feature maps are all organized 2D-3D correspondences \wrtthe image, we employ a simple yet effective 2D convolutional Patch-PnnP module to directly regress the 6D object pose from M2D-3DM_{\text{2D-3D}} and MSRAM_{\text{SRA}}.

The Patch-PnnP module consists of three convolutional layers with kernel size 3×33\times 3 and stride=2\text{stride}=2, each followed by Group Normalization and ReLU activation. Two Fully Connected (FC) layers are then applied to the flattened feature, reducing the dimension from 8192 to 256. Finally, two parallel FC layers output the 3D rotation R\mathbf{R} parameterized as R6d\mathbf{R}_{\text{6d}} (Eq. 1) and 3D translation t\mathbf{t} parameterized as tSITE\mathbf{t}_{\text{SITE}} (Eq. 4), respectively.

Dense Correspondences Maps (M2D-3DM_{\text{2D-3D}}). In order to compute the Dense Correspondences Maps M2D-3DM_{\text{2D-3D}}, we first estimate the underlying Dense Coordinates Maps (MXYZM_{\text{XYZ}}). M2D-3DM_{\text{2D-3D}} can then be derived by stacking MXYZM_{\text{XYZ}} onto the corresponding 2D pixel coordinates. In particular, given the CAD model of an object, MXYZM_{\text{XYZ}} can be obtained by rendering the model’s 3D object coordinates given the associated pose. Similar to , we let the network predict a normalized representation of MXYZM_{\text{XYZ}}. Concretely, each channel of MXYZM_{\text{XYZ}} is normalized within $byby(l_{x},l_{y},l_{z})$, which is the size of corresponding tight 3D bounding box of the CAD model.

Notice that M2D-3DM_{\text{2D-3D}} does not only encode the 2D-3D correspondences, but also explicitly reflect the geometric shape information of objects. Moreover, as previously mentioned, since M2D-3DM_{\text{2D-3D}} is regular \wrtthe image, we are capable of learning the 6D object pose via a simple 2D convolutional neural network (Patch-PnnP).

Surface Region Attention Maps (MSRAM_{\text{SRA}}). Inspired by , we let the network predict the surface regions as additional ambiguity-aware supervision. However, instead of coupling them with RANSAC, we use them within our Patch-PnnP framework.

Essentially, the ground-truth regions MSRAM_{\text{SRA}} can be derived from MXYZM_{\text{XYZ}} employing farthest points sampling.

For each pixel we classify the corresponding regions, thus the probabilities in the predicted MSRAM_{\text{SRA}} implicitly represent the symmetry of an object. For instance, if a pixel is assigned to two potential fragments due to a plane of symmetry, Minimizing this assignment will return a probability of 0.5 for each fragment. Moreover, leveraging MSRAM_{\text{SRA}} not only mitigates the influence of ambiguities but also acts as an auxiliary task on top of M3DM_{\text{3D}}. In other words, it eases the learning of M3DM_{\text{3D}} by first locating coarse regions and then regressing finer coordinates. We utilize MSRAM_{\text{SRA}} as a symmetry-aware attention to guide the learning of Patch-PnnP.

Geometry-guided 6D Object Pose Regression. The presented image-based geometric feature patches, \ie, M2D-3DM_{\text{2D-3D}} and MSRAM_{\text{SRA}}, are then utilized to guide our proposed Patch-PnnP for direct 6D object pose regression as

We employ L1L_{1} loss for normalized MXYZM_{\text{XYZ}} and visible masks MvisM_{\text{vis}}, and cross-entropy loss (CECE) for MSRAM_{\text{SRA}}.

Thereby, ⊙\odot denotes element-wise multiplication and we only supervise MXYZM_{\text{XYZ}} and MSRAM_{\text{SRA}} using the visible region.

The overall loss for GDR-Net can be summarized as LGDR=LPose+LGeom.\mathcal{L}_{\text{GDR}}=\mathcal{L}_{\text{Pose}}+\mathcal{L}_{\text{Geom}}. Notice that our GDR-Net can be trained end-to-end, without requiring any three-stage training strategy as in .

Decoupling Detection and 6D Object Pose Estimation. Similar to , we mainly focus on the network for 6D object pose estimation and make use of an existing 2D object detector to obtain the zoomed-in input RoIs. This allows us to directly make use of the advances in runtime and accuracy within the rapidly growing field of 2D object detection, without having to change or re-train the pose network. Therefore, we adopt a simplified Dynamic Zoom-In (DZI) to decouple the training of our GDR-Net and object detectors. During training, we first uniformly shift the center and scale of the ground-truth bounding boxes by a ratio of 25%. We then zoom in the input RoIs with a ratio of r=1.5r=1.5 while maintaing the original aspect ratio. This ensures that the area containing the object is approximately half the RoI. DZI can also circumvent the need of dealing with varying object sizes.

Noteworthy, although we employ a two-stage approach, one could also implement GDR-Net on top of any object detector and train it in an end-to-end manner.

Experiments

In this section we first introduce our experimental setup, and then present the evaluation results for several commonly employed benchmark datasets. Thereby, we first present experiments on a synthetic toy dataset, which clearly demonstrates the benefit of our Patch-PnnP compared to the classic optimization-driven PnnP. Additionally, we demonstrate the effectiveness of our individual components by performing an ablative study on LM . Finally, we compare GDR-Net with state-of-the-art methods on two challenging datasets, \ieLM-O and YCB-V .

Implementation Details. All our experiments are implemented using PyTorch . We train all our networks end-to-end using the Ranger optimizer Ranger means the RAdam optimizer combined with Lookahead and Gradient Centralization . with a batch size of 24 and a base learning rate of 1e-4, which we anneal at 72% of the training phase using a cosine schedule .

Datasets. We conduct our experiments on four datasets: Synthetic Sphere , LM , LM-O , and YCB-V . The Synthetic Sphere dataset contains 20k samples for training and 2k for testing, created by randomly capturing a unit sphere model using a virtual calibrated camera with focal length 800, resolution 640×\times480, and the principal point located at the image center. The Rotations and translations are uniformly sampled in 3D space, and within an interval of ××\times\times, respectively. LM dataset consists of 13 sequences, each containing ≈\approx 1.2k images with ground-truth poses for a single object with clutter and mild occlusion. We follow and employ ≈\approx15% of the RGB images for training and 85% for testing. We additionally use 1k rendered RGB images for each object during training as in . LM-O consists of 1214 images from a LM sequence, where the ground-truth poses of 8 visible objects with more occlusion are provided for testing. Apart from the real data from LM, we also leverage synthetic data for training. Similarly, we render 10k synthetic images (syn) for each object as in . YCB-V is a very challenging dataset exhibiting strong occlusion, clutter and several symmetric objects. It comprises over 110k real images captured with 21 objects, both with and without texture. For both LM-O and YCB-V, we also leverage the publicly available synthetic data using physically-based rendering (pbr) for training.

Evaluation Metrics. We use two common metrics for 6D object pose evaluation, \ieADD(-S) , and n°,n cmn\degree,n~{}\text{cm} . The ADD metric measures whether the average deviation of the transformed model points is less than 10%10\% of the object’s diameter (0.1d). For symmetric objects, the ADD-S metric is employed to measure the error as the average distance to the closest model point . When evaluating on YCB-V, we also compute the AUC (area under curve) of ADD(-S) by varying the distance threshold with a maximum of 10 cm . The n°,n cmn\degree,n~{}\text{cm} metric measures whether the rotation error is less than n°n\degree and the translation error is below nn cm. Notice that to account for symmetries, n°,n cmn\degree,n~{}\text{cm} is computed \wrtthe smallest error for all possible ground-truth poses .

2 Toy Experiment on Synthetic Sphere

We conduct a toy experiment comparing our approach with PnnP/RANSAC and on the Synthetic Sphere dataset. We generate MXYZM_{\text{XYZ}} from the provided poses and feed them to our Patch-PnnP. For fairness, MSRAM_{\text{SRA}} is excluded from the input. Following , during training, we randomly add Gaussian noise N(0,σ2)\mathcal{N}(0,\sigma^{2}) with σ∈U[0,0.03]\sigma\in\mathcal{U}[0,0.03] to each point of the dense coordinates maps. Since the coordinates maps are normalized in $,wechoose0.03asitreflectsapproximatelythesamelevelofnoiseasin.Additionally,werandomlygenerated0, we choose 0.03 as it reflects approximately the same level of noise as in . Additionally, we randomly generated 0% to 30% of outliers forM_{\text{XYZ}}$ (Fig. 3(c)). During testing, we report the relative ADD error \wrtthe sphere’s diameter on the test set with different levels of noise and outliers.

Comparison with PnnP/RANSAC and . In Fig. 3, we demonstrate the effectiveness and robustness of our approach by comparing Patch-PnnP with the traditional RANSAC-based EPnnP and the learning-based PnnP from ). As depicted in Fig. 3, while RANSAC-based EPnnPWe follow the state-of-the-art method CDPN for the implementation and hyper-parameters of PnnP/RANSAC in all our experiments, see supplement for details. is more accurate when noise is unrealistically minimal, learning-based PnnP methods are much more accurate and robust as the level of noise increases. Moreover, Patch-PnnP is significantly more robust than Single-Stage \wrtto noise and outliers, thanks to our geometrically rich and dense correspondences maps.

3 Ablation Study on LM

We present several ablation experiments for the widely used LM dataset . We train a single GDR-Net for all objects for 160 epochs without applying any color augmentation. For fairness in evaluation, we leverage the detection results from Faster-RCNN as provided by .

Number of Regions in MSRAM_{\text{SRA}}. In Tab. 4(a), we show results for different numbers of regions in MSRAM_{\text{SRA}}. Thereby, without our attention MSRAM_{\text{SRA}} (number of regions = 0), the accuracy is deliberately good, which suggests the effectiveness and versatility of Patch-PnnP. Nevertheless, the overall accuracy can be further improved with increasing number of regions in MSRAM_{\text{SRA}}, despite starting to saturate around 64 regions. Thus, we use 64 regions for MSRAM_{\text{SRA}} in all other experiments as trade-off between accuracy and memory.

Effectiveness of Patch-PnnP. We demonstrate the effectiveness of the image-like geometric features (M2D-3D,MSRAM_{\text{2D-3D}},M_{\text{SRA}}) by comparing our Patch-PnnP with traditional PnnP/RANSAC , the PointNet-like PnnP from , and a differentiable PnnP (BPnnP ). For PointNet-like PnnP, we extend the PointNet in to account for dense correspondences. Specifically, we utilize PointNet to pointwisely transform the spatially flattened geometric features (M2D-3DM_{\text{2D-3D}} and MSRAM_{\text{SRA}}) and directly predict the 6D pose with global max pooling followed by two FC layers. Since the correspondences are explicitly encoded in M2D-3DM_{\text{2D-3D}}, no special attention is needed for the keypoint orders as in . For BPnnP , we replace the Patch-PnnP in our framework with their implementation of BPnnPhttps://github.com/BoChenYS/BPnP. As BPnnP was originally designed for sparse keypoints, we further adapt it appropriately to deal with dense coordinates.

As shown in Tab. 4(b), Patch-PnnP is more accurate than traditional PnnP/RANSAC (B0 \vsA0), PointNet-like PnnP (B0 \vsC0) and BPnnP (B0 \vsC1) in estimating the 6D pose. Furthermore, in terms of rotation, our Patch-PnnP outperforms PointNet-like PnnP by a large margin, which proves the importance of exploiting the ordering within the correspondences. Noteworthy, Patch-PnnP is much faster in inference and up to 4×\times faster in training than BPnnP, since the latter relies on PnnP/RANSAC for both phases.

Parameterization of 6D Pose. In Tab. 4(b), we illustrate the impact of our proposed 6D pose parameterization. In particular, the 6-dimensional R6d\mathbf{R}_{\text{6d}} (Eq. 1) achieves a much more accurate estimate of R\mathbf{R} than commonly used representations such as unit quaternions , log quaternions and the Lie algebra-based vectors (c.f. B0 \vsD1-D3, and G0 \vsG2). Moreover, we can deduce that the allocentric representation is significantly stronger than the egocentric formulation (B0 \vsD0).

Similarly, the parameterization of the 3D translation is of high importance. Essentially, directly predicting t\mathbf{t} in 3D space leads to worse results than leveraging the scale-invariant formulation tSITE\mathbf{t}_{\text{SITE}} (E0 \vsB0). Additionally, replacing the scale-invariant δz\delta_{z} in tSITE\mathbf{t}_{\text{SITE}} with the absolute distance tzt_{z} or directly regressing the object center (ox,oz)(o_{x},o_{z}) leads to inferior poses \wrttranslation (B0 \vsE1, E2). Hence, when dealing with zoomed-in RoIs, it is essential to parameterize the 3D translation in a scale-invariant fashion.

Ablation on Pose Loss. As mentioned in Section 3.1, the loss function has an impact on direct 6D pose regression. In Tab. 4(b), we compare our disentangled LPose\mathcal{L}_{\text{Pose}} to a simple angular loss and the Point-Matching loss (F0). Furthermore, we present its disentangled versions following . As shown in (B0 and F0-F4), all variants of the PM loss are clearly better than the angular loss in terms of rotation estimation. In addition, disentangling the rotation R\mathbf{R} and distance tzt_{z} in LPM\mathcal{L}_{\text{PM}} largely enhances the rotation accuracy. Nonetheless, the overall performance is slightly inferior to our disentangled formulation LPose\mathcal{L}_{\text{Pose}}, which disentangles tSITE\mathbf{t}_{\text{SITE}} rather than the 3D translation t\mathbf{t}. It is worth noting that LR,sym\mathcal{L}_{\mathbf{R},\text{sym}} has a rather insignificant contribution compared with LR\mathcal{L}_{\mathbf{R}}. This can be accounted to the lack of severe symmetries in LM and to our proposed surface region attention MSRAM_{\text{SRA}}.

Effectiveness of Geometry-Guided Direct Regression. Furthermore, we train GDR-Net leveraging only our pose loss LPose\mathcal{L}_{\text{Pose}} by discarding the geometric supervision LGeom\mathcal{L}_{\text{Geom}}. Surprisingly, even the simple version outperforms CDPN \wrtADD(-S) 0.1d, when employing R6d\mathbf{R}_{\text{6d}} for rotation (Tab. 4(b) G0 \vsA0). Yet, we clearly outperform our baseline using GDR-Net with explicit geometric guidance. If we predict the rotation as allocentric quaternions, the accuracy decreases (G2 \vsG0), which can partially account for the weak performance of previous direct methods . Moreover, when we remove the guidance of M2DM_{\text{2D}}, the accuracy drops significantly (G0 \vsG1). Based on these results, we can see that an appropriate geometric guidance is essential for direct 6D pose regression.

Direct pose regression also enhances the learning of geometric features as the error signal from pose can be backpropagated. Tab. 4(b) (B1, B2) shows that when evaluating GDR-Net with PnnP/RANSAC from the predicted M2D-3DM_{\text{2D-3D}}, the overall performance exceeds CDPN . Similar to CDPN, we run tests using PnnP/RANSAC for R\mathbf{R} and Patch-PnnP for t\mathbf{t}, which achieves the overall best accuracy (B2). This demonstrates that our unified GDR-Net can leverage the best of both worlds, namely, geometry-based indirect methods and direct methods.

Effectiveness of Detection and Pose Decoupling. Similar to CDPN , we decouple the detector and GDR-Net by means of Dynamic Zoom-In (DZI). When evaluating GDR-Net with the Yolov3 detections from , the overall accuracy only drops slightly while the accuracy for ADD(-S) 0.1d almost remains unchanged (Tab. 4(b) H0).

4 Comparison with State of the Art

We compare our approach with state-of-the-art methods on the LM-O and YCB-V datasets.We follow the most commonly used evaluation protocol for LM-O and YCB-V, which has also been employed by another learned PnnP and many other works such as . We kindly refer the readers to our supplement for the results under BOP setup. During training, we apply similar color augmentation as in to prevent overfitting. For YCB-V, due to the large number of symmetric objects, the symmetric variant for the pose loss LR,sym\mathcal{L}_{\mathbf{R},\text{sym}} is employed. During testing, for LM-O, we employ Faster-RCNN to obtain 2D detections from the RGB images; for YCB-V, we utilize the publicly available detections from FCOS https://github.com/LZGMatrix/BOP19_CDPN_2019ICCV.

Results on LM-O. Tab. 2 presents the results of GDR-Net compared with state-of-the-art methods on LM-O. When trained with “real+syn”, our single GDR-Net is comparable to . Nevertheless, using one network per object, we easily surpass state of the art without refinement. Moreover, our GDR-Net trained with “real+pbr” even outperforms the refinement-based method DeepIM .

Results on YCB-V. We compare GDR-Net to state-of-the-art approaches on YCB-V in Tab. 3 (see supplement for detailed results). Our GDR-Net trained one network per object exceeds again state of the art, even without leveraging any refinement. Our single model for all objects is also comparable to the refinement-based methods such as \wrtAUC of ADD-S metric. Noteworthy, our approach runs much faster than the methods requiring refinement.

5 Runtime Analysis

On a desktop with an Intel 3.40GHz CPU and an NVIDIA 2080Ti GPU, given a 640×480640\times 480 image, using the Yolov3 detector, our approach takes ≈\approx 22ms for a single object and ≈\approx 35ms for 8 objects, including 15ms for detection.

Conclusion

In this work, we revisited the ingredients of direct 6D pose regression and proposed a novel GDR-Net to unify direct and geometry-based indirect methods. The key idea is to exploit the intermediate geometric features regarding 2D-3D correspondences organized regularly as image-like 2D patches, which facilitates us to utilize a simple yet effective 2D convolutional Patch-PnnP to directly regress 6D pose from geometric guidance. Our approach achieves real-time, accurate and robust monocular 6D object pose estimation. In the future, we want to extend our work to more challenging scenarios, such as the lack of annotated real data and unseen object categories or instances .

Acknowledgement We thank Zhigang Li, Xingyu Liu at Tsinghua University for their helpful discussion, and Nikolas Brasch at Technical University of Munich for proofreading. We also thank anonymous reviewers for their constructive comments. This work was supported in part by China Scholarship Council (CSC) Grant #201906210393, and in part by the National Key R&D Program of China under Grant 2018AAA0102801.

References

Appendix A Details of Pn𝑛nP/RANSAC

The implementation and hyper-parameters of PnnP/RANSAC follow the state-of-the-art method CDPN for all our experiments. Specifically, we leverage EPnnP together with 100 RANSAC iterations using a reprojection error threshold of 3 and confidence threshold of 0.99.

Appendix B Detailed Results of YCB-V

We present detailed evaluation results on YCB-V for our GDR-Net in Tab. B.1 and Tab. B.2 and compare them to state-of-the-art approaches \wrtADD(-S) and AUC of ADD-S/ADD(-S), respectively. As for methods trained simultaneously for all objects, our GDR-Net clearly outperforms all other state-of-the-art methods. Furthermore, when GDR-Net is trained separately for each individual object, we can even surpass refinement-based methods such as DeepIM \wrtAUC of ADD-S/ADD(-S) metric.

Appendix C BOP Results on LM-O and YCB-V

In the main paper, we have presented the results on LM-O and YCB-V following the most commonly used evaluation protocol following another learned PnnP and many other works such as . Nevertheless, the evaluation protocol of BOP Challenge has recently become more popular. Therefore, we also present the results of our GDR-Net on LM-O and YCB-V under the BOP setup.

The BOP evaluation protocol differs from the former in three main aspects as follows. i) No real data should be used for LM-O, thus we only employ the provided synthetic pbr data for training on LM-O; ii) The number of test images for both LM-O and YCB-V is smaller, \ie, they only contains a subset of the original test images; iii) The evaluation metric is different. Thereby, for each dataset, an Average Recall (AR) score is reported by calculating the mean Average Recall of three different metrics: AR=(ARMSPD+ARMSSD+ARVSD)/3\text{AR}=(\text{AR}_{\text{{\tiny MSPD}}}+\text{AR}_{\text{{\tiny MSSD}}}+\text{AR}_{\text{{\tiny VSD}}})/3. Please refer to for the detailed explanation of these metrics.

Tab. C.3 presents the results of our GDR-Net on LM-O and YCB-V compared with other state-of-the-art RGB-based methods under BOP setup. Since our method is built on top of CDPN , we follow to train one network per object for the sake of fairness. We utilize the publicly available detections from FCOS https://github.com/LZGMatrix/BOP19_CDPN_2019ICCV following CDPNv2 . We can see that our GDR-Net significantly outperforms all other state-of-the-art methods without refinement. It is worth noting that most of these top-performing methods rely on the indirect PnnP/RANSAC solver, while ours directly regresses the 6D object pose leveraging geometric guidance, which again demonstrates the effectiveness of our proposed learning-based Patch-PnnP. Our GDR-Net even outperforms the state-of-the-art refinement-based method CosyPose on LM-O. On YCB-V, ours is worse than CosyPose but far better than all other methods without refinement. Nevertheless, our method runs much faster than CosyPose as no refinement step is needed. Moreover, our method can be combined with an additional refiner such as CosyPose to achieve better results.

Appendix D Qualitative Results

We demonstrated additional qualitative results for LM , LM-O , and YCB-V in Fig. 4, Fig. D.1 and Fig. D.2, respectively. Thereby, in Fig. 4, we visualize the 6D pose by overlaying the image with the corresponding transformed 3D bounding box. In Fig. D.1 and Fig. D.2, we illustrate the estimated 6D poses by rendering the 3D models on top of the input image and highlighting the respective contours. Note that while Blue constitutes the ground-truth poses, we demonstrate in Green the predicted poses from GDR-Net. For better visualization we cropped the images and zoomed into the area of interest.