DualPoseNet: Category-level 6D Object Pose and Size Estimation Using Dual Pose Network with Refined Learning of Pose Consistency

Jiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu, Kui Jia, Yuanqing Li

Introduction

Object detection in a 3D Euclidean space is demanded in many practical applications, such as augmented reality, robotic manipulation, and self-driving car. The field has been developing rapidly with the availability of benchmark datasets (e.g., KITTI and SUN RGB-D ), where carefully annotated 3D bounding boxes enclosing object instances of interest are prepared, which specify 7 degrees of freedom (7DoF) for the objects, including translation, size, and yaw angle around the gravity axis. This 7DoF setting of 3D object detection is aligned with common scenarios where the majority of object instances stand upright in the 3D space. However, 7DoF detection cannot precisely locate objects when the objects lean in the 3D space, where the most compact bounding boxes can only be determined given full pose configurations, i.e., with the additional two angles of rotation. Pose predictions of full configurations are important in safety-critical scenarios, e.g., autonomous driving, where the most precise and compact localization of objects enables better perception and decision making.

This task of pose prediction of full configurations (i.e., 6D pose and size) is formally introduced in as category-level 6D object pose and size estimation of novel instances from single, arbitrary views of RGB-D observations. It is closely related to category-level amodal 3D object detection (i.e., the above 7DoF setting) and instance-level 6D object pose estimation . Compared with them, the focused task in the present paper is more challenging due to learning and prediction in the full rotation space of SO(3)SO(3); more specifically, (1) the task is more involved in terms of both defining the category-level canonical poses (cf. Section 3 for definition of canonical poses) and aligning object instances with large intra-category shape variations , (2) deep learning precise rotations arguably requires learning rotation-equivariant shape features, which is less studied compared with the 2D counterpart of learning translation-invariant image features, and (3) compared with instance-level 6D pose estimation, due to the lack of testing CAD models, the focused task cannot leverage the privileged 3D shapes to directly refine pose predictions, as done in .

In this work, we propose a novel method for category-level 6D object pose and size estimation, which can partially address the second and third challenges mentioned above. Our method constructs two parallel pose decoders on top of a shared pose encoder; the two decoders predict poses with different working mechanisms, and the encoder is designed to learn pose-sensitive shape features. A refined learning that enforces the predicted pose consistency between the two decoders is activated during testing to further improve the prediction. We term our method as Dual Pose Network with refined learning of pose consistency, shortened as DualPoseNet. Fig. 1 gives an illustration.

For an observed RGB-D scene, DualPoseNet first employs an off-the-shelf model of instance segmentation (e.g., MaskRCNN ) in images to segment out the objects of interest. It then feeds each masked RGB-D region into the encoder. To learn pose-sensitive shape features, we construct our encoder based on spherical convolutions , which provably learn deep features of object surface shapes with the property of rotation equivariance on SO(3)SO(3). In this work, we design a novel module of Spherical Fusion to support a better embedding from the appearance and shape features of the input RGB-D region. With the learned pose-sensitive features, the two parallel decoders either make a pose prediction explicitly, or implicitly do so by reconstructing the input (partial) point cloud in its canonical pose; while the first pose prediction can be directly used as the result of DualPoseNet, the result is further refined during testing by fine-tuning the encoder using a self-adaptive loss term that enforces the pose consistency. The use of implicit decoder in DualPoseNet has two benefits that potentially improve the pose prediction: (1) it provides an auxiliary supervision on the training of pose encoder, and (2) it is the key to enable the refinement given no testing CAD models. We conduct thorough experiments on the benchmark category-level object pose datasets of CAMERA25 and REAL275 , and also apply our DualPoseNet to the instance-level ones of YCB-Video and LineMOD . Ablation studies confirm the efficacy of our novel designs. DualPoseNet outperforms existing methods in terms of more precise pose. Our technical contributions are summarized as follows:

We propose a new method of Dual Pose Network for category-level 6D object pose and size estimation. DualPoseNet stacks two parallel pose decoders on top of a shared pose encoder, where the implicit one predicts poses with a working mechanism different from that of the explicit one; the two decoders thus impose complementary supervision on training of the pose encoder.

In spite of the lack of testing CAD models, the use of implicit decoder in DualPoseNet enables a refined pose prediction during testing, by enforcing the predicted pose consistency between the two decoders using a self-adaptive loss term. This further improves the results of DualPoseNet.

We construct the encoder of DualPoseNet based on spherical convolutions to learn pose-sensitive shape features, and design a module of Spherical Fusion wherein, which is empirically shown to learn a better embedding from the appearance and shape features from the input RGB-D regions.

Related Work

Instance-level 6D Object Pose Estimation Traditional methods for instance-level 6D pose estimation include those based on template matching , and those by voting the matching results of point-pair features . More recent solutions build on the power of deep networks and can directly estimate object poses from RGB images alone or RGB-D ones . This task assumes the availability of object CAD models during both the training and test phases, and thus enables a common practice to refine the predicted pose by matching the CAD model with the (RGB and/or point cloud) observations .

Category-level 3D Object Detection Methods for category-level 3D object detection are mainly compared on benchmarks such as KITTI and SUN RGB-D . Earlier approach leverages the mature 2D detectors to first detect objects in RGB images, and learning of 3D detection is facilitated by focusing on point sets inside object frustums. Subsequent research proposes solutions to predict the 7DoF object bounding boxes directly from the observed scene points. However, the 7DoF configurations impose inherent constrains on the precise rotation prediction, with only one yaw angle predicted around the gravity direction.

Category-level 6D Object Pose and Size Estimation More recently, category-level 6D pose and size estimation is formally introduced in . Notably, Wang et al. propose a canonical shape representation called normalized object coordinate space (NOCS), and inference is made by first predicting NOCS maps for objects detected in RGB images, and then aligning them with the observed object depths to produce results of 6D pose and size; later, Tian et al. improve the predictions of canonical object models by deforming categorical shape priors. Instead, Chen et al. trained a variational auto-encoder (VAE) to capture pose-independent features, along with pose-dependent ones to directly predict the 6D poses. Besides, monocular methods are also explored in recent works .

Problem Statement

The Proposed Dual Pose Network with Refined Learning of Pose Consistency

2 The Pose Encoder ΦΦ\Phi

Precise prediction of object pose requires that the features learned by f=Φ(X,P)\bm{f}=\Phi(\mathcal{X},\mathcal{P}) are sensitive to the observed pose of the input P\mathcal{P}, especially to rotation, since translation and size are easier to infer from P\mathcal{P} (e.g., even simple localization of center point and calculation of 3D extensions give a good prediction of translation and scale). To this end, we implement our Φ\Phi based on spherical convolutions , which provably learn deep features of object surface shapes with the property of rotation equivariance on SO(3)SO(3). More specifically, we design Φ\Phi to have two parallel streams of spherical convolution layers that process the inputs X\mathcal{X} and P\mathcal{P} separately; the resulting features are intertwined in the intermediate layers via a proposed module of Spherical Fusion. We also use aggregation of multi-scale spherical features to enrich the pose information in f\bm{f}. Fig. 1 gives the illustration.

Aggregation of Multi-scale Spherical Features It is intuitive to enhance pose encoding by using spherical features at multiple scales. Since the representation S~lX,P\widetilde{\mathcal{S}}_{l}^{\mathcal{X},\mathcal{P}} computed by (2) fuses the appearance and geometry features at an intermediate layer ll, we technically aggregate multiple of them from spherical fusion modules respectively inserted at lower, middle, and higher layers of the two parallel streams, as illustrated in Fig. 1. In practice, we aggregate three of such feature representations as follows

where Flatten(⋅)\texttt{Flatten}(\cdot) denotes a flattening operation that reforms the feature tensor S~lX,P\widetilde{\mathcal{S}}^{\mathcal{X},\mathcal{P}}_{l} of dimension W×H×dlW\times H\times d_{l} as a feature vector, MLP denotes a subnetwork of Multi-Layer Perceptron (MLP), and MaxPool(fl,fl′,fl′′)\texttt{MaxPool}\left(\bm{f}_{l},\bm{f}_{l^{\prime}},\bm{f}_{l^{\prime\prime}}\right) aggregates the three feature vectors by max-pooling over three entries for each feature channel; layer specifics of the two MLPs used in (3) are given in Fig. 1. We use f\bm{f} computed from (3) as the final output of the pose encoder Φ\Phi, i.e., f=Φ(X,P)\bm{f}=\Phi(\mathcal{X},\mathcal{P}).

Given f\bm{f} from the encoder Φ\Phi, we implement the explicit decoder Ψexp\Psi_{exp} simply as three parallel MLPs that are trained to directly regress the rotation R\bm{R}, translation t\bm{t}, and size s\bm{s}. Fig. 1 gives the illustration, where layer specifics of the three MLPs are also given. This gives a direct way of pose prediction from a cropped RGB-D region as (R,t,s)=Ψexp∘Φ(X,P)(\bm{R},\bm{t},\bm{s})=\Psi_{exp}\circ\Phi(\mathcal{X},\mathcal{P}).

For the observed point cloud P\mathcal{P}, assume that its counterpart Q\mathcal{Q} in the canonical pose is available. An affine transformation (R,t,s)(\bm{R},\bm{t},\bm{s}) between P\mathcal{P} and Q\mathcal{Q} can be established, which computes q=1∣∣s∣∣RT(p−t)\bm{q}=\frac{1}{||\bm{s}||}\bm{R}^{T}(\bm{p}-\bm{t}) for any corresponding pair of p∈P\bm{p}\in\mathcal{P} and q∈Q\bm{q}\in\mathcal{Q}. This implies an implicit way of obtaining predicted pose by learning to predict a canonical Q\mathcal{Q} from the observed P\mathcal{P}; upon prediction of Q\mathcal{Q}, the pose (R,t,s)(\bm{R},\bm{t},\bm{s}) can be obtained by solving the alignment problem via Umeyama algorithm . Since f=Φ(X,P)\bm{f}=\Phi(\mathcal{X},\mathcal{P}) has learned the pose-sensitive features, we expect the corresponding q\bm{q} can be estimated from p\bm{p} by learning a mapping from the concatenation of f\bm{f} and p\bm{p}. In DualPoseNet, we simply implement the learnable mapping as

Ψim\Psi_{im} applies to individual points of P\mathcal{P} in a point-wise manner. We write collectively as Q=Ψim(P,f)\mathcal{Q}=\Psi_{im}(\mathcal{P},\bm{f}).

We note that an equivalent representation of normalized object coordinate space (NOCS) is learned in for a subsequent computation of pose prediction. Different from NOCS, we use Ψim\Psi_{im} in an implicit way; it has two benefits that potentially improve the pose prediction: (1) it provides an auxiliary supervision on the training of pose encoder Ψ\Psi (note that the training ground truth of Q\mathcal{Q} can be transformed from P\mathcal{P} using the annotated pose and size), and (2) it enables a refined pose prediction by enforcing the consistency between the outputs of Ψexp\Psi_{exp} and Ψim\Psi_{im}, as explained shortly in Section 4.6. We empirically verify both the benefits in Section 5.1.1, and show that the use of Ψim\Psi_{im} improves pose predictions of Ψexp∘Φ(X,P)\Psi_{exp}\circ\Phi(\mathcal{X},\mathcal{P}) in DualPoseNet.

5 Training of Dual Pose Network

Given the ground-truth pose annotation (R∗,t∗,s∗)(\bm{R}^{*},\bm{t}^{*},\bm{s}^{*}) Following , we use canonical R∗\bm{R}^{*} for symmetic objects to handle ambiguities of symmetry. for a cropped (X,P)(\mathcal{X},\mathcal{P}), we use the following training objective on top of the explicit decoder Ψexp\Psi_{exp}:

where ρ(R)\rho(\bm{R}) is the quaternion representation of rotation R\bm{R}.

Since individual points in the predicted Q={qi}i=1N\mathcal{Q}=\{\bm{q}_{i}\}_{i=1}^{N} from Ψim\Psi_{im} are respectively corresponded to those in the observed P={pi}i=1N\mathcal{P}=\{\bm{p}_{i}\}_{i=1}^{N}, we simply use the following loss on top of the implicit decoder

The overall training objective combines (5) and (6), resulting in the optimization problem

6 The Refined Learning of Pose Consistency

For instance-level 6D pose estimation, it is a common practice to refine an initial or predicted pose by a post-registration or post-optimization ; such a practice is possible since CAD model of the instance is available, which can guide the refinement by matching the CAD model with the (RGB and/or point cloud) observations. For our focused category-level problem, however, CAD models of testing instances are not provided. This creates a challenge in case that more precise predictions on certain testing instances are demanded.

Thanks to the dual pose predictions from Ψexp\Psi_{exp} and Ψim\Psi_{im}, we are able to make a pose refinement by learning to enforce their pose consistency. More specifically, we freeze the parameters of Ψexp\Psi_{exp} and Ψim\Psi_{im}, while fine-tuning those of the encoder Φ\Phi, by optimizing the following problem

where Q={qi}i=1N=Ψim∘Φ(X,P)\mathcal{Q}=\{\bm{q}_{i}\}_{i=1}^{N}=\Psi_{im}\circ\Phi(\mathcal{X},\mathcal{P}) and (R,t,s)=Ψexp∘Φ(X,P)(\bm{R},\bm{t},\bm{s})=\Psi_{exp}\circ\Phi(\mathcal{X},\mathcal{P}) are the outputs of the two decoders. Note that during training, the two decoders are consistent in terms of pose prediction, since both of them are trained to match their outputs with the ground truths. During testing, due to an inevitable generalization gap, inconsistency between outputs of the two decoders always exists, and our proposed refinement (8) is expected to close the gap. An improved prediction relies on a better pose-sensitive encoding f=Φ(X,P)\bm{f}=\Phi(\mathcal{X},\mathcal{P}); the refinement (8) thus updates parameters of Φ\Phi to achieve the goal. Empirical results in Section 5.1.1 verify that the refined poses are indeed towards more precise ones. In practice, we set a loss tolerance ϵ\epsilon as the stopping criterion when fine-tuning LΦRefine\mathcal{L}_{\Phi}^{Refine} (i.e., the refinement stops when LΦRefine≤ϵ\mathcal{L}_{\Phi}^{Refine}\leq\epsilon), with fast convergence and negligible cost.

Experiments

Datasets We conduct experiments using the benchmark CAMERA25 and REAL275 datasets for category-level 6D object pose and size estimation. CAMERA25 is a synthetic dataset generated by a context-aware mixed reality approach from 66 object categories; it includes 300,000300,000 composite images of 1,0851,085 object instances, among which 25,00025,000 images of 184184 instances are used for evaluation. REAL275 is a more challenging real-world dataset captured with clutter, occlusion and various lighting conditions; its training set contains 4,3004,300 images of 77 scenes, and the test set contains 2,7502,750 images of 66 scenes. Note that CAMERA25 and REAL275 share the same object categories, which enables a combined use of the two datasets for model training, as done in .

We also evaluate the advantages of DualPoseNet on the benchmark instance-level object pose datasets of YCB-Video and LineMOD , which consist of 2121 and 1313 different object instances respectively.

Implementation Details We employ a MaskRCNN implemented by to segment out the objects of interest from input scenes. For each segmented object, its RGB-D crop is converted as spherical signals with a sampling resolution 64×6464\times 64, and is then fed into our DualPoseNet. Configurations of DualPoseNet, including channel numbers of spherical convolutions and MLPs, have been specified in Fig. 1. We use ADAM to train DualPoseNet, with an initial learning rate of 0.00010.0001. The learning rate is halved every 50,00050,000 iterations until a total number of 300,000300,000 ones. We set the batch size as 6464, and the penalty parameter in Eq. (7) as λ=10\lambda=10. For refined learning of pose consistency, we use a learning rate 1×10−61\times 10^{-6} and a loss tolerance ϵ=5×10−5\epsilon=5\times 10^{-5}. For the instance-level task, we additionally adopt a similar 2nd-stage iterative refinement of residual pose as did; more details are shown in the supplementary material.

Evaluation Metrics For category-level pose estimation, we follow to report mean Average Precision (mAP) at different thresholds of intersection over union (IoU) for object detection, and mAP at n°n\degree mm cm for pose estimation. However, those metrics are not precise enough to simultaneously evaluate 6D pose and object size estimation, since IoU alone may fail to characterize precise object poses (a rotated bounding box may give a similar IoU value). To evaluate the problem nature of simultaneous predictions of pose and size, in this work, we also propose a new and more strict metric based on a combination of IoU, error of rotation, and error of relative translation, where for the last one, we use the relative version since absolute translations make less sense for objects of varying sizes. For the three errors, we consider respective thresholds of {50%,75%}\{50\%,75\%\} (i.e., IoU50 and IoU75), {5°,10°}\{5\degree,10\degree\}, and {5%,10%,20%}\{5\%,10\%,20\%\}, whose combinations can evaluate the predictions across a range of precisions. For instance-level pose estimation, we follow and evaluate the results of YCB-Video and LineMOD datasets by ADD-S and ADD(S) metrics, respectively.

We first conduct ablation studies to evaluate the efficacy of individual components proposed in DualPoseNet. These studies are conducted on the REAL275 dataset .

We use both Ψexp\Psi_{exp} and Ψim\Psi_{im} for pose decoding from DualPoseNet; Ψexp\Psi_{exp} produces the pose predictions directly, which are also used as the results of DualPoseNet both with and without the refined learning, while Ψim\Psi_{im} is an implicit one whose outputs can translate as the results by solving an alignment problem. To verify the usefulness of Ψim\Psi_{im}, we report the results of DualPoseNet with or without the use of Ψim\Psi_{im} in Table 1, in terms of the pose precision from Ψexp\Psi_{exp} before the refined learning. We observe that the use of Ψim\Psi_{im} improves the performance of Ψexp\Psi_{exp} by large margins under all the metrics; for example, the mAP improvement of (IoU50,10°,10%{}_{50},10\degree,10\%) reaches 5.8%5.8\%, and that of (IoU75,5°,10%{}_{75},5\degree,10\%) reaches 4.1%4.1\%. These performance gains suggest that Ψim\Psi_{im} not only enables the subsequent refined learning of pose consistency, but also provides an auxiliary supervision on the training of pose encoder Φ\Phi and results in a better pose-sensitive embedding, implying the key role of Ψim\Psi_{im} in DualPoseNet.

To evaluate the efficacy of our proposed spherical fusion based encoder Φ\Phi, we compare with three alternative encoders: (1) a baseline of Densefusion , a pose encoder that fuses the learned RGB features from CNNs and point features from PointNet in a point-wise manner; (2) SCNN-EarlyFusion, which takes as input the concatenation of SX\mathcal{S}^{\mathcal{X}} and SP\mathcal{S}^{\mathcal{P}} and feeds it into a multi-scale spherical CNN, followed by an MLP; (3) SCNN-LateFusion, which first feeds SX\mathcal{S}^{\mathcal{X}} and SP\mathcal{S}^{\mathcal{P}} into two separate multi-scale spherical CNNs and applies an MLP to the concatenation of the two output features. The used multi-scale spherical CNN is constructed by 88 spherical convolution layers, with aggregation of multi-scale spherical features similar to Φ\Phi. We conduct ablation experiments by replacing Φ\Phi with the above encoders, while keeping Ψexp\Psi_{exp} and Ψim\Psi_{im} as remained. Results (without the refined learning of pose consistency) in Table 1 show that the three alternative encoders perform worse than our proposed Φ\Phi with spherical fusion. Compared with the densefusion baseline, those based on spherical convolutions enjoy the property of rotation equivariance on SO(3)SO(3), and thus achieve higher mAPs. With spherical fusion, our proposed pose encoder Φ\Phi enables information communication progressively along the hierarchy, outperforming either SCNN-EarlyFusion with feature fusion at the very beginning or SCNN-LateFusion at the end.

We finally investigate the benefit from the proposed refined learning of pose consistency. Results in Table 1 show that with the refined learning, pose precisions improve stably across the full spectrum of evaluation metrics, and the improvements increase when coarser metrics are used; this suggests that the refinement indeed attracts the learning towards more accurate regions in the solution space of pose prediction. Examples in Fig. 3 give corroborative evidence of the efficacy of the refined learning. In fact, the refinement process is a trade-off between the improved precision and refining efficiency. As mentioned in Section 4.6, the refining efficiency depends on the learning rate and number of iterations for fine-tuning the encoder with the objective (8). In Fig. 4, we plot curves of mAP of (IoU50, 10°10\degree, 20%20\%) against the number of iterations when using different learning rates. It shows faster convergence when using larger learning rates, which, however, may end with less mature final results, even with overfitting. In practice, one may set a proper tolerance ϵ\epsilon for a balanced efficiency and accuracy. We set the learning rate as 1×10−61\times 10^{-6} and ϵ=5×10−5\epsilon=5\times 10^{-5} for results of DualPoseNet reported in the present paper. Under this setting, it costs negligible 0.20.2 seconds per instance on a server with Intel E5-2683 CPU and GTX 1080ti GPU.

1.2 Comparisons with Existing Methods

We compare our proposed DualPoseNet with the existing methods, including NOCS , SPD , and CASS , on the CAMERA25 and REAL275 datasets. Note that NOCS and SPD are designed to first predict canonical versions of the observed point clouds, and the poses are obtained from post-alignment by solving a Umeyama algorithm . Quantitative results in Table 2 show the superiority of our proposed DualPoseNet on both datasets, especially for the metrics of high precisions. For completeness, we also present in Table 2 the comparative results under the original evaluation metrics proposed in ; our results are better than existing ones at all but one rather coarse metric of IoU50, which is in fact a metric less sensitive to object pose. Qualitative results of different methods are shown in Fig. 5. Comparative advantages of our method over the existing ones are consistent with those observed in Table 2. For example, the bounding boxes of laptops in the figure generated by NOCS and SPD are obviously bigger than the exact extensions of laptops, while our method predicts more compact bounding boxes with precise poses and sizes. More comparative results are shown in the supplementary material.

2 Instance-level 6D Pose Estimation

We apply DualPoseNet to YCB-Video and LineMOD datasets for the instance-level task. Results in Table 3 confirm the efficacy of our individual components (the encoder Φ\Phi, the implicit decoder Ψim{\Psi}_{im}, and the refined learning of pose consistency); following , we also augment our DualPoseNet with a 2nd-stage module for iterative refinement of residual pose, denoted as DualPoseNet(Iterative), to further improve the performance. As shown in Table 4, DualPoseNet(Iterative) achieves comparable results against other methods, showing its potential for use in instance-level tasks. More quantitative and qualitative results are shown in the supplementary material.

Acknowledgement

This work was partially supported by the Guangdong R&\&D key project of China (No.: 2019B010155001), the National Natural Science Foundation of China (No.: 61771201), and the Program for Guangdong Introducing Innovative and Entrepreneurial Teams (No.: 2017ZT07X183).

References