Learning Canonical Shape Space for Category-Level 6D Object Pose and Size Estimation

Dengsheng Chen, Jun Li, Zheng Wang, Kai Xu

Introduction

6D object pose estimation based on a single-view RGB(D) image is an essential building block for several real-world applications ranging from robotic navigation and manipulation to augmented reality. Most existing works have so far been addressing instance-level 6D pose estimation where each target object has a corresponding CAD model with exact shape and size . Thereby, the problem is largely reduced to finding sparse or dense correspondence between the target object and the stock 3D model. Pose hypotheses can then be generated and verified based on the correspondences. Although enjoying high pose accuracy, the requirement of exact CAD models by these techniques hinders their practical use in many application scenarios.

Recently, category-level 6D object pose estimation starts to gain attention . In this problem, the target object of a shape category is unseen before and no CAD model is available, although some other instances of the same category may have been seen. Therefore, the major challenge is how to deal with intra-class variation . In general, household objects could exhibit significant variations in color, texture, shape and size even within the same category. Without an exactly same CAD model, correspondence-based approach would find difficulty given the considerable intra-class shape variation.

To resolve this, a unified representation for a variety of instances of an object category is needed, to which the target object in observation could be “matched”. The recently proposed Normalized Object Coordinate Space (NOCS) is a nice example of such unified representation. Based on this representation, category-level object poses can be estimated with high accuracy through mapping each pixel of the input image to a point in NOCS. Finding dense mappings between NOCS and an unseen object, however, is an ill-posed problem under significant shape variation. Therefore, a generalizable mapping function accommodating large amount of unknown shape variants can be difficult to learn.

In this work, we propose to learn a canonical shape space (CASS) as our unified representation. CASS is modeled by the latent space of a deep generative model of canonical 3D shapes with normalized pose and actual metric size. In particular, we train a variational auto-encoder (VAE) for generating 3D point clouds in the canonical space from an RGBD image. The VAE is trained in a cross-category fashion, exploiting the publicly available large 3D shape repositories. Since the 3D point cloud is generated with normalized pose and metric size, the encoder of the VAE learns view-factorized RGBD embedding. It maps an RGBD image in arbitrary view into a pose-independent 3D shape representation. Object pose can then be estimated via contrasting it with a pose-dependent feature of the input RGBD extracted with a separate deep neural networks (Figure 1). This circumvents the difficulty in estimating dense correspondence between two representations as in other methods .

We integrate the learning of CASS and pose and size estimation into an end-to-end trainable network which involves several key designs. First, through learning the canonical shape space with plentiful shape variants, we obtain a unified representation encompassing adequate shape variations. Second, to overcome the lack of real-world training images with 3D point clouds, we enhance the encoder of the VAE to take both RGBD images and 3D shapes as input. This allows us to train the VAE through exploiting off-the-shelf 3D shape repositories. Third, in realizing pose estimation, we opt for feature contrasting over dense correspondence, leading to better generality to unseen instances. Meanwhile, to match the distributions of pose-dependent and pose-independent features so that pose estimation can be easily trained, we propose a few crucial designs, e.g., network weight sharing and training batch mixing. Last, our VAE model is able to reconstruct a 3D point cloud of the target object with metric size, which reduces the learning difficulty by decoupling the estimation of pose and size.

Through evaluating on public category-level datasets, we show that our method archives the state-of-the-art pose accuracy and comparably high size accuracy. Our work makes the following contributions:

We propose a novel correspondence-free approach to category-level object pose and size estimation based on learned canonical shape space.

We design an end-to-end trainable deep neural network for jointly learning the canonical shape space and estimating object pose and size.

We devise several key designs to ease the network training such as distribution matching between pose-dependent and pose-independent features.

Related Work

Many works on instance-level 6D pose estimation adopt template-based methods . In these methods, a set of RGB(D) templates rendered from CAD models in various poses are matched against the input image in a sliding-window fashion, based on hand-designed or learned feature descriptors. The final pose is retrieved from the best matched template or estimated by 3D model registration. Another group of works pursue to match the target object to the corresponding 3D model. Depending on the input modality, the core task is to find 2D-to-3D , 2.5D(depth)-to-3D or 3D-to-3D correspondence . Several other works opt to learn a 6D object pose regressor directly from the feature descriptions . Brachmann et al. learn to regress an object coordinate representation which can then be used in pose estimation.

Learning effective feature representation using convolutional neural networks (CNNs) for robust matching has become a main focus of the recent literature . Another line of works utilize CNNs to detect feature points or corner points . PVNet is a unique approach of feature point detection using CNNs: A vector field is estimated for the input RGB image based on which the feature points are voted. Some other works choose to learn an end-to-end deep model that can directly regress 6D object pose from the raw RGB(D) input . SSD-6D combines single-shot object detection in RGB images with pose hypothesis regression and verification. Similar approach has also been used for multi-view active pose estimation . Wang et al. propose DenseFusion to learn pixel-wise feature extraction and pose estimation. The network can also predict a confidence for each pose hypothesis for final pose selection.

Xiao et al. train a CNN that takes as input both an image and a CAD model, and outputs object pose with respect to the 3D model. This model can generalize to the target objects which are unseen during training. However, they still require the target CAD model during inference. Thus, we classify it as an instance-level work.

Category-level approaches

There has been a large body of works on category-level object detection and 2D/3D/4D pose estimation. However, methods designed for estimating 6D poses is still scarce . Sahin et al. introduce a part-based random forest approach for this task. In their method, parts extracted from CAD model instances of some category are represented with skeletons, which are fed into a random forest for hypothesizing 6D poses. Due to the reliance on purely geometric features, this method mainly deals with depth input. Wang et al. introduce Normalized Object Coordinate Space (NOCS) as a shared canonical representation of object instances within a category. They train a region-based neural network to directly infer the pixel-wise correspondence between an RGB image to NOCS. Together with instance mask and depth map, 6D object pose is estimated using shape matching. Recently, Wang et al. realized category-level 6D pose tracking based on keypoint matching.

CASS vs. NOCS

Although both CASS and NOCS can be regarded as a unified shape space spanning intra-class variations, there are several substantial differences. First, NOCS is explicitly defined through consistently aligning all object instances of a category in a normalized 3D space. Our CASS is a shape embedding space implicitly learned with a generative model. Second, when conducting object pose estimation, NOCS is used as the target of pixel-wise correspondence, based on which 6D pose is computed geometrically. In contrast, CASS is treated as a normalized, holistic shape representation from which pose is estimated in an end-to-end and correspondence-free manner. Third, different from NOCS where the coordinates are regressed only for visible area, our network learns to reconstruct a complete 3D shape in CASS which is a global shape understanding beneficial to pose estimation.

Model

Our model is an end-to-end trainable network integrating the learning of both shape space and pose estimation. We first provide an overview of the network architecture and then elaborate the various network modules. Training details such as loss functions, parameter setting and training protocol will then follow.

Our core network is composed of three modules responsible for 1) Canonical Shape Space learning, view-factorizing RGBD embedding and point cloud reconstruction, 2) pose-dependent feature extraction and 3) pose estimation, respectively. The three components are tightly coupled and jointly trained using both synthetic and real-world data. Next, we elaborate the design of the three modules.

1 Canonical Shape Space and View Factorization

Our goal is to learn a shape space spanning as many shape variants of a category as possible, where all shapes are pose-normalized but with actual metric size. Moreover, to work with RGBD inputs, we also need a function to map an RGBG image to the point in that space representing the corresponding full shape in normalized pose and metric size. Such a mapping factorizes the view in the RGBD image so that the RGBD feature embedding is view-factorized.

We model the space of canonical shapes with the latent space of a deep generative model of pose-normalized shapes. In achieving so, we leverage the publicly available 3D shape repositories such as ShapeNet . The 3D models in ShapeNet are consistently oriented and properly scaled within each category. We sample each model into a point cloud of M=500M=500 points. The point sampled 3D shapes, X3DX_{\text{3D}}, are used to train a variational auto-encoder (VAE). The encoder employs the geometric embedding network for 3D point clouds proposed in , which is a variant of PointNet . Based on the learned feature, the decoder warps a point cloud of 3D ellipsoid to match the shape of the input point cloud. We turn the auto-encoder into a VAE by adding a sampling layer between the encoder and decoder. The learned posterior distribution z∼p(z∣X3D)z\sim p(z|X_{\text{3D}}) models the space of canonical shapes.

Learning view-factorizing RGBD embedding

Having learned the CASS, our next task is to project an RGBD image in arbitrary view to the space so that the projector functions as view factorization. Such cross-modality data projection task could be finished with the help of data correspondence between the two modalities , where metric loss is used to optimize the projector. Let us refer to this solution as correspondence-based projection. This approach can be adapted to VAE straightforwardly where a projector is trained to map data in one modality to the latent space learned for another modality based on cross-modality data correspondence. However, we found that this method leads to suboptimal point cloud reconstruction and pose estimation due to 1) possibly incorrect correspondences and 2) the compromise between the metric loss and other losses.

To address these issues, we opt for a joint embedding approach. Specifically, we learn a VAE which has two encoders mapping both RGBD images and 3D point clouds to a shared latent space. Whilst the 3D encoder adopts PointNet, the RGBD encoder employs the dense fusion architecture proposed in (we use the global feature for the whole image instead of the pixel-wise features). Our crucial design is that the two encoders, albeit having different network architectures, are trained with mixed training batch and shared training gradients. The latter means that gradients computed for either modality are back-propagated to tune both encoders. Through such mixed training, the learned shared latent space spans the joint feature space of the both modalities.

Compared to correspondence-based approaches, our joint embedding has the following advantages: Firstly, our model can be trained in a correspondence-free or unpaired fashion. This means the two modalities do not have to share object instances: It is unnecessary for the object in an RGBD image to have a corresponding 3D model in the training shape set. Secondly, our model introduces no extra loss function other than the basic ones of conventional VAEs. Thirdly and most importantly, the mixed training of the two encoders help to match the feature distributions of the two data modalities (see Figure 4), leading to better model generality and domain transferability.

In summary, the learning of the CASS and the RGBD feature embedding (denoted by FvfF_{\text{vf}}) optimizes the following loss functions:

where LreconL_{\text{recon}} and LKLL_{\text{KL}} are the reconstruction loss and KL divergence loss, respectively. X3DX_{\text{3D}} and X3DRX^{\text{R}}_{\text{3D}} are the input and reconstructed 3D point clouds, and XrgbdRX^{\text{R}}_{\text{rgbd}} and X3D∗X^{\ast}_{\text{3D}} the 3D point cloud reconstructed from the input RGBD and the corresponding ground-truth, respectively. We use Chamfer distance to measure reconstruction loss. During test, the 3D point cloud encoder (corresponding to the network branch with red arrows in Figure 2) is discarded and only the RGBD encoder is used for feature extraction.

The RGBD encoder factorizes image view, resulting in pose-independent RGBD feature (CASS code). However, the 3D encoder does not factorize object pose or size. This is because both the input and output of the 3D encoder are pose-normalized and metrically sized. It simply maps a pose-normalized shape to the canonical shape space, without processing on its pose or size. Figure 3 gives an illustrative summary of the view/pose-factorization ability for all network modules, and the pose-dependency of all involved data and features. Figure 5 (top row) shows the t-SNE visualization of the two feature embeddings. In the plot of view-factorized RGBD features, objects are clustered by category with different poses mixed together, indicating the factorization of view. The plot of geometric features of canonical point clouds (no pose) also demonstrates category-based clustering effect.

2 Pose-Dependent Feature Extraction

To facilitate pose estimation from an input RGBD image, we also extract pose-dependent features for the RGBD image. We devise two networks for extracting photometric and geometric features separately, based on the RGB and the depth images, respectively. In our network, these features are used for pose estimation through comparing against pose-dependent features, they are expected to encode the information of pose-color and pose-geometry correlations, respectively. Figure 5 (bottom row) shows the t-SNE plots of the two features, which both exhibit pose-induced subspace clustering effect.

Given the image crop containing the object of interest, we train a fully convolutional network that processes the color information into a color feature FphoF_{\text{pho}}. Similar to , the image embedding network is an auto-encoder architecture that maps an image of size H×W×3H\times W\times 3 to a pixel-wise feature map H×W×NH\times W\times N. Each pixel has a NN-dimensional vector. We then perform an average pooling over all pixel-wise features, obtaining a NN-dimensional feature for the full image.

Geometric feature extraction

Given the corresponding point patch, we utilize point-based CNNs to extract an NN-dim geometric feature FgeoF_{\text{geo}}. Here, a key design is that this point-based feature extractor can share the same network of the PointNet-based geometric feature encoder trained for CASS learning. As mentioned above, the geometry encoder is not pose-factorizing. Consequently, it can be used to extract pose-dependent geometric features. Consequently, we have a Siamese network of PointNet-based encoders, one for pose-independent CASS embedding and the other for pose-dependent geometric feature extraction (see Figure 2). Having these two tasks share network weights reduces the amount of parameters to be learned. Furthermore, it helps to match the distributions of the CASS codes and the geometric features. This makes them more comparable in facilitating feature-comparison-based pose estimation.

3 Pose and Size Estimation

We concatenate FvfF_{\text{vf}}, FphoF_{\text{pho}} and FgeoF_{\text{geo}} into a feature vector of 3N3N length and then feed it into a CNN with 1D convolutions. The output contains a rotation represented by a quaternion qq and a 3D translation vector tt. The loss function for pose prediction is defined as the discrepancy between the object point clouds transformed by the ground-truth pose and by the predicted one:

where xix_{i} is the ii-th point of the M=500M=500 sampled points for the object. [R∗∣t∗][R^{\ast}|t^{\ast}] and [R∣t][R|t] are the ground-truth and predicted poses, respectively. To handle the alignment ambiguity of symmetric objects, we relax the point-wise matching loss to Chamfer distance, similar to . Object size is calculated as the dimension of the axis-aligned bounding box (AABB) of the reconstructed 3D point cloud.

4 Training Details

The input to our method is a 640×480640\times 480 RGBD image. With the RGB image, we perform object detection and segmentation. Any off-the-shelf method can be used. For example, we utilize Mask-RCNN for the CAMERA dataset. The image crops do not need to be resize as they are fed into pixel-wise CNNs. All point patches and 3D models are re-sampled into 500500 points for PointNet feature encoding. The dimensionality of CASS code and all other features is N=1024N=1024. The DenseFusion and FoldingNet modules involved in the various network components use the same network configuration as the original works. The configuration of all other network modules such as CNNs and MLPs are given in Figure 2 (e.g., “4L” means four layers and “1D” means 1D convolutional layers). For each convolutional layer in the various modules, we add a Batch Normalization layer followed by an ReLU nonlinearity. See more details in the supplemental material.

Training protocol

We adopt a three-stage training. The first stage trains the VAE for CASS learning and view-factorizing RGBD embedding (the part shaded in light blue in Figure 2) for 8080K iterations. The size of mixed batch is 88 which randomly mixes the training data of RGBD encoding and 3D encoding. In the second stage, we fix the VAE and jointly train pose-dependent feature extraction (the light green part) and pose estimation (the light red part) for 8080K iterations. The third stage then jointly fine-tunes all parts for 4040K iterations. All training batch has the size of 88. We use an initial learning rate of 0.00010.0001 and the ADAM optimizer (β1=0.9\beta_{1}=0.9 and β2=0.999\beta_{2}=0.999) with a 1×10−61\times 10^{-6} weight decay. In each stage, we decrease the learning rate by a factor of 1010 for every 4040K iterations.

Results and evaluations

In this section, we aim to answer the following questions with both qualitative and quantitative evaluations. 1) Whether are the various network modules and design choices necessary? 2) How does our method perform in terms of pose accuracy and when does it outperform the state-of-the-arts? 3) How capable is our network in terms of single-view shape reconstruction?

We use the datasets from NOCS which contains six categories: bottle, bowl, camera, can, laptop, and mug. The dataset has two parts: a real-world dataset with 4.3K RGBD images from 7 scene videos (3 instances per category) and a synthetic dataset with 275K rendered images generated with 1085 model instances from ShapeNetCore under random views. We evaluate our method on the NOCS-REAL275 dataset, which contains 2.75K real scene images with 3 unseen instances per category. In learning the CASS, we also utilized the 3D models from the ShapeNetCore dataset.

2 Evaluation Metrics

We follow the evaluation metrics in NOCS which jointly measure the object detection and pose estimation:

IoU25 & IoU50: the average precision of object instances for which the 3D overlap between the two bounding boxes is larger than 25%25\% or 50%50\% under predicted and ground truth poses respectively;

5∘5cm, 10∘5cm & 10∘10cm: the average precision of object instances for which the the error is less than n∘n^{\circ} for rotation and mm cm for translation. We choose 5∘5cm, 10∘5cm, 10∘10cm similar to .

Additionally, we employ the Chamfer Distance (CD) and Earth Mover’s Distance (EMD) to evaluate shape reconstruction from single-view RGBD images.

3 Evaluation on NOCS-REAL275 Dataset

In Table 1, we compare our method against NOCS , which is the state-of-the-art method for category-level 6D object pose and size estimation. In their method, the network is trained to find a normalized coordinate for each pixel and then solve for the pose and size with the help of depth map. On the contrary, our method directly regresses the 6D pose by comparing pose-independent and pose-dependent features. We report the results of NOCS with 3232 pose classification bins, which is its best-performing variant. Like NOCS, our results were not post-processed, e.g., by ICP refinement, although that is potentially facilitated by our point cloud reconstruction. The results show that our method outperforms NOCS in all metrics except the IoU metrics. The slightly lower IoU values are caused by our less accurate size calculation based on point cloud reconstruction. Direct regression of 6D pose is a hard problem. The success of our method is mainly attributed to the strong view-factorized (pose-independent) feature learning with the help of CASS learning and its RGBD embedding. Figure 6 shows more detailed analysis with category-wise plots of the various evaluation metrics.

Shape reconstruction

Table 2 reports a quantitative evaluation of 3D point cloud reconstruction from RGBD input, based on the test set of NOCS-REAL275. From the table, batch mixing leads to much higher reconstruction accuracy in terms of both Chamfer Distance (CD) and Earth Mover’s Distance (EMD) metrics. This is because batch mixing ensures the distributions matching between the RGBD embedding and canonical point cloud embedding. This leads to a more accurate RGBD projection (embedding) into the latent shape space.

4 Ablation Studies

To experimentally justify the various design choices of our method, we make the following ablations (or their combinations) to our model:

w/o CASS. Train the pose & size estimation network without the CASS code as an input.

w/o Distribution Matching (DM). Replace the Siamese network with two independent modules, without matching the distribution of the CASS codes and the geometric features.

w/o Batch Mixing (BM). Remove batch mixing and use L2L_{2} distance as an additional loss to train the projection from an RGBD image in arbitrary view to the canonical shape space.

From the results reported in Table 3, we can see that CASS learning is the most important component for our method. Without CASS learning, the accuracy drops the most especially on the xx∘xx^{\circ}yyyy-cm metrics. Next to CASS learning, batch mixing is also very important factors. VAE is beneficial to model generalization to unseen objects since it helps learning a well-spanned CASS space with the normal distribution prior. However, it does lead to blurred 3D reconstruction at the same time, which may sacrifices size accuracy (see the IoU comparison in Table 1). Nevertheless, all factors together contributes the high-precision (5∘5^{\circ}55cm) estimation of pose and size.

5 Qualitative results

Figure 7 shows some visual comparisons between our method and NOCS. According to the estimate pose and scale, we draw orientated bounding box for each detected instance overlaid on top of the input RGB images. As can be observed, our method achieves better accuracy especially for size estimation under object occlusion and background distraction.

Visual results of shape reconstruction

Figure 8 shows visual results of shape reconstruction. Our method is able to reconstruct full 3D shapes in point cloud from single-view RGBD images, contrasting them with the point clouds unprojected from depth maps.

Conclusion

We have presented a novel correspondence-free approach to category-level object pose and size estimation. This is achieved by learning a shape space of 3D models in normalized pose and metric size based on deep generative model. The input RGBD image is embedded into the shape space, extracting pose-independent features. Pose estimation is realized by comparing the pose-independent and pose-dependent features. Evaluation shows that our method arrives at the state-of-the-art performance.

Our current method has a few limitations on which we aim to improve as future work. First, our method cannot handle well very complex shapes due to the difficulty in reconstructing shapes with complicated geometry (e.g. high genus). In this aspect, our method can be enhanced by learning a more powerful shape reconstruction with, e.g., volumetric 3D representation. Second, our current method does not close the loop in terms of utilizing the reconstructed shape geometry to guide/supervise the training of pose estimation. This may lead to a unsupervised or self-taught approach which we plan to investigate in a future work. Third, our method still cannot achieve very high precision, as reflected by the relatively lower accuracy for the 5∘5^{\circ}55cm metric. This may be an inherent limitation for a correspondence-free or sparse approach. Note, however, our method did not use ICP to refine the pose or size. Last, we plan to extend our current framework to online object pose tracking similar to .

Acknowledgement

We thank the anonymous reviewers for the valuable suggestions. We are grateful to Chen Wang, one of the authors of DenseFusion, for the help and discussion. This work was supported in part by the National Key Research and Development Program of China (No. 2018AAA0102200), the NSFC (61572507, 61532003, 61622212, 61902419), the NUDT Research Grants (No.ZK19-30) and the Natural Science Foundation of Hunan Province for Distinguished Young Scientists (2017JJ1002).

References