Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation

Meng Tian, Marcelo H Ang, Gim Hee Lee

Introduction

Accurate 6D object pose estimation plays an important role in a variety of tasks, such as augmented reality, robotic manipulation, scene understanding, etc. In recent years, substantial progress has been made for instance-level 6D object pose estimation, where the exact 3D object models for pose estimation are given. Unfortunately, these methods cannot be directly generalized to category-level 6D object pose estimation on new object instances with unknown 3D models. Consequently, the category, 6D pose and size of the objects have to be concurrently estimated. Although some other object instances from each category are provided as priors, the high variation of object shapes within the same category makes generalization to new object instances extremely challenging.

To the best of our knowledge, is the first work to address the 6D object pose estimation problem at category-level. This approach defines 6D pose on semantically selected centers and trains a part-based random forest to recover the pose. However, building part representations upon 3D skeleton structures limits the generalization capability across unseen object instances. Another work proposes the first data-driven solution and creates a benchmark dataset for this task. They introduce the Normalized Object Coordinate Space (NOCS) to represent different object instances within a category in a unified manner. A region-based network is trained to infer correspondences from object pixels to the points in NOCS. Class label and instance mask of each object are also obtained at the same time. These predictions are used together with the depth map to estimate the 6D pose and size of the object via point matching. However, the lack of explicit representation of shape variations limits their performance.

In this work, we propose to reconstruct the complete object models in the NOCS to capture the intra-class shape variation. More specifically, we first learn the categorical shape priors from the given object instances, and then train a network to estimate the deformation field of the shape prior (that is used to get the reconstructed object model) and the correspondences between object observation and the reconstructed model. The shape prior serves as prior knowledge of the category and encodes geometrical characteristics that are shared by objects of a given category. The deformation predicted by our network captures the instance-specific shape details, i.e. shape variation of that particular instance. We present a method which is applicable across different object categories and data representations to learn the shape priors. In particular, an autoencoder is trained on a collection of object models from various categories. For each category, we compute the mean latent embedding over all instances in the respective category. The categorical shape prior is constructed by passing the mean embedding through a decoder. Note that there is no restriction on the data representation (point cloud, mesh, or voxel) of shape priors or collected models as long as we choose a proper architecture for the encoder and decoder.

We use the Umeyama algorithm to recover the 6D pose and metric size of the object from the correspondences estimated by our network that maps the point cloud obtained from the observed depth map to the points in NOCS. We evaluate our method on two standard benchmarks. Extensive experiments demonstrate the advantage of our network and prove the effectiveness of explicitly modeling the deformation. In summary, the main contributions of this work are:

We propose a novel deep network for category-level 6D object pose and size estimation; our network explicitly models the deformation from the categorical shape prior to the object model.

We present a learning-based method which utilizes the latent embeddings to construct the shape prior; our method is applicable across different categories and data representations.

Our network achieves significantly higher mean average precisions on both synthetic and real-world benchmark datasets.

Related Work

Instance-Level Pose Estimation. Existing instance-level pose estimation approaches broadly fall into three categories. The first category casts votes in the pose space and further refines coarse pose with algorithms such as iterative closest point. LINEMOD uses holistic template matching to find the nearest viewpoint. generates a latent code for the input image and search for its nearest neighbor in the codebook. aggregate the 6D votes cast by locally-sampled RGB-D patches. The second category directly maps input image to object pose. extend 2D object detection network such that it can predict orientation as an add-on to the identity and 2D bounding box of the object. regress 6D pose from RGB-D images in an end-to-end framework. The third category relies on establishing point correspondences. regress the corresponding 3D object coordinate for each foreground pixel. detect the keypoints of the object on image and then solve a Perspective-n-Point problem. estimates a dense 2D-3D correspondence map between the input image and object model. Although our approach follows the approach from the third category, our task focuses on a more general setting where the object models are not available during inference.

Category-Level Object Detection. The task of 3D object detection aims to estimate 3D bounding boxes of objects in the scene. runs sliding windows in 3D space and generates amodal proposals for objects. first generate 2D object proposals and then lift the proposals to 3D space. are single-stage detectors which directly detect objects from 3D data. Although the above mentioned methods address the problem at category-level, the considered objects are usually constrained to the ground surface, e.g. instances of typical furniture classes in indoor scenes, cars, pedestrians, and cyclists in outdoor scenes. Consequently, the assumption that rotation is constrained to be only along the gravity direction has to be made. On the contrary, our approach can recover the full 6D pose of objects.

Category-Level Pose Estimation. There are only a few pioneering works focusing on estimating 6D pose of unseen objects. leverages a generative representation of 3D objects and produces a multimodal distribution over poses with mixture density network. However, only rotation is considered in their work. introduces a part-based random forest which employs simple depth comparison features, but our approach deals with RGB-D images. proposes a canonical representation for all instances within an object category. Our approach also makes use of this representation. Instead of directly regressing the coordinates in NOCS, we account for intra-class shape variation by explicitly modeling the deformation from shape prior to object model. trains a variational autoencoder to generate the complete object model. However, the reconstructed shape is not utilized for pose estimation. In our network, shape reconstruction and pose estimation are integrated together. proposes the first category-level pose tracker, while our approach performs pose estimation without using temporal information.

Shape Deformation. 3D deformation is commonly applied for object reconstruction from a single image. use free-form deformation in conjunction with voxel, mesh and point cloud representations, respectively. starts from a coarse shape and predicts a series of deformations to progressively improve the geometry. Similar to us, also supervise the deformation with global feature of the target. However, we circumvent the fixed topology assumption of mesh representation by using point clouds instead.

Our Method

Background. Given an RGB-D image as the input, our goal is to detect and localize all visible object instances in the 3D space. The object instances are not seen previously, but must come from known categories. Each object is represented by a class label and an amodal 3D bounding box parameterized by its 6D pose and size. The 6D pose is defined to be the rigid transformation (i.e. rotation and translation) that transforms the object from the reference to the camera coordinate frame. It is common to choose the coordinate frame of the given 3D object models as reference in instance-level 6D object pose estimation. Unfortunately, this is not viable for our category-level task since the instances of the 3D models are not available. To mitigate this problem, we leverage on the Normalized Object Coordinate Space (NOCS) – a shared canonical representation for all possible object instances within a category proposed in . The categorical 6D object pose and size estimation problem is then reduced to finding the similarity transformation between the observed depth map of each object instance and its corresponding points in the NOCS (i.e. NOCS coordinates).

Overview. In contrast to that directly outputs the NOCS coordinates from a Convolutional Neural Network (CNN), we propose an intermediate step to estimate the deformation of a pre-learned shape prior to improve the learning of intra-class shape variation. Our shape priors are learned from a collection of models spanning all categories (Section 3.1). As shown in Fig. 1, our approach consists of three stages. The first stage performs instance segmentation on color image using an off-the-shelf network (e.g. Mask R-CNN ). Next we convert the masked depth map into a point cloud with camera intrinsic parameters for each instance and crop an image patch according to the bounding box of the mask. Taking the point cloud, image patch, and the corresponding shape prior as inputs, our network outputs a deformation field that deforms the shape prior into the shape of the desired object instance (a.k.a. reconstructed model). Furthermore, our network outputs a set of correspondences that associates each point in the point cloud obtained from the observed depth map of the object instance with the points of the reconstructed model. This set of correspondences is used to mask the reconstructed model into the NOCS coordinates (Section 3.2). Finally, the 6D pose and size of the object can be estimated by registering the NOCS coordinates and the point cloud obtained from the observed depth map (Section 3.3).

Although object shape varies among different instances, an investigation over the 3D models reveals that objects of the same category (especially artificially generated objects) tend to have semantically and geometrically similar components. For example, cameras are usually made up of a nearly cuboid body and a cylindrical lens; and mugs are typically cylindrical with a handle. These categorical characteristics provide strong priors on the shape reconstruction of novel instances. We propose the learning of a mean shape to capture the high-level characteristics from all the available models for each respective category. To this end, we first train an autoencoder with all available object models and then compute the mean latent embedding of each object category with the encoder. These latent embeddings are passed into the decoder to get the mean shape priors for each object category. Unlike methods such as simple averaging and principal component analysis (PCA) that operate on voxel representations, our autoencoder framework can be easily altered to take any 3D representations.

Given a shape collection M={Mci ∣ i=1,2,⋯ ,N; c=1,2,⋯ ,C}\mathcal{M}=\{M_{c}^{i}\,\mid\,i=1,2,\cdots,N;\,c=1,2,\cdots,C\}, where MciM_{c}^{i} is the 3D point cloud model of instance ii from category cc, we independently apply a similarity transformation to each model such that it is properly aligned in the NOCS. This step ensures that the learned shape prior has the same scale and orientation as the target shape to be reconstructed. The encoder Φ\Phi takes the point cloud and outputs a low-dimensional feature vector, i.e. the latent embedding zci∈Rnz_{c}^{i}\in\mathcal{R}^{n}. The decoder Ψ\Psi takes this feature vector and outputs a point cloud that reconstructs the input:

Specifically, we adopt the PointNet-like encoder proposed in , and a three-layer fully-connected decoder as shown in Fig. 2a. The reconstruction error is measured by the Chamfer distance:

The autoencoder is trained on a shape collection by minimizing the reconstruction error. Once the training is converged, we obtain the latent embeddings {zci}\{z_{c}^{i}\} of all instances in M\mathcal{M}. Although not explicitly enforced during training, these latent embeddings form clusters in the latent space according to their categories. Fig. 2b visualizes the clustering effect of the embeddings. We use T-SNE to further embed these features in R2\mathcal{R}^{2} for visualization. Similar clustering results are also observed on a different set of models . Based on this observation, we compute the mean latent embedding for each category and then pass it through the decoder to construct the shape prior:

The resulting categorical shape priors {Mc}\{M_{c}\} are shown in Fig. 2c.

2 Our Network Architecture

We denote the observation of an object instance as (V,I)(V,I), where V∈RNv×3V\in\mathcal{R}^{N_{v}\times 3} is the point cloud and I∈RH×W×3I\in\mathcal{R}^{H\times W\times 3} is the image patch. NvN_{v} denotes the number of foreground pixels with a valid depth value. The corresponding shape prior is Mc∈RNc×3M_{c}\in\mathcal{R}^{N_{c}\times 3}, where NcN_{c} is the number of points in McM_{c}. Our network takes VV, II and McM_{c} as inputs, and outputs the per-point deformation field D∈RNc×3D\in\mathcal{R}^{N_{c}\times 3} and a correspondence matrix A∈RNv×NcA\in\mathcal{R}^{N_{v}\times N_{c}}. The final reconstructed model is M=Mc+DM=M_{c}+D. Each row of AA sums to 1 since it represents the soft correspondences between a point in VV and all points in MM. As shown in Fig. 3, our network is composed of four parts: (1) extracts features from the object instance; (2) extracts features from the shape prior; (3) regresses the deformation field DD; and (4) estimates the correspondence matrix AA.

On the consideration that the depth and color are two different modalities, we follow the pixel-wise dense fusion approach proposed in to effectively extract RGB-D features from the observation. For point cloud VV, we use an embedding network similar to PointNet to generate per-point geometric features by mapping each point in VV to the dgd_{g}-dimensional feature space. The image patch II is processed with a fully convolutional network which follows an encoder-decoder architecture and maps II to RH×W×dc\mathcal{R}^{H\times W\times d_{c}}. Next we associate the geometric feature of each point with its corresponding color feature and concatenate the feature pairs. Since each point in VV has a corresponding pixel in II, not vice versa, redundant color features are discarded. The concatenated features are termed “instance point features” and fed to another shared multi-layer perceptron. An average pooling layer is used to generate the “instance global feature”. The categorical shape prior McM_{c} is a point cloud with purely geometric information. We apply a simpler embedding network to extract the “category point features” and the “category global feature”.

The shape prior McM_{c} provides the prior knowledge of the category, i.e. the coarse shape geometry and canonical pose. Although the observation (V,I)(V,I) is partial, it provides instance-specific details of the target shape. A natural way to reconstruct the object in NOCS is to deform McM_{c} under the guidance of (V,I)(V,I). Consequently, we concatenate the category and instance global features, and enrich the category point features with the concatenated features. The obtained feature vectors are successively convolved with 1×11\times 1 kernels to generate the deformation field DD. Similar intuition and feature concatenation strategy also apply to the estimation of AA. We combine the instance point features and global feature to aggregate both local and global information for each point. Each point in VV gets mapped to the points of the reconstructed model through concatenation with the category global feature. We obtain the NOCS coordinates, denoted as PP, of the points in VV by multiplying AA and MM, i.e.

3 6D Pose Estimation

Our goal is to estimate the 6D pose and size of the object instance. Given depth observation VV and its NOCS coordinates PP, the optimal similarity transformation parameters (rotation, translation, and scaling) can be computed by solving the absolute orientation problem using Umeyama algorithm . We also implement the RANSAC algorithm for robust estimation.

4 Loss Functions

In this section, we define the loss functions used to train our network, and explain how we handle object symmetry during training.

Reconstruction Loss. We assume that ground-truth model MgtM_{gt} is available during training. The deformation field DD is supervised indirectly by minimizing the Chamfer distance (c.f. Eq. 2) between MM and MgtM_{gt}, i.e. Lcd=dCD(M,Mgt)=dCD(Mc+D,Mgt)L_{\text{cd}}=d_{\text{CD}}(M,M_{gt})=d_{\text{CD}}(M_{c}+D,M_{gt}).

Correspondence Loss. It is impractical and unnecessary to pre-compute the ground-truth value for AA. Alternatively, we supervise AA indirectly via the NOCS coordinates PP (which is a result of applying the correspondence matrix on the reconstructed model) since the ground-truth NOCS coordinates PgtP_{gt} can be obtained easily from the object model and its 6D pose through image rendering. We use the smooth L1L_{1} loss function:

where x=(x1,x2,x3)∈P\mathbf{x}=(x_{1},x_{2},x_{3})\in P, and y=(y1,y2,y3)∈Pgt\mathbf{y}=(y_{1},y_{2},y_{3})\in P_{gt}.

Object symmetry is an inevitable problem for pose estimation algorithms, especially for those that require supervised training. For symmetrical objects, there exists at least one rotation such that appearance of the object is preserved under this rotation. In other words, two observations of a symmetric object can be very similar but with different rotation labels. We follow the solution proposed by to map ambiguous rotations to a canonical one. More specifically, the Map operator for an arbitrary rotation R∈SO(3)R\in SO(3) is defined as:

where the proper symmetry group S(Mci)\mathcal{S}(M_{c}^{i}) is the set of rotations which preserve the appearance of a given object MciM_{c}^{i}. The experimental datasets assume continuous symmetry and the axis of symmetry is the y-axis of the NOCS. Hence, S^\hat{S} takes the following form:

where R11R_{11}, R13R_{13}, R31R_{31}, and R33R_{33} are the elements of RR. During training, we apply the Map operator to the rotation label: Rgt←RgtS^R_{gt}\leftarrow R_{gt}\hat{S} to eliminate the rotation ambiguity of any symmetric object with the ground-truth pose (Rgt,Tgt)(R_{gt},T_{gt}). In practice, our network is supervised by ground-truth NOCS coordinates PgtP_{gt}. Equivalently, we transform PgtP_{gt} by S^T\hat{S}^{T}: Pgt←S^TPgtP_{gt}\leftarrow\hat{S}^{T}P_{gt}.

Regularization Losses. Row AiA_{i} of the matrix AA represents the distribution over the correspondences between ii-th point of VV and the points in MM. AiA_{i} can be understood as a relaxed one-hot vector, since each point of VV usually can be well approximated by at most three points of MM. We encourage AiA_{i} to be a peaked distribution by minimizing the average cross entropy: Lentropy=1Nv∑i∑j−Ai,jlog⁡Ai,jL_{\text{entropy}}=\frac{1}{N_{v}}\sum_{i}\sum_{j}-A_{i,j}\log A_{i,j}. We also regularize DD to discourage large deformations: Ldef=1NC∑di∈D∥di∥2L_{\text{def}}=\frac{1}{N_{C}}\sum_{\mathbf{d}_{i}\in D}\|\mathbf{d}_{i}\|_{2}. Minimal deformation preserves the semantic consistency between shape prior and the reconstructed model. For example, we want that the point belongs to the handle of the “mug” prior remains in the handle after deformation. This consistency loss is beneficial for the prediction of the correspondence matrix AA.

In summary, the overall objective is a weighted sum of all four losses:

Experiments

Datasets. The CAMERA dataset is generated by rendering and compositing synthetic objects into real scenes in a context-aware manner. In total, there are 300K composite images, where 25K are set aside for evaluation. The training set contains 1085 object instances selected from 6 different categories - bottle, bowl, camera, can, laptop and mug. The evaluation set contains 184 different instances. The REAL dataset is complementary to the CAMERA. It captures 4300 real-world images of 7 scenes for training, and 2750 real-world images of 6 scenes for evaluation. Each set contains 18 real objects spanning the 6 categories. The two evaluation sets are referred to as CAMERA25 and REAL275.

Evaluation Metric. Following , we independently evaluate the performance of 3D object detection and 6D pose estimation. We report the average precision at different Intersection-Over-Union (IoU) thresholds for object detection. For 6D pose evaluation, the average precision is computed at n∘ mcm\text{n}^{\circ}\,\text{m}cm. We ignore the rotational error around the axis of symmetry for symmetrical object categories (e.g. bottle, bowl, and can). Specially, we treat mug as symmetric object in the absence of the handle, and asymmetric object otherwise.

Baseline. Wang et al. is currently the only publicly available code and datasets for the 6D object pose and size estimation task. Futhermore, it is also the state-of-the-art performance on the task. Hence, we choose as our baseline for comparison.

2 Implementation Details

We collect all the instances in the CAMERA training dataset to train the autoencoder. Shape priors are learned from this collection and used in all experiments. Each prior consists of 1024 points. We use the Mask R-CNN implemented by matterport for instance segmentation. For each detected instance, we resize the image patch to 192×192192\times 192, and randomly sample 1024 points by repetition (if insufficient point count) or downsampling (if sufficient point count). To extract instance color features, we choose the PSPNet with ResNet-18 as backbone. We randomly select 5 point-pairs to generate a hypothesis for the RANSAC-based pose fitting. The maximum number of iteration is 128 and inlier threshold is set to 10% of the object diameter. For the hyperparameters of the total loss, we empirically find that λ1=5.0\lambda_{1}=5.0, λ2=1.0\lambda_{2}=1.0, λ3=1e−4\lambda_{3}=1e-4, and λ4=0.01\lambda_{4}=0.01 are good choices.

3 Comparison to Baseline

We compare our approach to the Baseline on CAMERA25 and REAL275. Quantitative results are summarized in Table 1.

CAMERA25. In the setting of estimating 6D object pose and size from an RGB-D image, we achieve a mAP of 83.1% for 3D IoU at 0.75, and a mAP of 54.3% for 6D pose at 5∘ 2cm5^{\circ}\,2\text{cm}. Our results are 14% and 22% higher than the Baseline , respectively. We naively remove the depth input and related sub-networks in our network (i.e. RGB image as the only input) to make fair comparisons with the Baseline , which takes an RGB image as its input. As shown in Table 1, our results without depth input are still significantly better than the Baseline (i.e. +15.5% and +17.9%). On one hand, this experiment shows the advantage of explicit handling of the intra-class shape variation, and the effectiveness of our method which reconstructs the object via deformation. On the other hand, it also shows that adding depth to the network does help to improve overall performance, although our improved performance does not rely on it solely. Given that depth image is required to uniquely determine the scale of the object, we recommend it in practical applications. The top row of Fig. 4 shows the average precision at different error thresholds for all 6 object categories. It provides independent analysis for 3D IoU, rotation, and translation error.

REAL275. The REAL training set only contains 3 object instances per category, we enlarge this training set such that the network can generalize well to unseen objects. Following the Baseline , we randomly select data from CAMERA and REAL training set according to a ratio of 3:13:1. In fair comparison to the Baseline , our approach improves the mAP by 23.1% for 3D IoU at 0.75 and 12.1% for 6D pose at 5∘ 2cm5^{\circ}\,2\text{cm}. In strict comparison, we can still outperform by 16.4% and 8.5%, respectively. These results provide further evidence to support our approach. Fig. 4 (bottom) shows more detailed analysis of the errors.

4 Evaluation of Shape Reconstruction

To evaluate the quality of the reconstruction, we compute the CD metric (c.f. Eq. 2) of the reconstructed model from our method with the ground truth model in the NOCS. We get a CD metric of 1.97 on CAMERA25 and 3.17 on REAL275. In comparison, the CD metrics are 3.70 and 4.41 on the respective dataset for the shape priors from our autoencoder. The better CD metrics of the reconstructed models compared to the shape priors show that the deformation estimation in our framework improves the quality of the 3D model reconstruction. Table 2 shows the CD metric of our reconstructed models and the shape priors for each category.

5 Ablation Studies

Different shape priors. We first evaluate how different shape priors influence the performance. All settings are kept the same in this experiment except for the choices of the priors. Results are summarized in Table 3 and 4. “Embedding” refers to the priors obtained from decoding the mean latent embeddings. We also try the instance whose latent embedding has the minimum L2L_{2} distance to the mean latent embedding (denoted as “NN”). In addition, we explore random selection of one instance per category from the shape collection to compose our priors (denoted as “Random”). In general, our approach remains stable under different priors. Our network can adapt to different shape priors because the deformation is explicitly estimated. We achieve the best result for accurate pose (i.e. 5∘ 2cm5^{\circ}\,2\text{cm}) estimation when the learned categorical shape prior is used. Since our main target is to recover the 6D pose, we choose “Embedding” as our best model. To validate whether the priors are necessary, we use a point cloud uniformly sampled from a sphere of diameter one as our prior (denoted as “None”). The mAP decreases by 3.7% on real dataset when there are no priors, but the best result is achieved for object size estimation. Although shape priors are beneficial for estimating 6D pose, they sometimes bias shape reconstruction.

Directly regress the NOCS Coordinates? As indicated by Eq. 4, our approach decouples the NOCS coordinates PP to shape reconstruction MM and dense correspondences AA. However, both the network architecture and the training will be much simpler when we follow to regress PP directly (denoted as “Regression” in Table 3 and 4). For 6D pose estimation, the mAP of “Regression” at 5∘ 2cm5^{\circ}\,2\text{cm} is notably lower than “Embedding” on CAMERA25 (-3.1%) and REAL275 (-5.6%). This result further supports the benefit of handling shape variation via reconstruction over naive regression of the NOCS coordinates. “Regression” achieves slightly better mAP for object size estimation since it only finds the NOCS coordinates for the observed part, while “Embedding” needs to complete the unknown part of the object.

Regularization losses. To validate the necessity of the two regularization losses, we train the network without regularizing deformation or correspondence, while keeping all the rest same as ”Embedding”. The mean average precisions of both variants are still comparable to ”Embedding” on synthetic dataset. However, mAP of 6D pose at 5∘ 2cm5^{\circ}\,2\text{cm} drops noticeably (-5.9% and -3.6%) on the more difficult real dataset.

6 Qualitative Results.

In Fig. 5, we provide several qualitative results from both synthetic and real instances. The 6D pose and object size can be reliably recovered from noisy point correspondences using RANSAC-based pose fitting. Shape reconstruction can capture the variations between instances. The qualities of our predictions are generally better on synthetic data than real data, which is an indication that observation noise needs more attention in our future work. Out of the six categories, camera shows less accurate reconstruction due to its more complicated and varying geometry.

Conclusions

We present a novel approach for category-level 6D object pose estimation. Our network explicitly models intra-class shape variation by the estimation of the deformation from a shape prior to the object model. Shape priors are learned form a collection of object models and constructed in the latent space. Experiments on synthetic and real datasets demonstrate the advantage of our proposed approach.

This research is supported in parts by the Singapore MOE Tier 1 grant R-252-000-A65-114, and the Agency for Science, Technology and Research (A*STAR) under its AME Programmatic Funding Scheme (Project #A18A2b0046).

References

A Comparison to CASS

CASS is the latest work on category-level 6D object pose and size estimation. Similar to our work, they reconstruct the complete object model in the canonical space as a by-product. However, they train a variational autoencoder to generate the point cloud, while we estimate the deformation field of the corresponding shape prior. In addition, they directly regresses the pose and size by comparing pose-independent and pose-dependent features, while we recover the pose by establishing dense correspondences. As shown in Table 5, our approach significantly outperforms CASS in pose accuracy. This demonstrates the superiority of our correspondence-based approach over direct pose regression.

B Comparison to 6-PACK

6-PACK is the state-of-the-art category-level 6D pose tracker. Although our approach does not require pose initialization nor leverages on temporal consistency, we still achieve comparable accuracy on REAL275 at 5∘ 5cm5^{\circ}\,5\text{cm} (30.4% compared to 33.3%). More importantly, the accuracy of 6-PACK drops below 30% when the first 40 frames of a sequence (460 frames on average) are excluded from evaluation. This indicates that 6-PACK is highly dependent on pose initialization for higher accuracy. In contrast, the accuracy of our method remains stable since it is a pose estimation method.

C Qualitative Results

Fig. 6 shows the per-frame pose detection results of our approach. The results are better on synthetic data (top two rows) than on real data (bottom two rows). This performance gap is mainly induced by the observation noise, which has greater influence on objects with complicated geometric shape (e.g. camera).

D Runtime Analysis

Given RGB-D images with resolution of 640×480640\times 480 and mean object count of 4, our implementation approximately runs at 4 FPS on a desktop with an Intel Core i7-5960X CPU (3.0 GHz) and a NVIDIA GTX 1080Ti GPU. Specifically, it takes an average time of 130 ms for instance segmentation, 100 ms for network inference, and 20 ms for pose alignment.

E Visualization of Shape Priors

In Fig. 7, we visualize the different categorical shape priors used in our ablation studies.

F Derivation of the Map Operator

We first give the proposition from , and then derive the Map operator used in our work as a corollary. Given an object MciM_{c}^{i}, the proper symmetry group S(Mci)\mathcal{S}(M_{c}^{i}) is defined as:

where I(Mci,p)\mathcal{I}(M_{c}^{i},\mathbf{p}) is the image of object MciM_{c}^{i} under pose p\mathbf{p}. Intuitively, S(Mci)\mathcal{S}(M_{c}^{i}) consists of rotations which preserve the appearance of a given object.

Given a proper symmetry group S(Mci)\mathcal{S}(M_{c}^{i}), ∀ R∈SO(3)\forall\,R\in SO(3), the Map operator is defined as:

where I3I_{3} is an 3×33\times 3 identity matrix. Then, Map(R1)=Map(R2)⟺I(Mci,R1)=I(Mci,R2)\textup{Map}(R_{1})=\textup{Map}(R_{2})\Longleftrightarrow\mathcal{I}(M_{c}^{i},R_{1})=\mathcal{I}(M_{c}^{i},R_{2}) .

The proof is omitted for brevity, refer to for the details. The Map operator used in our work can then be derived directly from Proposition 1.

The Map operator for symmetrical objects around the y-axis is given by:

Assuming the object MciM_{c}^{i} is symmetrical around the y-axis of the object coordinate system, then SS has the following form:

We minimize the Froebenius norm over θ\theta to solve for the Map: