Pose from Shape: Deep Pose Estimation for Arbitrary 3D Objects
Yang Xiao, Xuchong Qiu, Pierre-Alain Langlois, Mathieu Aubry, Renaud Marlet
Introduction
Imagine a robot that needs to interact with a new type of object not belonging to any pre-defined category, such as a newly manufactured object in a workshop. Using existing single-view pose estimation approaches for this new object would require stopping the robot and training a specific network for this object before taking any further action. Here we propose an approach that can directly take as input a 3D model of the new object and estimate the pose of the object in images relatively to this model, without any additional training procedure. We argue that such a capability is necessary for applications such as robotics “in the wild”, where new objects of unfamiliar categories can occur routinely at any time and have to be manipulated or taken into account for action. It also applies to virtual reality with similar circumstances.
To overcome the fact that deep pose estimation methods were category-specific, i.e., predicted different orientations according to object category, recent works [Grabner et al.(2018)Grabner, Roth, and Lepetit, Zhou et al.(2018)Zhou, Karpur, Luo, and Huang] have proposed to perform category-agnostic pose estimation on rigid objects, producing a single prediction. However, [Grabner et al.(2018)Grabner, Roth, and Lepetit] only evaluated on object categories that were included in the training data, while [Zhou et al.(2018)Zhou, Karpur, Luo, and Huang] required the testing categories to be similar to the training data. On the contrary, we want to stress that our method works on novel objects that can be widely different from those seen at training time. For example, we can train only on man-made objects, but still be able to estimate the pose of animals such as horses, whereas not a single animal has been seen in the training data (cf. Fig. 1 and 3). Our method is similar to category-agnostic approaches in that it only produces one pose prediction and does not require additional training to produce predictions on novel categories. However, it is also instance-specific, because it takes as input a 3D model of the object of interest.
Indeed, our key idea is that viewpoint is better defined for a single object instance given its 3D shape than for whole object categories. Our work can be viewed as leveraging the recent advances in deep 3D model representations [Su et al.(2015a)Su, Maji, Kalogerakis, and Learned-Miller, Qi et al.(2017a)Qi, Su, Mo, and Guibas, Qi et al.(2017b)Qi, Yi, Su, and Guibas] for the problem of pose estimation. We show that using 3D model information also boosts performances on known categories, even when the information is only approximate, as in the Pascal3D+ [Xiang et al.(2014)Xiang, Mottaghi, and Savarese] dataset.
When an exact 3D model of the object is known, as in the LINEMOD [Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab] dataset, state-of-the-art results are typically obtained by first performing a coarse viewpoint estimation and then applying a pose-refinement approach, typically matching rendered images of the 3D model to the target image. Our method is designed to perform the coarse alignment. Pose-refinement can be performed after applying our method using a classical approach based on ICP or the recent DeepIM [Li et al.(2018b)Li, Wang, Ji, Xiang, and Fox] method. Note that while DeepIM only performs refinement, it is similar to our work in the sense that it is category agnostic and leverages some knowledge of the 3D model, using a view rendered in the estimated pose, to predict its pose update.
To the best of our knowledge, we present the first deep learning approach to category-free viewpoint estimation, which can estimate the pose of any object conditioned only on its 3D model, whether or not it is similar to objects seen at training time.
We can learn with and use “shapes in the wild”, whose reference frame do not have to be consistent with a canonical orientation, simplifying pose supervision.
We demonstrate on a large variety of datasets [Xiang et al.(2014)Xiang, Mottaghi, and Savarese, Xiang et al.(2016)Xiang, Kim, Chen, Ji, Choy, Su, Mottaghi, Guibas, and Savarese, Sun et al.(2018)Sun, Wu, Zhang, Zhang, Zhang, Xue, Tenenbaum, and Freeman, Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab] that adding 3D knowledge to pose estimation networks provides performance boosts when applied to objects of known categories, and meaningful performances on previously unseen objects.
Related Work
In this section, we discuss pose estimation of a rigid object from a single RGB image first in the case where the 3D model of the object is known, then when the 3D model is unknown.
Traditional methods to estimate the pose of a given 3D shape in an image can be roughly divided into feature-matching methods and template-matching methods. Feature-matching methods try to extract local features from the image, match them to the given object 3D model and then use a variant of PnP algorithm to recover the 6D pose based on estimated 2D-to-3D correspondences. Increasingly robust local feature descriptors [Lowe(2004), Tola et al.(2010)Tola, Lepetit, and Fua, Tulsiani and Malik(2015), Pavlakos et al.(2017)Pavlakos, Zhou, Chan, Derpanis, and Daniilidis] and more effective variants of PnP algorithms [Lepetit et al.(2009)Lepetit, Moreno-Noguer, and Fua, Zheng et al.(2013)Zheng, Kuang, Sugimoto, Astrom, and Okutomi, Li et al.(2012)Li, Xu, and Xie, Ferraz et al.(2014)Ferraz, Binefa, and Moreno-Noguer] have been used in this type of pipeline. Pixel-level prediction, rather than detected features, has also been proposed [Brachmann et al.(2016)Brachmann, Michel, Krull, Yang, Gumhold, and Rother]. Although performing well on textured objects, these methods usually struggle with poorly-textured objects. To deal with this type of objects, template-matching methods try to match the observed object to a stored template [Li et al.(2011)Li, Wang, Yin, and Wang, Lowe(1991), Hinterstoisser et al.(2012a)Hinterstoisser, Cagniart, Ilic, Sturm, Navab, Fua, and Lepetit, Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab]. However, they perform badly in the case of partial occlusion or truncation.
More recently, deep models have been trained for pose estimation from an image of a known or estimated 3D model. Most methods estimate the 2D position in the test image of the projections of the object 3D bounding box [Rad and Lepetit(2017), Tekin et al.(2018)Tekin, Sinha, and Fua, Oberweger et al.(2018)Oberweger, Rad, and Lepetit, Grabner et al.(2018)Grabner, Roth, and Lepetit] or object semantic keypoints [Pavlakos et al.(2017)Pavlakos, Zhou, Chan, Derpanis, and Daniilidis, Georgakis et al.(2018)Georgakis, Karanam, Wu, and Kosecka] to find 2D-to-3D correspondences and then apply a variant of the PnP algorithm, as feature-matching methods. Once a coarse pose has been estimated, deep refinement approaches in the spirit of template-based methods have also been proposed [Manhardt et al.(2018)Manhardt, Kehl, Navab, and Tombari, Li et al.(2018b)Li, Wang, Ji, Xiang, and Fox].
Pose estimation not explicitly using object shape.
In recent years, with the release of large-scale datasets [Geiger et al.(2012)Geiger, Lenz, and Urtasun, Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab, Xiang et al.(2014)Xiang, Mottaghi, and Savarese, Xiang et al.(2016)Xiang, Kim, Chen, Ji, Choy, Su, Mottaghi, Guibas, and Savarese, Sun et al.(2018)Sun, Wu, Zhang, Zhang, Zhang, Xue, Tenenbaum, and Freeman], data-driven learning methods (on real and/or synthetic data) have been introduced which do not rely on an explicit knowledge of the 3D models. These can roughly be separated into methods that estimate the pose of any object of a training category and methods that focus on a single object or scene. For category-wise pose estimation, a canonical view is required for each category with respect to which the viewpoint is estimated. The prediction can be cast as a regression problem [Osadchy et al.(2007)Osadchy, Cun, and Miller, Penedones et al.(2012)Penedones, Collobert, Fleuret, and Grangier, Massa et al.(2016)Massa, Marlet, and Aubry], a classification problem [Tulsiani and Malik(2015), Su et al.(2015b)Su, Qi, Li, and Guibas, Elhoseiny et al.(2016)Elhoseiny, El-Gaaly, Bakry, and Elgammal] or a combination of both [Mousavian et al.(2017)Mousavian, Anguelov, Flynn, and Kosecka, Güler et al.(2017)Güler, Trigeorgis, Antonakos, Snape, Zafeiriou, and Kokkinos, Li et al.(2018a)Li, Bai, and Hager, Mahendran et al.(2018)Mahendran, Ali, and Vidal]. Besides, Zhou et al\bmvaOneDotdirectly regress category-agnostic 3D keypoints and estimate a similarity between image and world coordinate systems [Zhou et al.(2018)Zhou, Karpur, Luo, and Huang]. Following the same strategy, it is also possible to estimate the pose of a camera with respect to a single 3D model but without actually using the 3D model information. Many recent works have applied this strategy to recover the full 6-DoF pose for object [Tjaden et al.(2017)Tjaden, Schwanecke, and Schömer, Mousavian et al.(2017)Mousavian, Anguelov, Flynn, and Kosecka, Kehl et al.(2017)Kehl, Manhardt, Tombari, Ilic, and Navab, Xiang et al.(2018)Xiang, Schmidt, Narayanan, and Fox, Li et al.(2018a)Li, Bai, and Hager] and camera re-localization in the scene [Kendall et al.(2015)Kendall, Grimes, and Cipolla, Kendall and Cipolla(2017)].
In this work, we propose to merge the two lines of work described above. We cast pose estimation as a prediction problem, similar to deep learning methods that do not explicitly leverage viewpoint information. However, we condition our network on the 3D model of a single instance, represented either by a set of views or a point cloud, allowing our network to rely on the exact 3D model, similarly to the feature and template matching methods. To the best of our knowledge, we are the first to combine image and shape information as input to a network to estimate the relative orientation of the depicted object with respect to the shape.
Network Architecture and Training
Our approach consists in extracting deep features from both the image and the shape, and using them jointly to estimate a relative orientation. An overview is shown in Fig. 2. In this section, we present in more details our architecture, our loss function and our training strategy, as well as a data augmentation scheme specifically designed for our approach.
The first part of the network consists of two independent modules: (i) image feature extraction and (ii) 3D shape feature extraction. For image features, we use a standard CNN, namely ResNet-18 [He et al.(2016)He, Zhang, Ren, and Sun]. For 3D shape features, we experimented with two approaches depicted in Fig. 2(b) which are state-of-the-art 3D shape description networks.
First, we used the point set embedding network PointNet [Qi et al.(2017a)Qi, Su, Mo, and Guibas], which has been successfully used as a point cloud encoder for many tasks [Engelmann et al.(2017)Engelmann, Kontogianni, Hermans, and Leibe, Groueix et al.(2018)Groueix, Fisher, Kim, Russell, and Aubry, Qi et al.(2018)Qi, Liu, Wu, Su, and Guibas, Wang et al.(2018)Wang, Sun, Liu, Sarma, Bronstein, and Solomon, Xu et al.(2018)Xu, Anguelov, and Jain].
Second, we tried to represent the shape using rendered views, similar to [Su et al.(2015a)Su, Maji, Kalogerakis, and Learned-Miller]. Virtual cameras are placed around the 3D shape, pointing towards the centroid of the model; the associated rendered images are taken as input by CNNs, sharing weights for all viewpoints, which extract image descriptors; a global feature vector is obtained by concatenation. We considered variants of this architecture using extra input channels for depth and/or surface normal orientation but this did not improve our results significantly. Ideally, we would consider viewpoints on the whole sphere around the object with any orientation. In practice however, many objects have a strong bias regarding verticality and are generally seen only from the side/top. In our experiments, we thus only considered viewpoints on the top hemisphere and sampled evenly a fixed number of azimuths and elevations.
Orientation estimation.
The object orientation is estimated from both the image and 3D shape features by a multi-layer perceptron (MLP) with three hidden layers of size 800-400-200. Each fully connected layer is followed by a batch normalization, and a ReLU activation.
As output, we estimate the three Euler angles of the camera, azimuth (), elevation () and in-plane rotation (), with respect to the shape reference frame. Each of these angles is estimated using a mixed classification-and-regression approach, which computes both angular bin classification scores and offset information within each bin. Concretely, we split each angle uniformly in bins. For each -bin , the network outputs a probability using a softmax non-linearity on the -bin classification scores, and an offset relatively to the center of -bin , obtained by a hyperbolic tangent non-linearity. The network thus has outputs.
Loss function.
As we combine classification and regression, our network has two types of outputs (probabilities and offsets), that are combined into a single loss that is the sum of a cross-entropy loss for classification and Huber loss [Huber(1992)] for regression .
More formally, we assume we are given training data consisting of input images , associated object shapes and corresponding orientations . We convert the value of the Euler angles into a bin label encoded as a one-hot vector and relative offsets within the bins. The network parameters are learned by minimizing:
where are the probabilities predicted by the network for angle , input image and input shape , and the predicted offset within the ground truth bin.
Data augmentation.
We perform standard data augmentation on the input images: horizontal flip, 2D bounding box jittering, color jittering.
In addition, we introduce a new data augmentation, specific to our approach, designed to avoid the network to overfit the 3D model orientation, which is usually consistent in training data since most models are aligned. On the contrary, we want our network to be category-agnostic and to always predict the pose of the object with respect to the reference 3D model. We thus add random rotations to the input shapes, and modify the orientation labels accordingly. In our experiments, we restrict our rotations to azimuth changes, again because of the strong verticality bias in the benchmarks, but could theoretically apply it to all angles. Because of objects with symmetries, typically at or , we also restrict azimuthal randomization to a uniform sampling in , which allows to keep the bias of the annotations. See supplementary material for details and parameter study.
Implementation details.
For all our experiments, we set the batch size as 16 and trained our network using the Adam optimizer [Kingma and Ba(2014)] with a learning rate of for 100 epochs then for an additional 100 epochs. Compared to a shape-less baseline method, the training of our method with the shape encoded from 12 rendered views is about 8 times slower, on a TITAN X GPU.
Experiments
Given an RGB image of an object and a 3D model of that object, our method estimates its 3D orientation in the image. In this section, we first give an overview of the datasets we used, and explain our baseline methods. We then evaluate our method in two test scenarios: object belonging to a category known at training time, or unknown.
We experimented with four main datasets. Pascal3D+ [Xiang et al.(2014)Xiang, Mottaghi, and Savarese], ObjectNet3D [Xiang et al.(2016)Xiang, Kim, Chen, Ji, Choy, Su, Mottaghi, Guibas, and Savarese] and Pix3D [Sun et al.(2018)Sun, Wu, Zhang, Zhang, Zhang, Xue, Tenenbaum, and Freeman] feature various objects in various environments, allowing benchmarks for object pose estimation in the wild. On the contrary, LINEMOD [Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab] focuses on few objects with little environment variations, targeting robotic manipulation. Pascal3D+ and ObjectNet3D only provide approximate models and rough alignments while Pix3D and LINEMOD offer exact models and pixelwise alignments. We also used ShapeNetCore [Chang et al.(2015)Chang, Funkhouser, Guibas, Hanrahan, Huang, Li, Savarese, Savva, Song, Su, Xiao, Yi, and Yu] for training on synthetic data, with SUN397 backgrounds [Xiao et al.(2010)Xiao, Hays, Ehinger, Oliva, and Torralba], and tested on Pix3D and LINEMOD.
Unless otherwise stated, ground-truth bounding boxes are used in all experiments. We compute the most common metrics used with each dataset: is the percentage of estimations with rotation error less than ; MedErr is the median angular error (°); ADD-0.1d is the percentage of estimations for which the mean distance of the estimated 3D model points to the ground truth is smaller than 10% of the object diameter; ADD-S-0.1d is a variant of ADD-0.1d used for symmetric objects where the average is computed on the closest point distance. More details on the datasets and metrics are given in the supplementary material.
Baselines.
A natural baseline is to use the same architecture, data and training strategy as for our approach, but without using the 3D shape of the object. This is reported as ‘Baseline’ in our tables, and corresponds to the network of Fig. 2 without the shape encoder shown in light blue. We also report a second baseline, aiming at evaluating the importance of the precision of the 3D model for our approach to work. We used exactly our approach, but at testing time we replaced the 3D shape of the object in the test image by a random 3D shape of the same category. This is reported as ‘Ours (RS)’ in the tables.
1 Pose estimation on supervised categories
We first evaluate our method in case the categories of tested objects are covered by training data. We show that leveraging the 3D model of the object clearly improves pose estimation.
We evaluate our method on ObjectNet3D, which has the largest variety of object categories, 3D models and images. We report the results in Table 3 (top). First, an important result is that using the 3D model information, whether via a point cloud or rendered views, provides a very clear boost of the performance, which validates our approach. Second, results using rendered multiple views (MV) to represent the 3D model outperform the point-cloud-based (PC) representation [Qi et al.(2017a)Qi, Su, Mo, and Guibas]. We thus only evaluated Ours(MV) in the rest of this section. Third, testing the network with a random shape (RS) in the category instead of the ground truth shape, implicitly providing class information without providing fine-grained 3D information, leads to results better than the baseline but worst than using the ground truth model, demonstrating our method ability to exploit fine-grained 3D information. Finally, we found that even our baseline model already outperformed StarMap [Zhou et al.(2018)Zhou, Karpur, Luo, and Huang], mainly because of five categories (iron, knife, pen, rifle, slipper) on which StarMap completely fails, likely because a keypoint-based method is not adapted for small and narrow objects.
We then evaluate our approach on the standard Pascal3D+ dataset [Xiang et al.(2014)Xiang, Mottaghi, and Savarese]. Results are shown in Table 3 (top). Interestingly, while our baseline is far below state-of-the-art results, adding our shape analysis network provides again a very clear improvement, with results on par with the best category-specific approaches, and outperforming category agnostic methods. This is especially impressive considering the fact that the 3D models provided in Pascal3D+ are only extremely coarse approximations of the real 3D models. Again, as can be expected, using a random model from the same category provides intermediary results between the model-less baseline and using the actual 3D model.
Finally, we report results on Pix3D in Table 3 (top). Similar to the other methods, our model was purely trained on synthetic data and tested on real data, without any fine-tuning. Again, we can observe that adding 3D shape information brings a large performance boost, from to . Note that our method clearly improves even over category-specific baselines. We believe it is due to the much higher quality of the 3D models provided on Pix3D compared to ObjectNet3D and Pascal3D+. This hypothesis is supported by the fact that our results are much worse when a random model of the same category is provided.
These state-of-the-art results on the three standard datasets are thus consistent and validate (i) that using the 3D models provides a clear improvement (comparison to ‘Baseline’), and (ii) that our approach is able to leverage the fine-grained 3D information from the 3D model (comparison to estimating with a random shape ‘RS’ in the category).
2 Pose estimation on novel categories
We now focus on the generalization to unseen categories, which is the main focus of our method. We first discuss results on ObjectNet3D and Pix3D. We then show qualitative results on ImageNet horses images and quantitative results on the very different LINEMOD dataset.
Our results when testing on new categories from ObjectNet3D are shown in Table 3 (bottom). We use the same split between 80 training and 20 testing categories as [Zhou et al.(2018)Zhou, Karpur, Luo, and Huang]. As expected, the accuracy decreases for all methods when supervision is not provided on these latter categories. The fact that the Baseline performances are still much better than chance is accounted by the presence of similar categories is the training set. The advantage of our method is however even more pronounced than in the supervised case, and our multi-view approach (MV) still outperforms the point cloud (PC) approach by a small margin. Similarly, we removed from our ShapeNet [Chang et al.(2015)Chang, Funkhouser, Guibas, Hanrahan, Huang, Li, Savarese, Savva, Song, Su, Xiao, Yi, and Yu] synthetic training set the categories present in Pix3D, and reported in Table 3 (bottom) the results on Pix3D. Again, the accuracy drops for all methods, but the benefit from using the ground-truth 3D model increases.
In both ObjectNet and Pix3D experiments, the test categories were novel but still similar to the training ones. We now focus on evaluating our network, trained using synthetic images generated from man-made shapes from ShapeNetCore [Chang et al.(2015)Chang, Funkhouser, Guibas, Hanrahan, Huang, Li, Savarese, Savva, Song, Su, Xiao, Yi, and Yu], on completely different objects.
We first obtain qualitative results by using a fixed 3D model of horse from an online model repository [Free3D()] to estimate the pose of horses in ImageNet images. Indeed, compared to other animals, horses have more limited deformations. While this of course does not work for all images, the images for which the network provides the highest confidence are impressively good. On Figure 3, we show the most confident images for different poses, and we provide more results in the supplementary material. Note the very strong appearance gap between the rendered 3D models and the test images.
Finally, to further validate our network generalization ability, we evaluate it on the texture-less objects of LINEMOD [Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab], as reported in Table 4. This dataset focuses on very accurate alignment, and most approaches propose to first estimate a coarse alignment and then to refine it with a specific method. Our method provides a coarse alignment, and we complement it using the recent DeepIM [Li et al.(2018b)Li, Wang, Ji, Xiang, and Fox] refinement approach. Our method yields results below the state of the art, but they are nevertheless very impressive. Indeed, our network has never seen objects any similar the LINEMOD 3D models during training, while all the other baselines have been trained specifically for each object instance on real training images, except SSD-6D [Kehl et al.(2017)Kehl, Manhardt, Tombari, Ilic, and Navab] which uses the exact 3D model but no real image and for which coarse alignment performances are very low. Our method is thus very different from all the baselines in that it does not assume the test object to be available at training time, which we think is a much more realistic scenario for robotics applications. We actually believe that the fact our method provides a reasonable accuracy on this benchmark is a very strong result.
Conclusion
We have presented a new paradigm for deep pose estimation, taking the 3D object model as an input to the network. We demonstrated the benefits of this approach in terms of accuracy, and improved the state of the art on several standard pose estimation datasets. More importantly, we have shown that our approach holds the promise of a completely generic deep learning method for pose estimation, independent of the object category and training data, by showing encouraging results on the LINEMOD dataset without any specific training, and despite the domain gap between synthetic training data and real images for testing.
References
Supplementary Material
[Xiang et al.(2014)Xiang, Mottaghi, and Savarese] provides images with 3D annotations for 12 object categories. The images are selected from the training and validation set of PASCAL VOC 2012 [pascal-voc-2012] and ImageNet [Deng et al.(2009)Deng, Dong, Socher, Li, Li, and Fei-Fei], with 2k to 4k images in the wild per category. An approximate 3D CAD model is provided for each object as well as its 3D orientation in the image. Following the protocol of [Tulsiani and Malik(2015), Mousavian et al.(2017)Mousavian, Anguelov, Flynn, and Kosecka, Grabner et al.(2018)Grabner, Roth, and Lepetit], we use the ImageNet-trainval and Pascal-train images as training data, and the 2,113 non-occluded and non-truncated objects of the Pascal-val images as testing data. As in [Tulsiani and Malik(2015)], we use the metric , which measures the percentage of test samples having a pose prediction error smaller than : .
ObjectNet3D
[Xiang et al.(2016)Xiang, Kim, Chen, Ji, Choy, Su, Mottaghi, Guibas, and Savarese] is a large-scale 3D dataset similar to Pascal3D+ but with 100 categories, which provide a wider variety of shapes. To verify the generalization power of our method for unknown categories, we follow the protocol of StarMap [Zhou et al.(2018)Zhou, Karpur, Luo, and Huang]: we evenly hold out 20 categories (every 5 categories sorted in the alphabetical order) from the training data and only used them for testing. For a fair comparison, we actually use the same subset of training data as in [Xiang et al.(2016)Xiang, Kim, Chen, Ji, Choy, Su, Mottaghi, Guibas, and Savarese] (also containing keypoint annotations) and evaluate on the non-occluded and non-truncated images of the 20 categories, using the same metric.
Pix3D
[Sun et al.(2018)Sun, Wu, Zhang, Zhang, Zhang, Xue, Tenenbaum, and Freeman] is a recent dataset containing 5,711 non-occluded and non-truncated images of 395 CAD shapes among 9 categories. It mainly features furniture, with a strong bias towards chairs. But contrary to Pascal3D+ and ObjectNet3D, that only feature approximate models and rough alignments, Pix3D provides exact models and pixel-level accurate poses. Similar to the training paradigm of [Su et al.(2015b)Su, Qi, Li, and Guibas, Sun et al.(2018)Sun, Wu, Zhang, Zhang, Zhang, Xue, Tenenbaum, and Freeman], we train on ShapeNetCore [Chang et al.(2015)Chang, Funkhouser, Guibas, Hanrahan, Huang, Li, Savarese, Savva, Song, Su, Xiao, Yi, and Yu] with input images made of rendered views on random SUN397 backgrounds [Xiao et al.(2010)Xiao, Hays, Ehinger, Oliva, and Torralba] using random texture maps included in ShapeNetCore, and test on Pix3D real images and shapes.
ShapeNetCore
is a subset of ShapeNet [Chang et al.(2015)Chang, Funkhouser, Guibas, Hanrahan, Huang, Li, Savarese, Savva, Song, Su, Xiao, Yi, and Yu] containing 51k single clean 3D models, covering 55 common object categories of man-made artifacts. We exclude the categories containing mostly objects with rotational symmetry or small and narrow objects, which results in 30 remaining categories: airplane, bag, bathtub, bed, birdhouse, bookshelf, bus, cabinet, camera, car, chair, clock, dishwasher, display, faucet, lamp, laptop, speaker, mailbox, microwave, motorcycle, piano, pistol, printer, rifle, sofa, table, train, watercraft and washer. We randomly choose 200 models from each category and use Blender to render each model under 20 random views with various textures included in ShapeNetCore.
LINEMOD
[Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab] has become a standard benchmark for 6D pose estimation of textureless objects in cluttered scenes. It consists of 15 sequences featuring one object instance for each sequence to detect with ground truth 6D pose and object class. As other authors, we left out categories bowl and cup, that have a rotational symmetry, and consider only 13 classes. The common evaluation measure with LINEMOD is the ADD-0.1d metric [Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab]: a pose is considered correct if the average of the 3D distances between transformed object vertices by the ground truth transformation and the ones by estimated transformation is less than 10% of the object’s diameter. For the objects with ambiguous poses due to symmetries, [Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab] replaces this measure by ADD-S which is specially tailored for symmetric objects. We choose ADD-0.1d and ADD-S-0.1d as our evaluation metrics.
2 Evaluation Metrics
For results on LINEMOD, the ADD [Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab] metric is used to compute the averaged distance between points transformed using the estimated pose and the ground truth pose:
where is the number of points on the 3D object model, is the set of all 3D points of this model, is the ground truth pose and is the estimated pose. Following [Brachmann et al.(2016)Brachmann, Michel, Krull, Yang, Gumhold, and Rother], we compute the model diameter as the maximum distance between all pairs of points from the model. With this metric, a pose estimation is considered to be correct if the computed averaged distance is within 10% of the model diameter .
For the objects with ambiguous poses due to symmetries, [Hinterstoisser et al.(2012b)Hinterstoisser, Lepetit, Ilic, Holzer, Bradski, Konolige, and Navab] replaces this measure by ADD-S, which uses the closet point distance in computing the average distance for 6D pose evaluation as in:
3 Ablation and parameter study
Table 5 shows the experimental results of pose estimation on 20 novel categories of ObjectNet3D for different numbers and layouts of rendered images. The viewpoints are sampled evenly at azimuths and elevated at different elevations. represents respectively elevations at , , . The metric measures the percentage of testing samples with a angular error smaller than and MedErr is the median angular error (°) over all testing samples.
The table shows that using shape information encoded from rendered images (when ) can indeed help pose estimation on novel categories, i.e., that are not included in the training data. In the first column (0 rendered images) we show the performance of our baseline without using the 3D shape of the object, compared to this result, the network trained with only one rendered image has a clearly boosted accuracy.
The table also shows that more rendered images in the network input does not necessarily mean a better performance. In the table, the network trained with 12 rendered images elevated at 0°and 30°gives the best result. This may be because the ObjectNet3D dataset is highly biased towards low elevations on the hemisphere, which can be well represented without using the rendered image captured at high elevation such as 60°.
Parameter study on the azimuthal randomization strategy.
Table 6 summarizes the parameter study on the range of azimuthal jittering applied to input shapes during network training. The poor results obtained for and are due the objects with symmetries, typically at 90°or 180°.
4 Qualitative Results on LINEMOD
Some qualitative results for 13 LINEMOD objects are shown in Figure 4. Given object image and its shape, our approach gives a coarse pose estimate which is then refined by pose refinement method given by DeepIM [Li et al.(2018b)Li, Wang, Ji, Xiang, and Fox].