On Object Symmetries and 6D Pose Estimation from Images

Giorgia Pitteri, Michaël Ramamonjisoa, Slobodan Ilic, Vincent Lepetit

Introduction

3D object detection and pose estimation are of primary importance for tasks such as robotic manipulation, virtual and augmented reality and they have been the focus of intense research in recent years, mostly due to the advent of Deep Learning based approaches and the possibility of using large datasets for training such methods .

However, one challenge is often ignored in recent works. Many objects of our daily life or from industrial contexts exhibit symmetries, or at least ’quasi-symmetries’ when only a small detail prevents the object to have a perfect symmetry. These symmetries create ambiguities when aiming to estimate the 6D pose of the object from images, however only a few recent papers have considered the problems raised by object symmetries . In this paper, we first explain why exactly symmetries can be a problem for 6D pose estimation algorithms. We then provide a simple solution that is general and can be introduced in any 6D pose estimation algorithms.

To better understand the problem raised by the symmetries of an object, let’s first consider Fig. 1. The blue object has a rotational symmetry around the vertical axis: If we apply a rotation of 180°180\degree around this axis, this object has exactly the same appearance. More generally, when an object OO has some symmetry, there exist one or more rigid motions such that, if we apply these rigid motions to the object pose, the appearance of the object is preserved. Formally, we consider the set

where R(O,p)\mathcal{R}(O,\boldsymbol{p}) is the image of Object OO under pose p\boldsymbol{p} (ignoring lighting effects), m\boldsymbol{m} is a rigid motion related to the symmetry, and m.p\boldsymbol{m}.\boldsymbol{p} is the pose after applying motion m\boldsymbol{m}. M(O)\mathcal{M}(O) is thus the set of rigid motions m\boldsymbol{m} that preserve the visual aspect of a given object. It is easy to see that it forms a subgroup of SE(3)SE(3). calls the elements of M(O)\mathcal{M}(O) proper symmetries.

In other words, two images of a symmetrical object can be identical but not correspond to the same pose. If we consider an image I1=R(O,p)\boldsymbol{I}_{1}=\mathcal{R}(O,\boldsymbol{p}) of an object OO under pose PP and a motion m∈M(O)\boldsymbol{m}\in\mathcal{M}(O), then, the image I2\boldsymbol{I}_{2} of object OO under pose m.p\boldsymbol{m}.\boldsymbol{p} is equal to image I1\boldsymbol{I}_{1}, i.e. I2=R(O,m.p)=R(O,p)=I1\boldsymbol{I}_{2}=\mathcal{R}(O,\boldsymbol{m}.\boldsymbol{p})=\mathcal{R}(O,\boldsymbol{p})=\boldsymbol{I}_{1}. There is therefore no function

that can provide the pose p\boldsymbol{p} of object OO given an image I\boldsymbol{I}. Any attempt to learn such a function, for example with a Deep Network, would fail. For example, if a network is trained to predict the pose using the squared loss between the ground truth poses and the predicted poses, it would converge to a model predicting the average of the possible poses for an input image, which is of course meaningless.

Only few works consider the problem of symmetrical objects: Sundermeyer et al. solves this problem by learning a mapping to a latent representation of the pose; Bregier et al. introduced a representation of the pose that differs from rigid motions and suitable for their similarity metric between two poses; learns to predict several poses so that at least one pose corresponds to the ground truth; Rad and Lepetit rely on image mirroring to deal with some symmetries. While these papers propose interesting solutions, here, we consider a general analytic approach to the problem. It will give insights on the learning-based methods, and yields a simple method to solve the ambiguities due to symmetries.

In the remainder of the paper, we review the state-of-the-art on 3D object pose estimation from images, describe our method, and evaluate it on the T-Less dataset, which is made of very challenging objects and sequences.

Related Work

6-DoF pose estimation made significant progress recently. We discuss below mostly the most recent ones, and several techniques that have been proposed to specifically tackle objects with pose ambiguities.

Several recent works extend on deep architectures developed for 2D object detection by also predicting the 3D pose of objects. trained the SSD architecture to also predict the 3D rotations of the objects, and the depths of the objects. Deep-6DPose relies on Mask-RCNN instead of SSD. To improve robustness to partial occlusions, PoseCNN segments the objects’ masks and predicts the objects’ poses in the form of a 3D translation and a 3D rotation. Yolo6D relies on Yolo and predicts the object poses in the form of the 2D projections of the corners of the 3D bounding boxes, a 3D pose representation introduced in and .

Several works also attempted to be more robust to occlusions. first predict the 3D coordinates of the image locations lying on the objects, in the object coordinate system, and predict the 3D object pose through hypotheses sampling with preemptive RANSAC. predict the 2D projections of 3D points from image patches or local features, to avoid the effects of occluders when performing the prediction.

These works have been very successful at predicting the 3D pose of objects, however they mostly do not consider objects with symmetry. Our goal in this paper is not to propose another architecture for 3D pose prediction, but to study the effects of symmetries on the prediction process, and propose a general solution, which can be integrated in these previous works.

2 Ambiguity Aware Pose Estimation

is probably the first work that mentioned the difficulty of predicting the 3D pose of objects with symmetries using Deep Networks, and presents some results on the T-Less dataset. However, the paper does not provide many details about the method and the solution is not general. Our approach is related to the direction they point at, but we provide a general solution, with much more justifications.

learns a latent representation of the object pose using an auto-encoder. They show that their learned embedding is ambiguity agnostic, in the sense that visually ambiguous poses will map to the same code in the latent space. They perform pose estimation by matching the code obtained from an image of the object with a precomputed code table covering the 6D pose space. While this approach is very interesting, we consider here an orthogonal approach based on an analytical study of the ambiguities. Moreover, the code table introduces some discretization, while we predict a 3D pose that varies continuously with the input image.

learns to compare an input image with a set of renderings of the object under many views, to predict the most similar view and to predict the rotational symmetries of the object. This also requires to discretize the possible rotations, while we predict a continuous 3D pose.

also considers a learning-based approach, tackles ambiguities raised by partial occlusions in addition to rotational symmetries, i.e. when an occluder hides a part of an object, so that it is not possible to estimate the pose exactly anymore. This is done by training a network to predict multiple poses, so that only one has to correspond to the actual pose. At test time, the network predicts multiple poses, which are expected to represent the distribution over the possible poses. By contrast with this learning-based approach, we explicitly consider the ambiguities that can raise under symmetries.

introduced the concept of proper symmetries group in a survey that aims to cover ambiguities and a pose representation specific to a metric on 3D poses. We use this concept to solve the ambiguities created by symmetrical objects. The paper however does not consider pose prediction using regression or machine learning.

notices that symmetries produce multiple modes in the distribution Q(θ∣I)Q(\theta|\boldsymbol{I}) over 3D poses θ\theta. They therefore enforce a uniform prior P(θ)P(\theta) over symmetrical poses to successfully approximate QQ. However, they do not explicitly report results on (quasi)-symmetrical objects such as those of T-Less.

Method

We study below the effect of symmetries on algorithms aiming to learn the mapping between an image of an object and its 6D pose, and we show how we can derive a simple method for handling these symmetries. In the next section, we describe how this method can be integrated within a Faster-RCNN framework.

Let’s consider the set M(O)\mathcal{M}(O) already introduced in Eq. (1). In practice, the motions in M\mathcal{M} are usually in the form m=[R,0]\boldsymbol{m}=[R,\boldsymbol{0}] with R∈SO(3)R\in\text{SO}(3), i.e. objects have mostly rotational symmetries. A translation component different from 0\boldsymbol{0} would correspond to an object with translation symmetries, for example a long building with windows of similar appearances.

We thus first define the notion of ambiguous rotations: We say that two rotations R1R_{1} and R2R_{2} are ambiguous if they result in the same object appearance, i.e. if R(O,[R1,T1])=R(O,[R2,T2])\mathcal{R}(O,[R_{1},T_{1}])=\mathcal{R}(O,[R_{2},T_{2}]). This defines an equivalence relationship R1∼R2R_{1}\sim R_{2}. If R1∼R2R_{1}\sim R_{2}, then it is not possible from an image to distinguish between rotation R1R_{1} and R2R_{2} when predicting the pose. Predicting R1R_{1}, or R2R_{2}, or any rotation R∼R1R\sim R_{1} is equally good. This is in fact the idea behind the ADI metric .

As illustrated in Fig. 2, a natural idea to aim at preventing trouble during learning is therefore to first map equivalent rotations to a unique rotation, which we call a canonical rotation. This means that during training, training images with the same object appearance will be assigned the same rotation after mapping. The transformation F:I↦p\mathcal{F}:\boldsymbol{I}\mapsto\boldsymbol{p} of Eq. 2 will thus become a function and we will be able to learn it with a Deep Network for example. This implies that at inference, the network will predict the canonical rotation for a given input image, which is the best that can be done in presence of symmetries.

Given set M(O)\mathcal{M}(O) of the object’s proper symmetries, we are therefore looking for an operator Map(⋅)\text{Map}(\cdot) on SO(3)\text{SO}(3) that can map ambiguous 3D rotations to a single rotation such that Map(R1)=Map(R2)  ⟺  R1∼R2\text{Map}(R_{1})=\text{Map}(R_{2})\iff R_{1}\sim R_{2} (⋆\star) holds.

Given a proper symmetry group M(O)\mathcal{M}(O), let us define Map operator as:

where ∥.∥F\|.\|_{F} is the Froebenius norm. Then Map verifies the mapping property (⋆\star).

To simplify the notations, let us consider that M(O)\mathcal{M}(O) is made only of the rotation components. By definition of R1∼R2R_{1}\sim R_{2} and M(O)\mathcal{M}(O):

Let us consider the solution of the optimization problem in Eq. (4) for R1R_{1}:

We introduce variable TT such that S=S12TS=S_{12}T. Since SS and S12S_{12} belong to M(O)\mathcal{M}(O) and M(O)\mathcal{M}(O) is a group, TT also belongs to M(O)\mathcal{M}(O). We can therefore perform the following change of variable:

2 Implementing Map

If M\mathcal{M} is discrete, implementing operator Map is trivial, as it is only a matter of iterating over the elements of M\mathcal{M} to find the minimum. However, M\mathcal{M} can be continuous for some objects. This is the case for generalized cylinders and spheres . For spheres, Map is also trivial as it can always return the identity transformation, for example.

For generalized cylinders, implementing operator Map is more complex. In this case, M\mathcal{M} can be written as:

where RαuR^{\boldsymbol{u}}_{\alpha} is the rotation around axis u{\boldsymbol{u}} of amount α\alpha.

The Froebenius norm in Eq. (4) can be rewritten as

with D=S−1R−I3D=S^{-1}R-I_{3}. After some derivations:

The complete derivations can be found in the supplementary material.

Without loss of generality, we can assume that u{\boldsymbol{u}} is the z{\boldsymbol{z}}-axis of the object’s coordinate system: If it is not the case, the following still holds after applying a change of basis. Then, SS has the following form:

To implement Map, we need to solve the optimization problem S^=arg min⁡S∈M(O)Trace(DTD)\hat{S}=\operatorname*{arg\,min}\limits_{S\in\mathcal{M}(O)}\text{Trace}(D^{T}D), which can now be rewritten as a minimization over α\alpha:

This is solved analytically by solving ∂Trace(DTD)∂α=0\frac{\partial\text{Trace}(D^{T}D)}{\partial\alpha}=0 for α\alpha. The solution of Eq. (4) is then:

3 Discontinuities of ℱℱ\mathcal{F} After Mapping

After applying the Map operator, there are no pose ambiguities anymore, i.e. two similar images are assigned the same rotation. However, a new difficulty arises: The transformation F(I)→p\mathcal{F}(\boldsymbol{I})\rightarrow\boldsymbol{p} is now discontinuous around some rotations. This is problematic when using Deep Networks to learn F\mathcal{F}, as Deep Networks can only approximate continuous functions .

To understand why these discontinuities happen, let us consider an example, more exactly the rectangular object seen from the top as in Fig. 3. M(O)\mathcal{M}(O) is made of two rotations around the ZZ axis: The identity matrix, and the rotation of angle π\pi, and M(O)={I3,Rπu}\mathcal{M}(O)=\{I_{3},R^{\boldsymbol{u}}_{\pi}\}. If a training image is annotated with rotation Rπ/2+ϵzR_{\pi/2+\epsilon}^{\boldsymbol{z}}, this rotation will be mapped by operator Map to rotation Rϵ−π/2zR_{\epsilon-\pi/2}^{\boldsymbol{z}}; If a training image is annotated with pose Rπ/2−ϵzR_{\pi/2-\epsilon}^{\boldsymbol{z}}, this rotation will be mapped to itself i.e. Rπ/2−ϵzR_{\pi/2-\epsilon}^{\boldsymbol{z}}. By making ϵ\epsilon converge to 0, it can be seen that there is a discontinuity of F\mathcal{F} around images annotated with rotations π\pi before mapping.

Another way of looking at the problem is to notice that images of the object annotated with rotations Rϵ−π/2zR_{\epsilon-\pi/2}^{\boldsymbol{z}} and Rπ/2−ϵzR_{\pi/2-\epsilon}^{\boldsymbol{z}} look very similar, but with very different rotations. A Deep Network would have to learn to predict very different poses for very similar images.

4 Solving the Discontinuities

The discontinuities only occur when M\mathcal{M} is discrete: It can be seen from Eq. (18) that in the case of a generalized cylinder, the Map operator is continuous. Otherwise, we avoid these discontinuities by introducing a partition of SO(3)\text{SO}(3) made of two subsets Ω1\Omega_{1} and Ω2\Omega_{2}. For each subset, we train a different regressor to predict the pose. We will therefore have two regressors F1\mathcal{F}_{1} and F2\mathcal{F}_{2} instead of only one. In this way, both F1\mathcal{F}_{1} and F2\mathcal{F}_{2} will be continuous over their respective domains.

We describe below our method on an example, and then extend it to the general case.

Let us consider again the rectangular object pictured in Fig. 3, and already discussed in Section 3.3. For this object, we have M(O)={I3,Rπu}\mathcal{M}(O)=\{I_{3},R^{\boldsymbol{u}}_{\pi}\}. We can notice that M\mathcal{M} and Map generate a partition of SO(3)\text{SO}(3) made of two subsets:

where S^(R)\hat{S}(R) is the rotation of Eq. (4) when applying Map to RR.

However, this partition will not solve our problem: We already know that F\mathcal{F} is not continuous on Ω1\Omega_{1}. We must therefore introduce a new partition of SO(3)\text{SO}(3). For this partition, we consider the new set:

As shown in Fig. 4(b), no part Ω(k)\Omega^{(k)} include any discontinuity. Moreover, for a rotation in Ω(2)\Omega^{(2)}, there is another rotation in Ω(0)\Omega^{(0)} that generates the same object appearance. The same yields for Ω(3)\Omega^{(3)} and Ω(1)\Omega^{(1)}.

We therefore take Ω1=Ω(0)\Omega_{1}=\Omega^{(0)} for the domain of regressor F1\mathcal{F}_{1}, and Ω2=Ω(1)\Omega_{2}=\Omega^{(1)} for the domain of regressor F2\mathcal{F}_{2}. F1\mathcal{F}_{1} and F2\mathcal{F}_{2} thus do not suffer from discontinuities nor ambiguity. They are sufficient to estimate the object pose under any rotation, since we can map this rotation to a rotation either in Ω1\Omega_{1} or Ω2\Omega_{2} corresponding to the same appearance. To do so, we introduce a new mapping Map′\text{Map}^{\prime} derived from Map such that:

During training, given a training image I\boldsymbol{I} annotated with rotation RR, we compute (S^−1R,δ)←Map′(R)(\hat{S}^{-1}R,\delta)\leftarrow\text{Map}^{\prime}(R) and train regressor Fδ\mathcal{F}_{\delta} to predict rotation S^−1R\hat{S}^{-1}R from I\boldsymbol{I}.

During inference, given a test image I\boldsymbol{I} of an object, we need to know which regressor we should invoke to predict the pose. To do so, during training, we train a classifier C\mathcal{C} to predict which regressor we should invoke to compute the pose, that is we train C\mathcal{C} to predict δ\delta from I\boldsymbol{I}. For rotations close to the boundary between Ω1\Omega_{1} and Ω2\Omega_{2}, the prediction for C\mathcal{C} can become ambiguous. However, in this case, the ambiguity is not a problem in practice: Even if the classifier predicts the wrong regressor to use close to the boundary between Ω1\Omega_{1} and Ω2\Omega_{2}, both regressors can correctly predict poses close to this boundary.

4.2 One Symmetry Axis, Arbitrary M𝑀M

Let us now generalize to an object OO with an arbitrary amount of symmetries around a single axis uu. These symmetries are necessarily periodic around u{\boldsymbol{u}} with angular period fα=2π/Mf_{\alpha}=2\pi/M: Rotating OO around uu by any angle multiple of fαf_{\alpha} does not change its appearance. The proper symmetry group M(O)\mathcal{M}(O) for such an object is:

M(O)\sqrt{\mathcal{M}}(O) of Eq. (20)) becomes:

and mapping Map′\text{Map}^{\prime} of Eq. (22) becomes:

where Ω1={R:S^(R)=I3}\Omega_{1}=\{R:\hat{S}(R)=I_{3}\}.

We can use Map′\text{Map}^{\prime} the same way as in the previous subsection to train and use to regressors F1\mathcal{F}_{1} and F2\mathcal{F}_{2}.

4.3 General Case

In the general case, each rotation RR in M\mathcal{M} can be written in the form:

where u{\boldsymbol{u}}, v{\boldsymbol{v}}, etc. are rotation axes. Most common objects have at most 2 axes of symmetries, but it is possible to imagine objects with more, for example a golf ball. To keep the notations as simple as possible, we will stick to only two axes, as it is easy to extend to more axes from there.

and mapping Map′\text{Map}^{\prime} becomes:

where Ω1,1={R:S^(R)=I3}\Omega_{1,1}=\{R:\hat{S}(R)=I_{3}\}, Ω2,1={R:S^(R)=Rπ/Mu}\Omega_{2,1}=\{R:\hat{S}(R)=R^{\boldsymbol{u}}_{\pi/M}\}, and Ω1,2={R:S^(R)=Rπ/Nv}\Omega_{1,2}=\{R:\hat{S}(R)=R^{\boldsymbol{v}}_{\pi/N}\}. It means that in this case, we have to train 4 different regressors F1,1\mathcal{F}_{1,1}, F2,1\mathcal{F}_{2,1}, F1,2\mathcal{F}_{1,2}, and F2,2\mathcal{F}_{2,2} according to δ1\delta_{1} and δ2\delta_{2}, and the classifier C\mathcal{C} to predict a class index in [0;3][0;3].

5 Method Summary

The method developed above can be summarized as follow. We distinguish between generalized cylinders and objects with discrete symmetries.

If the object is a generalized cylinder, given a training image I\boldsymbol{I} annotated with rotation RR, we train a single regressor F\mathcal{F} to predict Map(R)\text{Map}(R) using Eq. 3 from I\boldsymbol{I}. At inference time, given a test image I\boldsymbol{I}, we simply have to invoke F\mathcal{F} to predict the object pose from I\boldsymbol{I}.

If the object has discrete symmetries, given a training image I\boldsymbol{I} annotated with rotation RR, we apply Map′\text{Map}^{\prime} to RR using Eq. (25) or Eq. (28) depending on the number of symmetry axes. Map′\text{Map}^{\prime} provides the rotation to be associated with I\boldsymbol{I} for training, as well as the index of the regressor Fi\mathcal{F}_{i} to train. In addition to training the regressors, we need to train classifier C\mathcal{C} to predict the index of the regressor to use. At inference time, we first invoke classifier C\mathcal{C} to predict which regressor we should use from I\boldsymbol{I}, and then, invoke this regressor to predict the object pose from I\boldsymbol{I}.

Integration into Faster-RCNN

We integrated our approach into Faster-RCNN . We keep the original architecture of to obtain region proposals and classify each of those regions with an object label: In the T-Less dataset , there exist 30 classes of object. We also keep the original loss terms for this part.

We chose to predict the objects’ 6D poses in the form of the 2D reprojections of the 8 corners of the 3D bounding boxes, as in for simplicity. From these 2D reprojections, it is possible to estimate a 6D pose using a PnnP algorithm . However, our approach is general, and using any other representation of the pose, with quaternions for example, is also possible.

Fig. 5 shows the different branches we added to the original Faster-RCNN architecture. We describe them below.

We add a specific branch to the Faster-RCNN architecture to predict the 2D coordinates of each 3D corner for each regressor. The output of each branch has size 16×3016\times 30, where 3030 is the number of object classes and 1616 accounts for the 8 2D coordinates to predict. This branch is implemented as a fully connected multi-layer perceptron and takes as input the output shared single channel feature-map. We then use an L1L_{1} or L2L_{2} loss on each coordinate. More details can be found in the supplementary materials.

We also added a specific multi-layer perceptron branch to Faster-RCNN to implement classifier C\mathcal{C}. Ground truth is obtained using Eq. (22).

Experiments

In this section, we detail how we evaluated our approach, and show its effectiveness on objects with various types of symmetries.

We use the objects of the T-Less dataset as they exhibit many different challenges due to symmetries, and are representative of objects in daily and industrial environments. However, the T-Less dataset does not provide many images for training, with a limited range of poses and illumination conditions. We therefore generated training and test images using the CAD models provided with T-Less, introducing partial occlusions and illumination variations. This dataset is made of 30K samples, generated using the CAD models provided in the original T-Less dataset with Cycles, a photorealistic rendering engine of the open source software Blender.

Each sample of our dataset is generated using a random set S\mathcal{S} of objects taken from the T-Less dataset, using random gray scale color (from dark-gray to white) for each of them. Each object of S\mathcal{S} is initially set with a random pose, and we let the objects fall down on a randomly textured plane, using Blender’s physics simulator. Because the objects can collide together, their final pose on the table is also random. Illumination randomization is performed by varying the level of ambient light and randomizing a point light source in terms of position, strength, and color. This often results in strong cast shadows, as can be seen in Fig. 6. A comparison with the original T-Less (primesense) dataset is given in Table 1. This dataset therefore provides challenging conditions, and allows us to focus on the challenges raised by the symmetries, without having to consider the domain gap between synthetic and real images.

2 Effectiveness of our Approach

As shown in Fig. 9, the loss of our Faster R-CNN -based implementation converges only when the rotations are normalized using our normalization procedure, indicating that something is incorrect in the loss function in absence of normalization. In Fig. 7, we show what happens in practice for three possible types of objects: Two generalized cylinders (objects 3030 and 33), an object with an axis of symmetry (object 2929), and an object without any symmetry (object 2626). When dealing with non-symmetrical objects, the network is able to learn the 6D pose with and without the normalization procedure. On the opposite, when the objects are symmetrical, without our normalization the network learns the average between all the possible poses ending up predicting a pose collapsed to the center of the object.

3 T-LESS Dataset: Comparison with [26]

We use the Visible Surface Discrepancy (VSD) error function introduced by . It compares the ground truth measured depth maps S^\hat{S} and the depth maps Sˉ\bar{S} rendered according to the estimated poses to evaluate the proportion of visible pixels for which the depth absolute discrepancy map ∣S^|\hat{S} - Sˉ∣\bar{S}| is below a threshold τ\tau. As in , we set τ=20\tau=20mm and report the recall of correct 6D object poses at evsd<0.3\mathbf{e_{vsd}<0.3}. This metric is not sensitive to visual symmetries, as they induce similar symmetries in depth maps.

Table 2 compares our method to the method of Sundermeyer et al . The object 3D orientation and translation along the x{\boldsymbol{x}}-and y{\boldsymbol{y}}-axes are typically well estimated. Although most of the translation error is along z{\boldsymbol{z}}-axis, it is unsurprising since we do not use or regress the depth information. In order to have a meaningful evaluation of our results in terms of VSD, we keep the ground truth of the translation along z{\boldsymbol{z}}-axis in our pose predictions.

Conclusion

In this paper, we studied the subtle problems that arise when training a machine learning method to predict the 6D pose of an object with symmetries. This leads to a simple method that is agnostic to the exact pose representation and the pose prediction model. Our method can therefore be included in current and future developments for properly handling objects with symmetries. A direct extension of our work could be to automatically detect the object symmetries.

References