Explaining the Ambiguity of Object Detection and 6D Pose From Visual Data

Fabian Manhardt, Diego Martin Arroyo, Christian Rupprecht, Benjamin Busam, Tolga Birdal, Nassir Navab, Federico Tombari

Introduction

Driven by deep learning, image-based object detection has recently made a tremendous leap forward in both accuracy as well as efficiency . An emerging research direction in this field is the estimation of the object’s pose in 3D space over the existing 6-Degrees-of-Freedom (DoF) rather than on the 2D image plane . This is motivated by a strong interest in achieving robust and accurate monocular 6D pose estimation for applications in the field of robotic grasping, scene understanding and augmented/mixed reality, where the use of a 3D sensor is not feasible .

Nevertheless, 6D pose estimation from RGB is a challenging problem due to the intrinsic ambiguity caused by visual appearance of objects under different viewpoints and occlusion. Indeed, most common objects exhibit shape ambiguities and repetitive patterns that cause their appearance to be very similar under different viewpoints, thus rendering pose estimation a problem with multiple correct solutions. Furthermore, also occlusion (from the same object or from others) can cause pose ambiguity.

For example, as illustrated in Figure 1, the cup is identical from every viewpoint in which the handle is not visible. Thus, from a single image, it is impossible to univocally estimate the current object pose. Moreover, object symmetry can also induce visual ambiguities leading to multiple poses with the same visual appearance. However, most datasets do not reflect this ambiguity, as the ground truth pose annotations are mostly uniquely defined at each frame. This is problematic for a proper optimization of the rotation, since a visually correct pose still results in a high loss. Thus, many recent 3D detectors avoid regressing the rotation directly and, instead, explicitly model the solution space in an unambiguous fashion .

Essentially, in , the authors train their convolutional neural network (CNN) by mapping all possible pose solutions for a certain viewpoint onto an unambiguous arc on the view sphere. Rad et al. employ a separate CNN solely trained to classify the symmetry in order to resolve these ambiguities. However, this simplification exhibits several downsides, such as the explicit inclusion of information about certain symmetries in each trained object. Moreover, this is not always easy to model, as e.g. in the case of partial view ambiguity. Further, all these approaches rely on prior knowledge and annotation of the object symmetries and aim to solve the ambiguity by providing a single outcome in terms of estimated pose and object. Added to this, these methods are also unable to deal with ambiguities generated by other common factors such as occlusion.

On the contrary, Sundermeyer et al. and Corona et al. recently proposed novel methods to conduct pose estimation in an ambiguity-free manner. In the core, both learn a feature embedding solely based on visual appearance. Nonetheless, although is able to deal with ambiguities implicitly, it does not model their detection and description explicitly. In contrast, also learns to classify the order of rotational symmetry, in particular the number of equivalent views around an axis of rotation. However, they require explicit hand-annotated labels and, in addition, cannot deal with ambiguities aside from these symmetry classes such as (self-) occlusion.

In this paper we propose to model the ambiguity of the object detection and pose estimation tasks directly by allowing our learned model to predict multiple solutions, or hypotheses, for a given object’s visual appearance (Fig 2). Inspired by Rupprecht et al. we propose a novel architecture and loss function for monocular 6D pose estimation by means of multiple predictions. Essentially, each predicted hypothesis itself corresponds to a 3D translation and rotation. When the visual appearance is ambiguous, the model predicts a point estimate of the distribution in 3D pose space. Conversely, when the object’s appearance is unique, the hypotheses will collapse into the same solution. Importantly, our model is capable of learning the distribution of these 6D hypotheses from one single ground truth pose per sample, without further supervision.

Besides providing more insight and a better explanation for the task at hand, the additional knowledge gained from rotation distributions can be exploited to improve the accuracy of the pose estimates. In essence, analyzing the distribution of the hypotheses enables us to classify if the current perceived viewpoint is ambiguous and to compute the axis of ambiguity for that specific object and viewpoint. Subsequently, when ambiguity is detected, we can employ mean shift clustering over the hypotheses in quaternion space to find the main modes for the current pose. A robust averaging in 3D rotation space for each mode then yields a highly accurate pose estimate. When the view is ambiguity-free, we can improve our pose estimates by robustly averaging over all 6D hypotheses, and by taking advantage of the predicted pose distribution as a confidence measure.

We propose a novel method for 6DoF pose estimation, which can deal with the inherent ambiguities in pose by means of multiple hypotheses.

Explicit detection of rotational ambiguities and characterization of the uncertainty in the problem without further annotation or supervision.

A mechanism to measure the reliability and to increase the robustness of the unambiguous 6D pose prediction.

Related Work

We first review recent work in object detection and pose estimation from 2D and 3D data. Afterwards, we discuss common grounds and main differences with approaches aimed at symmetry detection for 3D shapes.

Almost all current research focus on deep learning-based methods.

employ CNNs to learn an embedding space for the pose and class from RGB-D data, which can subsequently be utilized for retrieval. Notably, the majority of most recent deep learning based methods focus on RGB as input . Since utilizing pre-trained networks often accelerates convergence and leads to better local minima, these methods are usually grounded on state-of-the-art backbones for 2D object detection, such as Inception or ResNet . In particular, Kehl et al. employ SSD with an InceptionV4 backbone and extend it to also classify viewpoint and in-plane rotation. Similarly, Sundermeyer et al. also use SSD for localization, but employ an augmented auto-encoder for the unambiguous retrieval of the associated 6D pose. Rad et al. utilize VGG and augment it to provide the 2D projections of the 3D bounding box corners. A similar approach is chosen by , based on YOLO . Afterwards, both apply PnnP to fit the associated 3D bounding box into the regressed 2D projections, in order to estimate the 3D pose of the detection. In , Xiang et al. compute a shared feature embedding for subsequent object instance segmentation paired with pose estimation.

Finally, Do et al. extend Mask-RCNN with a third branch, which provides the 3D rotation and the distance to the camera for each prediction.

Oftentimes, object pose ambiguity arises from symmetric shapes. We review relevant methods that extract symmetry from 3D models to outline commonalities and differences with our approach.

To our knowledge, is the only method which estimates both: the 6D pose, and the symmetry of the perceived object. In particular, the network is trained to also predict the rotational order (i.e. the number of identical views), posing it as a classification task.

Generally, most methods for symmetry detection are found in the shape analysis community. Among the different kinds of symmetries, axial symmetries are of particular interest, and multiple approaches have been proposed. Most methods rely on feature matching or spectral analysis: treat the problem as a correspondence matching task between a series of keypoints on an object, determining the reflection symmetry hyperplane as an optimization problem. Elawady et al. rely on edge features extracted using a Log-Gabor filter in different scales and orientations coupled with a voting procedure on the computed histogram of local texture and color information. In addition, and are also grounded on wavelet-based approaches. Recently, neural network approaches have also been proposed. Ke et al. adapt an edge-detection architecture with multiple residual units and successfully apply it to symmetry detection using real-world images.

Notably, all these approaches aim at detecting symmetries of 3D shapes alone, while our focus is to model the ambiguity arising from objects under specific viewpoints with the goal of improving and explaining pose estimation.

Methodology

In this section we describe our method for handling symmetries and other ambiguities for object detection and pose estimation in detail. We will first define what we understand as an ambiguity.

Under ambiguities, a direct naive regression of the rotation as a quaternion will lead to poor results, as the network will learn to predict a rotation that is closest to all results in the symmetry group. This prediction can be seen as the (conditional) mean rotation. More formally, in a typical supervised setting we associate images IiI_{i} with poses pip_{i} in a dataset (Ii,pi)(I_{i},p_{i}) where i∈{1,…,N}i\in\{1,\ldots,N\}. To describe symmetries, we define for a given image IiI_{i}, the set S(Ii)\mathcal{S}(I_{i}) of poses pp that all have an identical image

Note that in the case of non-discrete symmetries the set S\mathcal{S} will contain infinitely many poses, which in turn transforms the sums of SS in the following to integrals. For the sake of a simpler notation and a finite training set in practice, we chose to continue with a notion of a finite ∣S∣|S|. The naive model f(I,θ)f(I,\theta), that directly regresses a pose p′p^{\prime} from II, optimizes a loss L(p,p′)\mathcal{L}(p,p^{\prime}) by minimizing

The key idea behind the proposed method is to model the ambiguity by allowing multiple pose predictions from the network. In order to predict MM pose hypotheses from ff, we extend the notation to fθ(I)=(fθ(1)(I),…,fθ(M)(I))f_{\theta}(I)=(f^{(1)}_{\theta}(I),\ldots,f^{(M)}_{\theta}(I)) where ff now returns MM pose hypotheses for each image II.

For training, the idea is not to punish all hypotheses given the current pose annotation, since they might be correct under ambiguities. Thus, we use a loss that optimizes only one of the MM hypotheses for each annotation. The most intuitive choice is to pick the closest one. We adapt the meta loss M\mathcal{M} from that operates on ff,

while we use the original pose loss L\mathcal{L} for each f(j)f^{(j)}

However, the hard selection of the minimum in equation 6 does not work in practice as some of the hypothesis functions fθ(j)(I)f_{\theta}^{(j)}(I) might never be updated if they are initialized far from the target values. We relax M^\hat{\mathcal{M}} to M\mathcal{M} by adding the average error for all hypotheses with an epsilon weight:

The normalization constants before the two components are designed to give a weight of (1−ϵ)(1-\epsilon) to M^\hat{\mathcal{M}} and ϵ\epsilon to the gradient distributed over all other hypotheses. When ϵ→0\epsilon\rightarrow 0, M→M^\mathcal{M}\rightarrow\hat{\mathcal{M}}. This is necessary since the average in the second term already contains the minimum from the first one.

2 Architecture

We employ SSD-300 with an extended InceptionV4 backbone and adjust it to also provide the 6D pose along with each detection. In particular, we append two more ’Reduction-B’ blocks to the backbone. Essentially, we branch off after each dimensionality reduction block and place in total 6.0996.099 anchor boxes to cover objects at different scales. Moreover, to include the unambiguous regression of the 6D pose, we modify the prediction kernel such that it provides C+M⋅PC+M\cdot P outputs for each anchor box. Thereby, CC denotes the number of classes, MM denotes the number of hypotheses, and PP denotes the number of parameters to describe the 6D pose. In our case, for each of the MM predicted hypotheses, we regress P=5P=5 values to characterize the 6D pose, composed of an explicitly normalized 4D quaternion for the 3D rotation and the object’s distance towards the camera. We can estimate the remaining two degrees-of-freedom by back-projecting the center of the 2D bounding box using the inferred depth.

Additionally, in line with we conduct hard negative mining to deal with foreground-background imbalances. Thus, given a set of positive boxes Pos and hard-mined negative boxes Neg for a training image, we minimize the following energy function:

For the class and the refinement of the anchor boxes, we employ the cross-entropy loss Lclass\mathcal{L}_{class} and the smooth L1-norm Lfit\mathcal{L}_{fit}, respectively. In order to compare the similarity of two quaternions, we compute the angle between the estimated rotation and the ground truth rotation according to

Additionally, we employ the smooth L1-norm as loss for the depth component Ldepth\mathcal{L}_{depth}.

Altogether, we define the final loss for each hypothesis jj and input image II as follows

3 Processing Multiple Hypotheses

During inference we further analyze the predicted multiple hypotheses in order to determine whether the pose of the object is ambiguous. Notice that prior to this, we first map all hypotheses to reside on the upper hemisphere. If we detect an ambiguity, we additionally exploit the multiple hypotheses to estimate the view-dependent axes of ambiguity.

We analyze the distribution of predicted hypotheses in quaternion space to determine whether the pose exhibits an ambiguity. To this end, Principal Component Analysis (PCA) is performed on the quaternion hypotheses qi\textbf{q}_{i}. The singular value decomposition of the data matrix indicates the ambiguity: if the dominant singular values σ1/2≫0\sigma_{1/2}\gg 0 (σi>σi+1 ∀i\sigma_{i}>\sigma_{i+1}\ \forall i), an ambiguity in the pose prediction is likely, while small singular values imply a collapse to a single unambiguous solution.

We determine the existence of ambiguity by thresholding the value of σ2\sigma_{2}. Empirically, we find the criteria σ2>0.8\sigma_{2}>0.8 to offer good estimations for ambiguity. It is noteworthy that we can learn to detect ambiguities without further supervision, directly from standard datasets.

As mentioned, very prominent representatives for visual ambiguities are symmetries in the objects of interest, as illustrated in Fig. 3 (left) and (mid). Nevertheless, for other objects such as cups, also (self-) occlusion can induce ambiguities in appearance (right).

To calculate a viewpoint dependant ambiguity axis, we take a closer look at the following scenario. A rotation qi=(qi1,qi2,qi3,qi4)\textbf{q}_{i}=\left(q_{i1},q_{i2},q_{i3},q_{i4}\right) rotates the camera c0c_{0} to cic_{i} around the rotation axis

All these rotation axes lie in the same plane which is perpendicular to the ambiguity axis s⊥ai ∀is\perp a_{i}\ \forall i. Thus, if we stack the rotation axes A=(a1T,a2T,⋯ ,anT)A=\left(a_{1}^{T},a_{2}^{T},\cdots,a_{n}^{T}\right), we can formulate the overdetermined linear equation system ATs=0A^{T}s=0. The ambiguity axis can be found as the solution to the optimization problem

4 From Multiple Hypotheses to 6D Pose

After analyzing the distribution of the hypotheses, we can robustly compute the associated 6D pose for each case.

In case of an unambiguous object pose, we utilize the multiple hypotheses as an input for a geometric median (geodesic L1L_{1}-mean ) to improve robustness of the overall estimation

The iterative calculation follows the Weiszfeld algorithm in the tangent spaces to the quaternion hypersphere . From a statistical perspective, our rotation measures are treated as inputs for an L1L_{1}-estimator to robustly detect the geometric median where dgeo\text{d}_{\text{geo}} gives the geodesic distance on the quaternion hypersphere. Note that Gramkow showed that locally, using the Euclidean distance in the ambient, quaternion space well approximates the Riemannian one. In addition, we compute the median depth of all hypotheses. Afterwards, we utilize the center of the 2D detection and backproject it into 3D to obtain the translation and therewith the full 6D pose of the detection.

As the number of possible 3D rotations is finite yet unknown, we employ mean shift to cluster the hypotheses in quaternion space. Specifically, we use the the angular distance of the quaternion vectors to measure similarity and the Weiszfeld algorithm to merge clusters inside mean shift. This yields either one cluster (if the poses are connected) or multiple (if they are unconnected) as illustrated in Fig. 3. For each cluster we compute a median rotation and the median depth to retrieve the associated 3D translation. Note that we only consider the depths of the hypotheses, which contributed to the corresponding cluster. We apply simple contour checks to find the best fitting cluster from which we extract the final 6D pose.

As noted in , domain adaptation between synthetically generated data samples and real-world images trivializes the collection of training data. We render CAD models in random poses and add a series of augmentations, such as illumination changes, shadows and blur, as well as background images taken from the MS COCO .

Evaluation

In this section, we first introduce our experimental setup. Following that, we clearly demonstrate the benefits of our method compared to typical pose estimation systems on a toy dataset. Next, we show robustness in determining whether a view exhibits an ambiguity. Fourth, we report our 6D pose estimation accuracy for the unambiguous and the ambiguous case on common benchmark datasets. Finally, we demonstrate how we can model reliability in pose estimation by analyzing the variance across hypotheses.

In order to properly assess the 6D pose performance, we distinguish between potentially ambiguous and non-ambiguous objects. When dealing with non-ambiguous objects, we report the absolute error for the 3D rotation in degrees and 3D translation in millimeters. We also show our accuracy using the Average Distance of Distinguishable Model Points (ADD) metric from , which measures if the average deviation of the transformed model points is less than 10%10\% of the object’s diameter.

For ‘ambiguous’ objects we rely on the Average Distance of Indistinguishable Model Points (ADI) metric, which extends ADD for ambiguity, measuring error as the average distance to the closest model point .

We also show our results for the Visual Surface Similarity (VSS) metric. As , we define VSS similar to the Visual Surface Discrepancy (VSD) , however, set τ=∞\tau=\infty. Hence, we measure the pixel-wise overlap of the rendered ground truth pose and the rendered prediction, which is not subject to ambiguities.

2 Synthetic Ambiguity Evaluation

We render a simple synthetic dataset of a rotating cup and cube. We compare the baseline with M=1M=1 hypothesis and our method with M=30M=30 hypotheses. The results are shown in Fig. 4, Tab.1, and the supplement. For the cup, both methods yield an ADI score of 100%. The single hypothesis approach SH⁡\operatorname{SH} is indeed able to compute visually correct poses even though it cannot model the pose distribution along an arc. It has learned the conditional mean pose where the handle is exactly opposite of the camera. Nonetheless, this is only one of the infinitely many possible solutions. In contrast, our method is able to predict the whole distribution as seen in the Bingham plots. This is essential for tasks such as next-best-view prediction or robotic manipulation. When there is no ambiguity, both methods predict only the one correct pose.

For the cube object, SH⁡\operatorname{SH} fails (red outline) with an ADI of only 15.6%. Here, the conditional mean is not inside the set of correct poses. Our method is again able to estimate the underlying distribution and can correctly estimate all four modes of correct poses. This yields a perfect ADI of 100%.

When applying our method to real data (Fig. 5), we achieve similar results. If there is a unique solution, the method is able to robustly estimate the correct pose. For ambiguous views, we retrieve the governing distribution as depicted by the viewpoint frustums and spherical plots.

3 Real World Datasets

To conduct evaluations on real data, we build two datasets addressing both unambiguous and ambiguous cases. In particular, for the former, we use the popular ‘LineMOD’ and ‘LineMOD Occlusion’ dataset . The authors of selected one sequence from the original ‘LineMOD’ dataset and labeled eight additional objects. Nevertheless, we moved the ‘glue’ and ‘eggbox’ object to the ambiguous dataset, since both exhibit several views (mostly from the top), which are not unique. Additionally, following we removed the ‘cup’ and ‘bowl’ objects, because no watertight CAD models are provided for them. We also discard the ‘lamp’ since the CAD model does not possess correct normal vectors for proper rendering. To the latter, the ambiguous dataset, besides the ‘glue’ and ‘bowl’ objects, we added several models from T-LESS to cover different types of ambiguities. In essence, T-LESS mostly consists of symmetric and textureless industrial objects. For our experiments we choose a subset that covers both cases: complete rotational symmetry along an axis (object 4) and objects with more than one rotational symmetry (object 5, 9, 10).

4 Ambiguity Detection Analysis

To evaluate the ability of our model to learn pose distributions, we manually labeled for each validation image of the ambiguous dataset, whether the current object view exhibits ambiguity based on the visible object texture and shape. This ground truth is used to quantitatively assess our capability of detecting pose ambiguity. Additionally, we compute the ground truth symmetry axis for each object. It is important to note that we do not conduct object symmetry detection, instead, we describe the perceived pose ambiguity in terms of a symmetry axis. These annotations are only used for evaluation and not during training.

For each detected ambiguity, we compute the average discrepancy of the computed symmetry axis from the ground truth annotation. For the ambiguity-free case, we achieve to report an accuracy of more than 99%, while for the ambiguous case we can also state a high accuracy of 82% correctly classified views. Furthermore, the mean axis only deviates by 24°, which shows that our formulation is able to precisely explain the perceived ambiguity.

In Fig. 6, we respectively show one sample of estimated ambiguity axis from ‘LineMOD’ and ‘T-LESS’. For each detection, we draw the estimated axis in red, while the green line denotes the hand-annotated groundtruth axis.

5 Comparison to State-of-the-Art

In Tab 2 and Tab 3, we report our results for the unambiguous subset for training with synthetic data and with the train data split from . Since the number of predicted hypotheses M\mathbf{M} is a hyperparameter, we will show an ablation in the supplement and only report our best results with M=5\mathbf{M=5} here.

For the case of synthetic training only, even for the single hypothesis case, our approach outperforms SSD-6D by more than 35% of relative error while also being more robust in terms of 2D detection. Comparing with Sundermeyer et al. we can report a relative improvement of approximately 50% referring to ADD. In addition, our averaging over all hypotheses leads to more robustness towards outliers and, thus, another improvement of all metrics.

When also employing real data, we can improve our results by approximately 9% to 44.4% and are on par with the state-of-the-art methods from and , even though we employ no crop and paste augmentations. Further, when using the more challenging ‘LineMOD Occlusion’ dataset, we can exceed Tekin et al. for all objects and overall almost triple their ADD score from 5.8% to 15.6%.

Referring to Tab 4, for the ambiguous ‘LineMOD’ objects, we attain a VSS score of 79% and an ADI score of 55%, which is a relative improvement of approximately 13% and 145% compared to SSD-6D. In the 6D setting, the multiple hypothesis detector overall achieves similar performance as the single hypothesis predictor. However, for the 2D detection case, we are able to increase the accuracy from 79% to 94%. As constituted, only a few views are ambiguous for these objects. Investigating the results, we discovered that the single hypothesis predictor is not able to understand exactly these views and tends to simply discard them. In contrast, the multiple hypotheses predictor is indeed able to understand these views and yields reliable pose predictions.

For all ambiguous ‘T-LESS’ objects (Tab 3), our multiple hypotheses approach surpasses the single hypothesis estimator, which, when trained and evaluated under the same conditions, is not able to capture the ambiguities in pose. Thus, the single hypothesis predictor is not able to produce equally accurate results, being only capable of computing precise poses for unambiguous views. Comparing with , we report similar performance in pose. Our ADI improves with 56.4%56.4\% compared to 50.6%50.6\% while VSS falls slightly behind by 2.8%2.8\%. For fairness, we only compare the 6D pose accuracy for correctly detected objects (i.e. IoU≤0.5IoU\leq 0.5) since trained their 2D detector for T-LESS on real data.

6 Measuring Reliability

To the best of our knowledge, there is no prior work capable of modelling the confidence in the continuous pose estimate. Yet, this information can highly improve the overall robustness and accuracy. In our case, we can utilize the different hypotheses to first determine whether the current view is unambiguous and subsequently employ them as a confidence measurement in the unambiguous 6D pose. To quantify the effect of this, we report our test results on the unambiguous subset of ‘LineMOD’ in Fig. 7 (top), where we compute a confidence measure via the standard deviation with respect to the Karcher mean .

Naturally, a lower standard deviation means more accurate poses. By only allowing poses with σ<0.1\sigma<0.1, all metrics improve, while only losing about 10.5%10.5\% of all estimates. The rotational error decreases by approximately 20%20\% and the translation error drops from 44.8mm to 43.0mm. Accordingly, using an even lower threshold (e.g. σ<0.05\sigma<0.05) gives another significant improvement for pose (especially in rotation), however, at the cost of rejecting more estimates.

The qualitative example image in Fig. 7 also confirms these results. The pose with the lowest standard deviation for the ‘driller’ is very accurate, and the one with the highest is rather imprecise. We experience the same behavior for all unambiguous ‘LineMOD’ objects.

Conclusion

We propose a new approach for pose estimation that implicitly models ambiguities without requiring any input pre-processing as well as the feasibility of domain adaptation between synthetic and real data. In addition, we can estimate the axis of rotational ambiguity and perform pose refinement based on clustering without knowing the number of clusters in advance. Our experiments show that our method is suitable for detecting both challenging objects with multiple rotational symmetries and datasets with little ambiguity. Lastly, we argue that our method constitutes a metric of reliability for the 6D pose.

In conclusion, we believe that the new formulation of the pose detection problem from images as an ambiguous task paves the way towards interesting applications in the domain of robotic interactions and automation.

We would like to thank Toyota Motor Corporation for funding and supporting this work and NVIDIA for the donation of a GPU.

Explaining the Ambiguity of Object Detection and 6D Pose from Visual Data

Datasets

In Fig 1 we would like to demonstrate all the objects we employed for our experiments. Thereby, the upper row illustrates all objects of the unambiguous dataset, taken from ‘LineMOD’ . These objects do not exhibit any views which might induce ambiguities. On the contrary, the lower row depicts all objects of the ambiguous dataset. While the first two objects also belong to the ‘LineMOD’ dataset, the last four accompany the T-LESS dataset . All these objects can induce ambiguities for certain viewpoints. For instance ‘obj 04’ is a symmetric screw, however, possessing distinct textures on its head. Due to this only the views from the bottom (which do not show the texture) are ambiguous. In contrast, for each viewpoint in ‘obj 09’ and ‘obj 10’, there exists always one identical viewpoint on the other side. Thus, these objects are never ambiguity-free.

Robust Ambiguity Detection and Estimation

Tab.1 shows our detailed ambiguity detection results for the unambiguous (top) and ambiguous (bottom) objects, respectively. In addition, we also report our individual results for the ambiguity axis estimation. We compute the mean deviation from the labeled ground truth. As a threshold for σ1\sigma_{1} we empirically find 0.80.8 to offer good accuracy. Fig. 2 demonstrates more qualitative results for ambiguity detection and the computation of the corresponding ambiguity axis.

2D Object Detection and 6D Pose Estimation

In this section, we present our detailed results for 6D pose estimation and 2D detection. As in the paper, for the unambiguous dataset we present our numbers with M=5M=5 and for the ambiguous dataset we set M=30M=30.

We present an ablation study for different numbers of hypotheses M\mathbf{M} in Tab. 2. We obtained our best results employing M=5\mathbf{M}=5 hypotheses. Below we show one qualitative sample for each object. In addition, on the right we also visualize the corresponding Bingham Distributions for visual validation. Lastly, we depict some qualitative results on the ‘LineMOD Occlusion’ dataset.

2 Ambiguous Object Detection and Pose Estimation

Since comparing against the ground truth is not suitable in a multiple hypothesis scenario, only metrics that do not rely on this value are apt for this case. We thus chose the Visual Surface Similarity and Average Distance of Indistinguishable points as metrics for pose. We always take the detection with the highest confidence. We present our individual scores for the ambiguous dataset in Tab 3. Additionally, below we show one qualitative sample for each object.

Employing Multiple Hypothesis as Measurement For Reliability

We would like to present more qualitative samples that the hypotheses can be employed as measurement for confidence. To this end, for each object of the unambiguous dataset we show the poses possessing the lowest and the highest standard deviation in the hypotheses.

Implementation Details

We implemented our method in TensorFlow v1.5 and conducted all experiments on an Intel i7-5820K@3.3GHz CPU with an NVIDIA GTX 1080 GPU. We train with a batch size of 10 and use Adam with a learning rate of 10−410^{-4}.

We decay the relaxation weight ϵ\epsilon from 0.050.05 to 0.010.01 during training (Eq. 7). Further, we empirically set α=1.0\alpha=1.0 and linearly increment β\beta from 33 to 1010 (Eq. 8). Finally, we set λ=3\lambda=3, which balances rotation and translation in the final loss (Eq. 10).

To avoid hypotheses to die due to bad initialization, besides sharing loss through the relaxation weight ϵ\epsilon, we also employ Hypotheses Dropout: during training we deactivate a hypothesis with a probability of p=50%p=50\% for the current image.

The mean shift and PCA implementations were taken from scikit-learn. We use verify_6D_poses in rendering/utils.py from ’s git repository to find the best cluster after mean shift.

To estimate and plot the Bingham distributions, we referred to this https://github.com/SebastianRiedel/bingham matlab implementation. Given a set of 4D quaternions, we compute the maximum likelihood Bingham distribution employing bingham_fit. We then render the sphere conducting an equatorial projection to 3D (plot_bingham_3d). Similarly, we also project the groundtruth and single hypothesis quaternions to 3D and superimpose them on the rendered sphere.

The pseudo-code below depicts the 6D pose inference procedure after the input image has been processed by the network.

Synthetic Training Samples

We generate synthetic samples by rendering objects with random poses onto images from the MS COCO dataset . Using OpenGL commands we generate a random pose from a valid range: 360º on the azimuth and altitude along a view sphere, and 180º for inplane rotation. We also vary the radius of the viewing sphere to enable multi-scale detection. In order to increase the variance of the dataset, we add random perturbations such as illumination and contrast changes, among others. This is a similar approach to . However, in contrast to them, for each assigned anchor box, we save exactly one 4D quaternion as the ground truth for the rotation, even if ambiguous.

References