6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints
Chen Wang, Roberto Martín-Martín, Danfei Xu, Jun Lv, Cewu Lu, Li Fei-Fei, Silvio Savarese, Yuke Zhu
I Introduction
Estimating 6D pose of objects, i.e., translation and orientation in 3D, offers a concise and informative form of state representation for robotic applications, such as manipulation and navigation . In robotic manipulation, the ability of tracking object 6D poses in real-time gives rise to fast feedback control . Pioneering work in 6D tracking has achieved remarkable accuracy and robustness given the 3D model of an object instance, often referred as instance-level 6D tracking. However, the assumption of known 3D model can be brittle in realistic settings, where perfect geometry of novel objects is hard to acquire. In this work, we propose to study the problem of category-level 6D tracking, where the goal is to develop category-level models capable of tracking novel object instances within a specific category.
The problem of category-level tracking has been studied extensively in 2D domains. Classical methods rely on handcrafted features as object representations for visual tracking . Recent work has embarked on an exploration of new computational tools, in particular, deep neural networks, and large amounts of training data to improve tracking performance under visual variations and heavy occlusions . However, a model for 6D pose tracking would have to handle the larger search space of all possible poses due to the increased dimensionality, leading to a substantial computational burden over 2D visual trackers.
One remedy is to reduce category-level 6D tracking to a 3D detection and 6D pose estimation problem. 3D detection and pose estimation have been studied in a large body of literature, especially in the context of autonomous driving . Most relevant to us is NOCS which introduced a category-level model to estimate the 6D pose of objects from RGB-D images. NOCS transforms every object pixel to a shared coordinate frame as keypoints for pose estimation. However, estimating poses from a large number of crude keypoints makes their method susceptible to noises from clutter and occlusion. Furthermore, these tracking-by-detection methods cannot leverage temporal information from previous frames. In contrast, we seek to develop a tracking model that learns compact and discriminative object representations for robust registration and leverages temporal consistency for efficient search.
To this end, we propose 6-PACK, a vision-based 6D-Pose Anchor-based Category-level Keypoint tracker. 6-PACK tracks a small set of keypoints in RGB-D videos and estimates object pose by accumulating relative pose changes over time (see Fig. 1). This method does not require the known 3D model. Instead, it circumvents the need of defining and estimating the absolute 6D pose via a novel anchor mechanism analogous to the proposal methodology used in 2D object detection . These anchors offer a base for generating 3D keypoints. Unlike previous methods that require manual keypoint annotations , we propose an unsupervised learning approach that discovers the optimal set of 3D keypoints for tracking. These keypoints serve as a compact representation of the object, from which its pose difference between two adjacent frames can be efficiently estimated . This keypoint-based representation leads to robust and real-time 6D pose tracking.
We show experimentally improved generalization and robustness in category-level 6D tracking compared to the traditional registration-based tracker , the state-of-the-art tracking-by-detection method , and other ablative baselines. 6-PACK substantially outperforms all baselines on the recently introduced NOCS-REAL275 dataset . The proposed tracker runs at 10Hz on a GTX1070 GPU. We deploy this tracker on a Toyota HSR Robot and demonstrate its utility in real-time manipulation tasks.
II Related Work
Historically, 6D object visual pose estimation has relied on matching the current view of an object to a given template model. The pose can then be recovered from solving an error minimization problem between correspondences, either a PnP, a reprojection or a distance error, depending on the nature of the template . While conceptually simple and efficient in practice, the performances of these methods degrade significantly under clutter or variable lighting conditions due to errors in feature matching.
Recent 6D object pose estimation methods instead learns to directly match input images with renderings or silhouette of template object models. PoseRBPF compares the latent code of the input image and that of the model rendering to recover the rotation part of the object pose. However, the requirements of knowing the object models confine these methods to known object instances.
The problem of category-level pose estimation has amassed significant interests in research areas such as autonomous driving due to the availability of large-scale datasets . Most relevant to us is a class of methods that combines the classical keypoint matching ideas and modern learning techniques by directly predicting either category-level semantic keypoints or 3D bounding box corners . However, supervised keypoints learning require large amounts of labelled data, and the manually annotated keypoints or the bounding box corners may not be the optimal landmarks to track. Moreover, defining features a priori (so-called feature engineering) has largely been outperformed by data-driven methods that learn the best features for the task .
A notable exception is NOCS , which learns to project every input object pixel into a category-level canonical 3D space as keypoint. The pose is then recovered by registering all projected keypoints to an “average” category object model. However, as we show experimentally, the reliance on dense keypoint correspondences makes NOCS susceptible to noise from occlusion. In contrast, our novel keypoint detector is trained end-to-end only with the final pose tracking objective to generate a small but robust set of 3D keypoints for tracking without direct keypoint supervision.
Our keypoint generator is closely related to the large corpus of keypoint detection and matching methods. Classical methods detect and match keypoints based on hand-designed local features . Modern learning approaches learns to detect keypoints via optimizing fully-supervised or semi-supervised objectives . Recently, KeypointNet shows the benefits of learning to generate keypoints in 3D space without supervision by exploiting geometric consistency across multiple views. The model generalizes to new object instances within known categories in a synthetic domain. However, as we show in our experimental evaluation, KeypointNet fails to scale to a real-world since the model has been trained to define keypoints on single models of objects centered at the origin of coordinates. This approach suffers when exposed to noisy RGB-D data from a real scene where the objects have occlusions and are not centered, increasing the boundaries of the 3D space where the keypoints could be found. Our method introduces a novel anchoring mechanism that allows the model to generate keypoints only in the most relevant subspace. The strategy significantly improves the quality of the generated keypoints and reduces the number of keypoints needed for estimating the pose, enabling our model to track objects in real-time.
III Problem Definition
The initial pose is the translation and rotation with respect to the camera frame of a canonical frame defined similarly for all instances of the same category. This setup was defined in NOCS for the related problem of category-level 6D pose estimation. For example, for the category camera, the frame is placed at the centroid of the object with the -axis pointing in the direction of the camera objective and the -axis pointing upwards. Similar to prior 6D tracking work , we assume the initial pose of the object is given. However, different from these approaches, our method is robust to errors in this initial pose, as shown in our experimental evaluation (Sec. V).
IV Model
6-PACK performs category-level 6D pose tracking in the following manner (Fig. 2, bottom-right). First, 6-PACK uses an attention mechanism over a grid of anchor points (Sec. IV-A) generated around the predicted pose of the object. Each anchor summarizes the volume around it with a distance-weighted sum of the individual features of the RGB-D points in its surrounding. This information allows to find a coarse centroid of the object in the new RGB-D frame and guide the following search of keypoints around it, which is more efficient than searching keypoints in the entire unconstrained 3D space of previous approaches .
Second, 6-PACK uses the anchor feature to generate keypoints for both symmetric and non-symmetric categories (Sec. IV-B and Fig. 2 left and top-right). Different from previous methods (e.g., kPAM ), these keypoints are learned in an unsupervised manner so they are the most robust and informative for tracking based on training data.
Finally, the keypoints of the current and previous frames are passed to a least-squares optimization that calculates the inter-frame change in pose. Based on this motion, we extrapolate the pose in the next frame to center the next distribution of anchor points.
The process starts at the location indicated by a given initial pose, . We allow for this initial pose to contain an error that we reject with an initial iterative procedure. We generate a set of keypoints and refine the given pose to be at the centroid of the set, . The centroid of the keypoints is close to the instance centroid as it is imposed in the training process (Sec. IV-B). We then run the generation and correction again, a total number of times. We select the refined pose that is closest to the centroid of generated keypoints as initial pose for the tracking procedure, since this refined pose is most likely to be the correct one for the instance of the class. The initialization process reduces the effort of providing a very accurate initial pose and increases the robustness of the category-level tracker.
Directly generating a set of ordered 3D keypoints for pose tracking is challenging due to the large output space. Prior work on automatic keypoint generation did not address this problem: in their setup, keypoints are generated for a single object within a bounded, origin-centered sphere. However, in our setup, the object can be anywhere in 3D space within the field of view of the RGB-D frame. Thus, the keypoints can be anywhere in the unbounded 3D space.
The stated problem of 3D keypoint generation in unbounded space resembles that of 2D object detection, where the goal is to draw a tight bounding box around the target object on the 2D image. A successful solution to this problem is to begin by drawing a grid of anchor points in the image and find the anchor that is closest to the object center. This procedure locates coarsely the object and simplifies the second step, the generation of a more accurate bounding box proposal around the anchor . Inspired by this idea, we propose an attention mechanism over a grid of 3D anchors around the predicted current object location. Each anchor contains a feature representation of the surrounding volume around it. The model learns to attend, based on this feature, to the anchor closest to the centroid of the object. The 3D keypoints could then be generated as offset points from the selected anchor (Sec. IV-B). By splitting the tracking problem into coarse attention-based anchor selection and fine-grained keypoint generation, our tracker has potential to tackle larger region of search space (improving robustness and perceivable inter-frame motion) while maintaining high tracking quality.
As mention before, each anchor contains a feature representing the 3D points in the volume surrounding it. As showed by prior work , it is possible to combine color and geometric information into a fused feature to be used for pose estimation. Therefore, we apply the DenseFusion feature embedding to all colored 3D points from the RGB-D image within the grid of anchors, and use a distance-weighted averaging to pool them into each anchor to generate the anchor embedding.
Once we have generated a per-anchor geometric and color feature, we train an attention network that learns to detect the one closest to the object centroid. Our attention network takes as input each anchor-level embedding and produces as output a confidence score per anchor. The attention network is trained to assign the highest confidence score to the anchor closest to the centroid of the object with supervision. Thus, given the ground-truth position of the object centroid , the loss function can be written as:
where refers to the minimum possible distance from the anchor to the object centroid. We use a two-layer MLP as the attention network. During the evaluation, 6-PACK selects the anchor with the highest confidence score and generates the keypoints as offsets to this anchor, as explained in the next section.
IV-B Unsupervised 3D Keypoints Generation
With the anchor-based attention mechanism, 6-PACK identified the anchor and associated feature with the highest confidence score. Now, 6-PACK will use this feature to generate the final set of 3D keypoints, , to track the instance of the object category. We propose a keypoint generation neural network that uses as input the anchor feature and generates a dimensional output containing an ordered list of keypoints. Since the list is ordered, we do not need to find correspondences between keypoints of consecutive frames to estimate the change in pose.
As mentioned before, we train our keypoint generation network in an unsupervised manner that does not require of manual annotation, and that, compared to supervised methods, has led to improved transferability on category-level orientation estimation . We render the unsupervised training of our keypoint generation network as optimizing the multi-view consistency between the keypoints generated in consecutive frames. In other words, suppose keypoints are generated in each of two consecutive frames; the training objective is to place the keypoints in the current view at the location that corresponds to the keypoints of the previous frame, transformed by the ground truth inter-frame motion. This objective can be formalized in the following multi-view consistency loss:
where is the ground truth inter-frame change in pose.
The multi-view consistency loss only guarantees inter-frame consistency between features locations independently of the perspective or the visible part of the object. However, this does not guarantee these locations are optimal for our final goal, to estimate the change of pose (e.g., all keypoints could end up all at the same location). To tackle this problem we turn the keypoint-based pose estimation step into a differentiable pose estimation loss function and combine it with the multi-view consistency loss in the training process . This loss function consists of a translation loss and rotation loss :
where and are the centroids of the keypoints in previous and current frames, and is the inter-frame change in orientation estimated based on the generated keypoint sets using least-squares optimization . Thus, these losses force the keypoints to be generated such that the ground truth change of pose can be computed from them.
We also integrate a separation loss and silhouette consistency loss defined by Suwajanakorn et al. . The separation loss forces keypoints to maintain some distance to each other to avoid degenerate configurations and improve the pose estimation. The silhouette consistency loss forces keypoints to be closer to the object’s surface to improve interpretability.
In addition to the main objectives introduced above, we impose a centroid loss that forces the centroid of the generated set of keypoints to be at the centroid of the object, , useful to correct for noise in the given initial pose as explained at the beginning of this section. The final overall training loss is a weighted sum of the 6 terms introduced above, where the weightings are determined by their relative magnitude and importance.
3D Keypoint Generation for Classes with Symmetry Axes: The presented multi-view consistency and pose loss functions do not handle well symmetries on instances of object categories since identifying the rotation along the symmetry axis is an unsolvable problem. We propose a coordinate system transformation that transforms the coordinates of points into a space that is rotation-invariant around the axis of symmetry. Fig. 3 illustrates the transformation for a case with five points on an instance of bowl. Suppose some common axis of symmetry for all instances in this category, , passing through the Cartesian origin of coordinates (y-axis in Fig. 3); we transform the location of a keypoint point from in the Cartesian coordinates to the triplet defined as:
Distance to the symmetry axis : The distance from point to the symmetry axis, .
Height along the symmetry axis : Distance between the orthogonal projection of onto the symmetry axis and the Cartesian origin of coordinates.
Relative angle : The separating angle between the radial vector connecting the point to the symmetry axis and the radial vector of the next keypoint encountered when advancing clockwise around the .
The new coordinates are closely related to cylindrical coordinates, where we replace the absolute rotation angle by a relative inter-keypoint angle. Based on the symmetry-invariant transformation, we redefine the multi-view consistency loss for symmetric categories as:
The pose estimation loss also needs to be adapted for categories with symmetry axes. While the translation loss remains unaltered, we redefine the rotation loss as the angular difference between the predicted change and the ground-truth change in orientation of the symmetry axis. The rotation loss for these categories is then simply:
V Experiments
In this section, we would like to answer the following questions: 1) Does our method indeed generate robust 3D keypoints that are suitable for 6D pose tracking? 2) Does our anchor-based attention mechanism improve the overall tracking performance? 3) How robust is our method against variable levels of noise in pose initialization, and 4) Is our method efficient enough for real-time applications such as closed-loop object manipulation?
To answer questions 1), 2) and 3), we evaluate our method and compare it to multiple baselines on the NOCS-REAL275 dataset, the only real-world benchmark dataset for category-level object 6D pose tracking. To answer question 4), we deploy our model on a real robot platform and test the model with another ten unseen objects and show that the model can be successfully used in a collaborative pouring and a “toasting” tasks.
Dataset: We use the NOCS-REAL275 dataset, which contains six categories including bottle, bowl, camera, can, laptop, and mug. Three of them are categories with axes of symmetry. The training set consists of 275K discrete frames of synthetic data generated with models of instances of the classes from ShapeNetCore with random poses, and seven real videos with ground truth poses depicting in total three instances of objects of each category. The testing set has six real videos depicting in total three different (unseen) instances for each object category with 3,200 frames in total.
Baselines: We compare three model variants with baselines to showcase the effectiveness of our design choices:
NOCS : The state-of-the-art category-level 6D pose estimation method that uses per-pixel prediction.
ICP : The standard point-to-plane ICP algorithm implemented in Open3D.
KeypointNet : The implementation of our model without the anchor-based attention mechanism, which direct generates 3D keypoints in the 3D space.
Ours without temporal prediction: An ablation of 6-PACK where the predicted pose in the next frame is the previous estimated pose.
Ours: 6-PACK where the predicted pose in the next frame extrapolates from the last estimated inter-frame change of pose (constant velocity model).
Evaluation Results on the NOCS-REAL275 dataset: Table I contains the results on the testing set of all six categories of the NOCS dataset. We compare two variants of our model (with and without temporal prediction) and several current state-of-the-art category-level 6D pose estimation methods. To evaluate the robustness against noisy initial poses, we inject up to 4cm of uniformly sampled random translation noise. We also measure robustness against missing frames by dropping 450 frames out of 3200 frames are uniformly from the testing videos.
Additionally, 6-PACK outperforms KeypointNet by a large margin in all metrics. The IoU25 metric reveals that KeypointNet frequently loses track of the object (50.3% across the evaluation set). As discussed in Sec. II, KeypointNet is not designed to handle the large unbounded keypoint generation space of our real-world setting. On the other hand, our method avoids losing track of category instances in evaluation (IoU25 >94%). Our anchor-based mechanism increases the space of search avoiding drifting.
We also evaluate the sensitivity of our proposed method in using different number of keypoints on the laptop category. Our model with 8-keypoints output achieves 62.4% in the 5°5cm metric, surpassing the 4-keypoints 55.2% and the 16-keypoints 48.6% variants. 8-keypoints offer the best trade-off between information compression and redundancy.
VI Conclusion
We presented 6-PACK, a category-level 6D object pose tracker. Our tracker is based on a novel anchor-based keypoint generation neural network that detects reliably the same keypoints on different instances of the same category and uses them to estimate the inter-frame change in pose. Our method is trained in an unsupervised manner to allow the network to select the best keypoints for tracking. We compared 6-PACK to 3D geometry methods and deep learning models, showing that our method achieves state-of-the-art performance on a challenging category-based 6D object pose tracking benchmark. Furthermore, we deployed 6-PACK on an HSR robot platform and showed that our method enables real-time tracking and robot interaction.
Acknowledgement
This work has been partially supported by JD.com American Technologies Corporation (“JD”) under the SAIL-JD AI Research Initiative and by an ONR MURI award (1186514-1-TBCJE). This article solely reflects the opinions and conclusions of its authors and not JD or any entity associated with JD.com. We also want to thank Toyota Research Institute for the Human Support Robot which we used to perform our real robot experiments.