StereOBJ-1M: Large-scale Stereo Image Dataset for 6D Object Pose Estimation
Xingyu Liu, Shun Iwase, Kris M. Kitani
Introduction
Effectively leveraging 3D cues from visual data to infer the pose of an object is crucial for applications such as augmented reality (AR) and robotic manipulation. Compared to objects with opaque and Lambertian surfaces, estimating the pose of transparent and reflective objects is especially challenging. To leverage depth information from sensors, previous approaches have explored deep models that take RGB-D maps as input . Unfortunately, as the experiments in have shown, existing commercial depth-sensing methods, such as time-of-flight (ToF) or projected light sensors, failed to capture the depths of transparent or reflective surfaces. As a result, monocular RGB-D maps cannot serve as a reliable input for object pose estimation models in these challenging scenarios. Based on this observation, we focus on using stereo RGB images as our input modality, allowing for object pose estimation on a wider range of objects, including transparent or highly reflective objects.
A major challenge in modern object pose estimation is that of acquiring a large-scale training dataset. To increase data size for training large-scale neural networks, previous works have explored leveraging synthetically rendered or augmented images with 3D mesh models. However, photorealistic rendering is still challenging with only basic graphics rendering tools and limited expertise. Synthetic image datasets that are currently available typically introduce a very large domain gap. This is especially true for transparent and reflective objects where variations in illumination and background scenes are crucial but difficult to model.
To address the challenge of costly pose data acquisition, and to enable further training and evaluation of modern object pose estimation models, we introduce a novel method for capturing and labeling a large-scale dataset with high efficiency and quality. Our method is based on multi-view geometry to accurately localize fiducial markers, cameras, and object keypoints in the scene. We use a hand-held stereo camera to record video data. With the help of two other static cameras mounted to tripods, the positions of a set of fiducial markers can be calculated on the fly, from which the pose of every recorded frame can be automatically computed. By annotating 2D object keypoints in just a few frames selected in a long recorded video, the 3D locations of the keypoints can be computed by triangulation. The 6D poses of objects can then be calculated by aligning 3D CAD models to the keypoints before being propagated to all other frames.
Using the procedure outlined above, we generate StereOBJ-1M dataset, the first pose dataset with stereo RGB as input modality with over 100K frames. It is also the largest 6D object pose dataset in history: it consists of 393,612 high-resolution stereo frames and over 1.5 million 6D pose annotations of 18 objects recorded in 182 indoor and outdoor scenes. The capacity of StereOBJ-1M is sufficient for training large-scale neural networks without additional synthetic images. The average labeling error of StereOBJ-1M is 2.3mm which is the best annotation precision among all public object pose datasets.
We implement two state-of-the-art methods as the baseline comparisons for 6D pose estimation using stereo on the StereOBJ-1M dataset. To handle 2D-3D correspondences predictions in two or more images, we propose a novel object-level 6D pose optimization approach named Object Triangulation. Contrary to classic triangulation that optimizes the 3D location of a point, we directly optimize the 6D pose of an object from 2D keypoint locations in multiple images. Experiment results show that Object Triangulation consistently improves pose estimation over monocular input while classic triangulation can yield worse results. With Object Triangulation, the stereo variants of both baseline methods significantly outperform their monocular counterparts on StereOBJ-1M, by at least 25% in ADD(-S) AUC and 14% in ADD(-S) accuracy, which highlights the importance of stereo modality in object pose estimation. We expect that StereOBJ-1M will serve as a common benchmark dataset for stereo RGB-based object pose estimation.
Related Work
Pose Annotation Methods. The first category of pose data annotation methods relies on capturing RGB-D images, reconstructing 3D point clouds, and labeling pose by constructing a 3D mesh , or fitting 3D object mesh models to 3D point clouds . However, this type of method cannot reliably deal with transparent objects where depth sensing is usually not possible. The second category of pose data annotation methods adopts keypoint as representation and leverages multi-view geometry for triangulation . Our novel data annotation method is keypoint and multi-view based. Different from previous methods, we record the scenes using a stereo RGB camera whose poses are computed on the fly based on fiducial markers whose locations are also computed on the fly.
Stereo Methods. Studying correspondence, depth, and other downstream tasks from two or multiple RGB images has been a long-standing topic in computer vision and robotics. Previous works have explored stereo-based methods for 3D object detection , disparity estimation , point-based 3D reconstruction and keypoint detection . Recently, multi-view based methods have been proposed for object pose estimation . Our 6D object pose dataset provides binocular stereo RGB images as the input modality, allowing stereo-based deep methods to be trained on object pose data. In addition, the annotation process of our dataset utilizes multi-view geometry.
Keypoint in Pose Representation. Keypoints is a popular intermediate representation for object or human pose. Previous work has explored deep learning methods for localing keypoints of an object or a human from a RGB image. Several public object pose estimation datasets are also constructed by using keypoints to simplify pose annotation . Our data annotation pipeline also uses keypoints as a bridge to the 6D pose, where the 3D positions of the keypoints are calculated through multi-view triangulation.
Related Dataset. Most existing pose datasets provide RGB-D as input modality . Since directly labeling 3D object pose in real RGB images is costly and inaccurate, most existing datasets rely on capturing RGB-D images and fitting 3D mesh models to 3D point clouds as their labeling method . TOD is the first object pose dataset with binocular stereo RGB as the input modality, and it uses a data labeling method based on multi-view geometry. However, TOD records in a studio environment and does not include occluded objects. Our dataset provides binocular stereo RGB as input modality and records objects with occlusion in 11 different real environments. A more comprehensive comparison of datasets is illustrated in Table 1.
Data Capturing & Labeling Pipeline
One of the major challenges in the pose estimation of 3D objects is the acquisition and annotation of large-scale and high-quality real object pose data. Limitations of previous efforts are in the following three aspects.
Sensor Modality. Most existing datasets such as only provide monocular RGBD from commercial depth sensors as 3D cue. These datasets and the associated labeling methods did not and could not handle transparent or reflective objects on which depth sensing is not reliable. Moreover, different technologies of depth sensing, e.g. infrared and LiDAR, may return different depths for the same objects and scenes. Thus a model trained on one RGBD-based pose dataset may not be able to generalize to another with a different depth-sensing technology.
Data Annotation. Existing data annotation methods usually require annotators to manually align object CAD model to 3D sensor signals, e.g. reconstructed 3D point clouds from depth maps, which are expensive and inaccurate. Limited by the cost of data annotation, the sizes of public real-world datasets such as are in the order of 10K or fewer images, which are insufficient for training large-scale deep neural net models. An alternative solution is to leverage synthetically rendered or augmented images. However, the problem of the domain gap is still yet to be solved and is especially challenging for transparent and reflective objects.
Scene Environment. Datasets such as were captured in a small number () of special indoor environments or studios and lack the diversity of real scenes. It is hard for models trained on such data to generalize to unseen environments, especially for transparent and reflective objects where background scenes and illumination are crucial.
To address the above problems, we propose a novel method for efficiently capturing and labeling 3D object pose data. We opt to use stereo RGB modality to provide 3D cues for the data. For labeling, our philosophy is to abandon depth sensing and utilize multi-view geometry for high-precision 3D localization of object keypoints for pose fitting. An overview of our pipeline is illustrated in Figure 2. It consists of the following seven steps which respectively correspond to Figure 2(a)-(g).
3. Scene Construction and Scanning. To construct the scene, we first place a few randomly selected objects from our dataset and mingle them with the small fiducial markers. Other occluding objects can also be included if necessary. Note that during this step, the positions of the small fiducial markers must remain unchanged while the static cameras can be removed. Then a human data collector holds a stereo RGB camera, slowly moves it around the scene, scans the objects from different viewpoints, and record a stereo video. The scanning paths are selected aiming to cover as many viewpoints as possible.
5. Keypoint Annotation. From all valid frames, we select a few to annotate the 2D locations of projected object keypoints on the images. The frames are selected using farthest point sampling (FPS) such that their camera translations are as far away from each other as possible. The keypoints of an object are defined by experts and are easy to be spotted and accurately located, e.g. corners. Note that it is possible that only a subset of the total keypoints are annotated in one particular frame.
6. Keypoint Triangulation. For each keypoint of an object, we retrieve all frames in which the keypoint is annotated. Using the moving camera pose and the 2D annotations, the 3D location of the keypoints in the world coordinate can be calculated by multi-view triangulation.
7. Pose Fitting. To obtain the 6D poses of the objects in the world coordinate, we solve an Orthogonal Procrustes problem to fit the object CAD models to the annotated 3D keypoints. Finally, the object poses are propagated to all valid frames via an inverse transform of the camera pose .
2 Labeling Error Analysis
An intriguing question that needs to be answered is: how accurate is our labeling method? We assume the error of the dimensions of the large fiducial marker array board and the small fiducial markers are negligible since they are both accurately measured by vernier caliper. Then the labeling error can come from two steps: automatic detection of small fiducial marker boards and the annotation of keypoint 2D locations, which contributes to the error in two nonlinear optimizations respectively: camera pose estimation and 3D point estimation from multiple views.
We use Monte Carlo simulation to quantify the pose annotation error with a similar procedure as in . Specifically, we dither the keypoint 2D projections according to the keypoint re-projection RMSE statistics and estimate the 3D keypoint error as an approximate of the labeling error. We report keypoint label error of 2.3mm RMSE as illustrated in Table 2. The reasons for label error improvement over are two folds. First, our stereo camera has a higher resolution than and allows more accurate labeling in 2D. Second, our object scanning paths are determined by human data collectors on the fly instead of being hard-coded and performed by robot , thus are more flexible and can adapt to specific scenes to cover more viewpoints and provide wider baselines for triangulation.
3 Comparison to Previous Labeling Methods
We point out that the idea of using multi-view and keypoints for pose labeling can also be found in human pose estimation scenarios such as the Panoptic Studio dataset . Unlike which relies on 480 fixed cameras mounted in a specially constructed studio for triangulation, our data acquisition method is affordable and portable — it only requires three cameras and two tripods, and can therefore be deployed in diverse indoor and outdoor environments. On the contrary, to construct datasets such as , a studio equipped with multiple sensors or robot assists has to be specially constructed. In addition to the logistic cost, such settings are not flexible enough for environments in the wild and therefore suffer from the lack of diversity of data.
TOD is the first object pose estimation dataset that provides stereo RGB modality. Our data capturing pipeline is different from in terms of moving camera pose calculation. Datasets such as rely on customized board printed with fiducial markers, and objects are placed near the center of the board. Thus only the simplest planar terrain can be used with the objects and lacks diversity. Instead, in our data pipeline, we distribute small fiducial markers into the scene and calculate their locations on the fly with the help of two static cameras. This allows the objects to be placed in more flexible and complex background terrains.
Our data pipeline has a much higher data efficiency than TOD . With the proposed data pipeline, in each constructed scene, we can capture and annotate more than 2,000 valid frames with a single scan. As a comparison, only captures 80 frames per scene with the help of a robot arm. An explanation is that in , the predefined automatic scanning path of the robot is limited by its operational space. In our data pipeline, the scanning is performed by humans and can adapt to different scenes, which results in (1) more valid frames per video; (2) larger coverage of viewpoints; and (3) wider baseline during triangulation and therefore higher precision.
StereOBJ-1M Dataset
With the proposed method, we construct StereOBJ-1M, a large-scale dataset, and benchmark for 3D object pose estimation from stereo RGB images. In this section, we provide technical details to our StereOBJ-1M dataset in terms of object 3D models and data sample illustration.
There are 18 objects included in our dataset. Among them, 10 objects are plastic tools used in biochemical labs and 8 objects are metal mechanics tools, which together include both transparent and reflective instances. We provide 3D CAD models of the 18 objects as illustrated in Figure 3. The CAD models are obtained using a high-precision EinScan Pro 2X Plus scanner which has a scan accuracy of 0.04mm. During scanning, reflective metallic parts of the objects are covered with white scanning spray. Among the 18 objects, there are 8 objects with discrete 2-fold rotational symmetry and one with continuous rotational symmetry. Among the 18 objects, microplate, tube_rack_2ml and tube_rack_50ml are transparent; centrifuge_tube, sterile_rack_10ml, sterile_rack_200ml and sterile_rack_1000ml are translucent.
The set of the objects used in our dataset has a special feature: it includes visually similar but different object instances. For example, as illustrated in Figure 3, the three pipettes are almost identical in their geometric features. In our dataset, we include image sequences where two or more similar but different object instances are present in the same scene. Thus it poses a new research question for the computer vision community: how to detect and estimate the poses of visually very similar but different objects? We expect that this question can be studied with our dataset.
2 Data Collection and Annotations
We collected the data in 8 real-life indoor environments, including desktop, washbasin, wooden floor etc. In addition to the indoor environments, we adopt 3 outdoor environments to enrich the diversity in background scenes. In each environment, we shuffle the objects and occlusion clutters several times to construct multiple scenes. In total, we constructed 182 scenes. A stereo video was recorded in every constructed scene. The lengths of the video range from 2 to 7 minutes. When sampled at 15 frames/sec, the recorded videos yield 393,612 stereo frames in total. On average, there are more than 2,100 stereo frames in every scene. Our dataset consists of 182 videos and contains over 1.5 million object pose annotations. The number of annotations of each object in each environment is illustrated in Figure 4.
Viewpoint coverage of each object is illustrated in Figure 5. For objects such as microplate and sterile tip racks, there is only one possible side when putting on a desktop, so at most 50% of viewpoint coverage. Annotations in our dataset are 6D poses for every object in the scene, from which object instance segmentation masks, 2D and 3D bounding boxes, and normalized coordinate maps can be inferred. We visualize some data samples from our dataset in Figure 6. As illustrated in the figure, the annotations of our dataset have high quality.
3 Benchmark and Evaluation
Train/Validation/Test Split. The image sequences are divided into train/validation and test sets such that scenes presented in the training set are held out in the validation and test set. The test set contains 32 image sequences that are selected to cover most environments and ensure every object is tested in at least 4,000 images across at least 3 different scenes. In the baseline experiments in Section 5, we did not render additional synthetic data except basic geometric and photometric augmentation, because the capacity of StereOBJ-1M training set is sufficient to train large deep models. However, users of the dataset can still opt to render additional data using the 3D mesh models we provide.
Among the objects, centrifuge_tube is the only category with multiple instances recorded in a scene and is used in the multi-object pose detection task. The rest of 17 objects are used in single-object pose estimation task which is the main focus of this paper. Results for pose detection of centrifuge_tube are provided in supplementary.
where and are the ground truth and estimated 6D poses. For symmetric objects, ADD-S is used instead. When computing ADD-S distance, the 3D distances are calculated as the average of each point’s closest distance to the other point set:
We use the following two evaluation metrics. (1) ADD(-S) accuracy: ADD(-S) accuracy measures the proportion of correct pose predictions. A pose prediction is considered correct if the ADD(-S) distance is less than the threshold of 10% of the model’s diameter. (2) ADD(-S) AUC: the area under ADD(-S) accuracy-threshold curve where the maximum threshold is set to 10cm.
Experiments
We implement and evaluate two methods on StereOBJ-1M dataset as baselines for future experiments. Specifically, we implement PVNet and KeyPose , two classic keypoint-based 6D pose estimation frameworks that have achieved state-of-the-art performance on various datasets. We train the two models on the merged train and validation sets and report their performance on the test set.
PVNet is a single-RGB keypoint-based method. It represents keypoints using a 2D direction field and estimates the 2D locations of the keypoints by RANSAC-based voting scheme . The 6D poses are determined by solving a Perspective-n-Point (PnP) problem .
KeyPose is a stereo-RGB keypoint-based method. Different from PVNet, it localizes object keypoint by predicting heatmaps in both stereo images. The 6D object poses are calculated by keypoint triangulation from two-view stereo and Orthogonal Procrustes pose fitting.
2 Monocular Image Experiments
We conduct monocular image experiments where only the left images are used as input to predict the 6D pose. The stereo method KeyPose is adapted to its monocular variant where only keypoints in the left stereo images are predicted using heatmaps, and the 6D poses are calculated by solving PnP problem . The results are illustrated in columns 1-2 in Tables 3 and 4. KeyPose and PVNet respectively achieve 36.03% and 36.92% in average ADD(-S) AUC, and 25.88% and 24.16% in average ADD(-S) accuracy. Among the objects, the performance of both baseline methods suffers especially on pipette categories, which highlights the challenge of pose estimation of visually similar but different object instances.
3 Stereo Image Experiments
Classic Triangulation. Given and for all , a naïve method to compute object 6D pose is to follow classic point-level triangulation used in KeyPose , i.e. triangulate 3D keypoints from stereo and fit them to canonical object 3D keypoints by solving an Orthogonal Procrustes problem, to obtain the estimated pose :
Object Triangulation. We propose a novel object-level triangulation approach as a stronger baseline. Compared to classic triangulation which optimizes the 3D location of a point, we directly optimize the 6D pose of an object from 2D keypoint predictions in both images. Mathematically, Object Triangulation combines the two steps in Equation (3) into one unified optimization of the 6D pose :
We use the Levenberg-Marquardt algorithm as the non-linear optimization method together with RANSAC . The results of the two baseline architectures with two pose optimization methods are illustrated in columns 3-6 in Tables 3 and 4. Baseline methods with Object Triangulation consistently improve over monocular variants on all object categories significantly while classic triangulation can yield worse results. With Object Triangulation, the stereo variants of both baseline methods significantly outperform their monocular counterparts on StereOBJ-1M, by at least 24% in ADD(-S) AUC and 13% in ADD(-S) accuracy.
Conclusions
In this work, we propose a novel object pose data capturing and annotation pipeline and present a large-scale object pose dataset with stereo RGB as input. We benchmark two state-of-the-art algorithms for 6D object pose estimation and propose a novel method for stereo object pose optimization that outperforms classic triangulation method. In addition to pose estimation, our dataset enables future research directions such as object reconstruction and scene flow estimation from stereo RGB.
Acknowledgement. This work is funded in part by JST AIP Acceleration, Grant Number JPMJCR20U1, Japan.
References
Appendix A Overview
In this document, we provide additional details on StereOBJ-1M dataset as presented in the main paper. We present additional baseline results on instance-level pose detection for centrifuge_tube class in Section B. In Section C, we provide details on the hardware of data capturing. In Section D, we provide more details on viewpoint distribution of each object class. Lastly, in Section E, we visualize more data samples from our dataset.
Appendix B Multi-instance Pose Detection Results
In the main paper, we report the results of two baselines on single-object pose estimation of 17 out of 18 objects on the test set where there is at most one object instance from a category in a scene. However, for centrifuge_tube, there are usually multiple instances recorded in a scene. Therefore, centrifuge_tube is used in multi-object pose detection task. In this task, the framework is supposed to perform instance-level detection and pose estimation simultaneously.
To adapt to instance-level pose detection, we modify the baseline formulation by introducing additional 2D object detection before pose estimation. Given a detected rough 2D bounding box of an object instance, we crop the image patch and send it to pose estimation baselines, i.e. PVNet and KeyPose , to estimate the 2D keypoint locations and therefore 6D pose of that object instance. The 2D object detector we used is Faster-RCNN .
We use Average Precision (AP) as the evaluation metrics of multi-instance pose detection. When calculating AP in 2D object detection, a detection result is considered correct if the IoU between the detected bounding box and a ground truth bounding box is larger than a threshold. Different from 2D object detection, we consider a pose detection result to be correct if the ADD(-S) distance between the detected 6D pose and a ground truth pose is smaller than a threshold. We use 10% of the object diameter as the threshold of ADD(-S). We report the pose detection results with single-RGB image as input in Table 5. We notice that the above baseline suffers when two or multiple instances object overlap in the image and are included in the same image patch. In this case, the pose estimation framework cannot distinguish different instances and fails in keypoint prediction.
Appendix C Data Capturing Hardware
We present the hardware used for capturing the data in Figure 7, including a large fiducial marker board, several small fiducial markers, two static cameras with two tripods, and one moving stereo camera. We used the same Weewiew stereo camera for all three cameras, though the two stereo cameras can be monocular. Weewiew stereo camera has a stereo baseline of approximately 4.5cm which is close to the distance between the two human eyes. All three cameras are calibrated.
The fiducial markers are the first 20 AprilTags . The large fiducial marker board is printed on a 20in 16in plastic picture frame. Though the large fiducial marker board needs to be accurately measured by its physical dimensions with a vernier caliper, the small fiducial markers do not.
Appendix D Viewpoint Coverage Distribution
Viewpoint coverage percentage is illustrated in Section 4.2 and Figure 5 of the main paper. We illustrate a more detailed viewpoint distribution for each object in Figure 8. The viewpoints are drawn as 3D points on the unit sphere centered at the object center. Their positions on the unit sphere are determined by the Azimuth and Elevation of the viewpoint. Their density on the sphere is shown by heatmap color. Notice that for objects such as microplate and tape_measure, there is no viewpoint distributed on the space, because there is only one possible side up when being put on a desktop.
Appendix E More Visualizations of Data Samples
We provide more visualizations of data samples from our dataset. As illustrated in Figure 9, the data annotation has high precision.
Appendix F Change Log
Mar 15, 2022: Updated dataset statistics and baseline performance results after cleaning up the data.