NeRF-Pose: A First-Reconstruct-Then-Regress Approach for Weakly-supervised 6D Object Pose Estimation
Fu Li, Hao Yu, Ivan Shugurov, Benjamin Busam, Shaowu Yang, Slobodan Ilic
Introduction
Several computer vision tasks, such as 2D object detection and semantic segmentation, have experienced tremendous progress in recent years thanks to the development of deep learning. However, 2D object detection alone is limited and insufficient for real-world applications such as Augmented Reality, Robotic Manipulation, Autonomous driving, etc., which often require the knowledge of the full 6 degrees of freedom (DoF) pose of the object in the scene. Therefore, the ability to recover the pose of the objects in 3D environments is essential for a better understanding of the 3D scene.
This fundamental 3D vision problem has been addressed by scholars for decades and is particularly difficult if the estimation is done only from a single RGB image. Recent research tackles the ill-posed nature of this problem in a data-driven way where the pose is usually computed with respect to a 3D CAD model. Most of the recent approaches require 6D pose labels as supervision signals. Moreover, most of the recent methods posed the pose estimation as a problem of correspondence estimation between a known 3D object model, such as a CAD model, and image pixels. However, it is hard and expensive to obtain accurate pose labels and fine-grained CAD models for all objects for real-world scenarios . On the other hand, synthetically generated images have the advantage of a potentially unlimited amount of labeled data. However, methods trained only on synthetic data have worse performance than their real counterparts due to the lack of realism. Moreover, precise textured CAD models are required to render synthetic data. We argue that it is much easier to obtain 2D image labels, such as segmentation masks, and relative camera poses. As widely used in the recent methods , segmentation mask can be obtained either manually or automatically with the off-the-shelf object segmentation approaches , few-shot segmentation, depth driven segmentation or background substaction, while relative camera poses can be obtained with Structure from Motion (SfM) , Simultaneous Localization And Mapping (SLAM) , Inertial Visual Odometry , or simply a marker board . Different from fully-supervised methods trained on all labels, e.g. CAD models, segmentation masks and 6D pose labels, the methods, that don’t use textured CAD models and 6D pose annotations, use weaker labels and, so, can be considered as weakly-supervised.
Our approach is to first recover implicit 3D object representation from training images containing weak labels: 2D segmentation masks and relative camera poses. Next, we use this implicit representation to supervise the regression of dense correspondences between training images and previously recovered object’s implicit 3D object representation. Therefore, we propose a first-reconstruct-then-regress training pipeline, named NeRF-Pose, which builds atop of the success of Neural Radiance Fields (NeRF) and its successors . We first reconstruct the object as a NeRF-based network trained with weak labels. Then, we train a pose regression network to regress the dense image pixel (2D)-object model (3D) correspondences.
As depicted in Fig. 1, during inference, we first detect the objects in a 2D image using an off-the-shelf 2D object detection network and then predict segmentation masks and dense correspondences represented in terms of Normalized Object Coordinates (NOCS) . With regressed correspondences and the NeRF object model, we propose a NeRF-enabled PnP+RANSAC method in order to compute the object pose in the end. Our key contributions can be summarized as follows:
A weakly-supervised object pose estimation approach, which is trained only with 2D annotations and relative camera poses, instead of relying on an explicit CAD model and accurate 6D pose labels.
OBJ-NeRF neural network encoding an implicit 3D object representation obtained from weak labels: segmentation masks and relative camera poses.
Pose regression network, which relies on the above object’s implicit NeRF-based representation.
NeRF-enabled PnP+RANSAC approach enabling highly accurate 6D pose computation.
An extension of the HomebrewedDB dataset containing real video sequences with weak labels (segmentation masks and relative camera poses).
We conduct experiments on LineMod (LM) and LineMod Occlusion (LMO) datasets, and extend the HomebrewedDB (HBD) dataset. Though training in weakly supervised settings, we have about 15% (LM) and 20% (LMO) improvement on ADD(-S) metric compared to the methods trained without CAD models but with accurate pose labels. We also achieve comparable results on LM, LMO, and HBD datasets in comparison to fully supervised methods. The experiments show that our weakly-supervised NeRF-Pose approach achieves accurate and robust object pose estimation.
Related Work
The first type of pose estimation methods that gained popularity in recent years are dense correspondence-based methods . While being different in implementation, their common denominator is the key idea to train a neural network to predict 2D-3D correspondences between each object pixel in the image and the 3D location of the corresponding point on the object’s surface. Those correspondences are consecutively used either with PnP+RANSAC or the Umeyama algorithm to compute the 6D object pose. DPOD is proposed to use discrete UV maps to uniquely parameterize the object surface. With this parameterization, the UNet-like network predicts two discrete UV coordinates for each visible pixel occupied by the object. Pix2Pose and CDPN leverage two-stage detectors and use 3D normalized vertex coordinates to parameterize the correspondences. Correspondences in CDPN are only used to estimate rotation, while translation there is predicted directly. NOCS uses Mask-RCNN and a normalized object coordinate space to predict correspondences. However, NOCS mainly focuses on estimating the scale and 6D pose for unseen objects. EPOS extends the idea of dense correspondences by parameterizing each 2D-3D correspondence with its location within a discrete object fragment. Multi-fragment perspectives and many-to-many 2D-3D correspondences enable it to handle symmetric objects more effectively.
Furthermore, CosyPose and Self6D use a similar pose parameterization which allows for direct pose prediction. They predict the pose by estimating the 2D center location of the object and its distance to camera center. Combined with the intrinsic parameters of the camera, it gives the estimate of the translation component of the 6D pose. Allocentric rotation parameterization is used for simpler rotation prediction. Self6D is first trained on synthetic data and then fine-tuned on real data without pose annotations in a self-supervised manner. AAE , leverages manifold learning to retrieve a descriptor of a given object patch from a pre-computed database consisting of descriptors of the same object under various rotation. The translation is estimated based on the object scale in the image. MHP predicts multiple pose hypothesis to estimate the pose of symmetric or occluded objects. GDR-Net and SO-Pose use the combination of dense correspondences and direct pose estimation by first predicting the dense correspondences and then performing the learning-based pose prediction, which they name patch-pnp. And, SO-Pose introduces self-occlusion information to make the correspondences more stable.
Methods above are all fully supervised with the accurate pose labels and object CAD models available. The following two methods relax the constraints and could be trained without object CAD models. LieNet directly regresses the pose with the Mask-RCNN as backbone network. With known object pose labels and camera intrinsics, Cai et al. supervise the coordinate prediction network with the multi-view consistency, minimizing re-projection error across different views. Limited by the accuracy of the coordinate prediction network, the reprojection error is insufficient to guide reliable correspondence learning. LatentFusion trains the implicit neural object representation, which, at reference, takes the multi-view RGBD images of well-calibrated unseen objects as input and reconstructs the object model. Then, the render-refine network is used to estimate the object pose from RGBD input iteratively. Bundle-SDF performs 6D tracking and reconstructs an object without assuming a CAD model from a video. WeLSA generates pose labels for weakly labeled data by training on very few labeled data using shape alignment and feature alignment. Although WeLSA works with very few labeled data, they employ depth maps for training the pipeline and estimating pose labels for weakly labeled data.
Concerning weakly supervised methods, there are few works related to ours. On the basis of NeRF, i-NeRF presents a differentiable camera pose refinement method, which treats the camera pose as network parameters and iteratively updates the pose by minimizing the discrepancy between the input images and the rendered outputs in particular views. However, because of the high computational burden, it takes about half a minute to process one image and is very sensitive to the pose initialization. BARF and NeRF– introduce the methods to estimate camera poses and train NeRF concurrently. It inspires us to optimize the object pose, while simultaneously reconstructing the object with known relative poses.
Different to the model-free approach , which optimizes the network via minimizing the re-projection error, we instead aim at densely generating accurate 2D-3D correspondences between input images and an implicit NeRF object representation obtained from weak labels, yielding better performance.
Methods
In this section, we present NeRF-Pose for 6D pose estimation, which only requires weak supervision. We assume that real images with 2D ground truth segmentation masks and relative camera poses are available during training. We first present OBJ-NeRF, an implicit 3D model representation, learned under above-defined constraints. This representation is then used to generate object correspondence maps, which is afterwards used to train our proposed pose regression network. Finally, the regressed correspondences are used in a novel NeRF-enabled PnP+RANSAC algorithm for iterative pose estimation.
NeRF and its followups recover the 3D scene from multiple views with known camera poses. Since we deal with the problem of 6D object pose estimation, our aim is to compute object-centric NeRF and use it as an implicit 3D model representation for object pose estimation. Thus, we propose to modify the original NeRF approach. We take as input the images, segmentation masks and relative camera poses, and output implicit object-specific NeRF representation, named OBJ-NeRF. As OBJ-NeRF reconstruction relies only on relative camera poses, it will produce a 3D model in some uncertain coordinate systems. In order to make use of OBJ-NeRF for the purpose of 6D pose estimation based on dense 2D-3D correspondences it is necessary to recover it with respect to some reference coordinate systems. To achieve this, simultaneously with the NeRF reconstruction, we propose to regress poses of the NeRF reconstructed object with respect to some chosen reference frames. These estimated poses will be used later for the generation of correspondence maps needed for training the correspondence estimation network.
Based on the NeRF scene representation, OBJ-NeRF reconstructs the object from multiple images, and optimizes the reference object pose at the same time. It takes images , object masks , camera intrinsics , relative poses as training input and outputs the rendered object view , its mask , correspondence map as well as the reference object pose in respect to the reference image . As shown in stage 1 in Fig. 2, the steps of the full procedure, from input images to rendering the outputs, can be named as Sample points, Transform points, Calculate, and Render.
Transform points. Different from scene-centric BARF and iNeRF , which optimize absolute camera poses in some arbitrary coordinate systems, we perform object-centric reconstruction, which instead estimates the object poses in the camera coordinate system. Therefore, the sampled points from camera coordinate system are transformed to the object coordinate system according to the object pose with respect to the image . However, poses of image objects are unavailable and thus need to be estimated. Notably, owing to our weakly-supervised settings, where relative camera poses from to , denoted by , are accessible, only needs to be optimized. Thus, arbitrary poses can then be computed by . We further define the transformation functions: and , which transform the coordinate and view direction into unified object-centric coordinates, with the parameters to be optimized.
Calculate. Until now, we have obtained the transformed ray points and their directions . Then, we feed into the original NeRF network to get color and volume density prediction . In summary, our OBJ-NeRF network can be formulated as:
where and are the network parameters.
Render. Akin to NeRF, we use volume rendering approach to map a set of calculated data to the image plane along the rays. The volume rendering function sums up the product of transmittance and alpha value of sampled points along the ray, which is differentiable. Following , the rendered color , mask and coordinates can be formulated as:
with , and , where is the sampling distance between sampled adjacent points.
Constrain object pose. With only 2D bounding box and segmentation, where object center information is unavailable, the object canonical pose center has no constraint. Accordingly, we constrain the object center by projecting the object center to each view. More specifically, we minimize the reprojection error between the projected object center and 2D bounding box center on the image plane by minimizing .
where iterates the training image views, and u iterates the image pixels.
In the end, the newly defined object canonical pose is determined by optimizing the object pose in multi-view settings. With the learned canonical pose and neural model, the NOCS map ground truth can be rendered for the next pose estimation stage.
2 Pose Estimation
The overall three-step pipeline is shown in Stage 2 in Fig. 2. The first two steps depend on separately trained convolutional neural networks. The third step is purely optimization-based and does not require training. The first step represents the off-the-shelf 2D detector trained on ground truth crops of the objects of interest. In practice, we use YOLOv3 . In the second step, a pose regression network is trained on images to predict object coordinates and segmentation mask . The third step is our NeRF-enabled PnP+RANSAC algorithm, to improve the performance of pose calculation by introducing the NeRF-Mask renderer.
Coordinates regression. Our pose regression network is inspired by DPoD and CDPN , the state-of-the-art dense correspondence-based methods for indirect pose estimation. We use ResNet as encoder backbone, and the decoder contains four upsampling layers. Our pose regression network outputs the predicted segmentation and the object coordinates , which encodes correspondences between input image pixels and the OBJ-NeRF 3D model representation.
The loss function is defined similarly to other depth regression works . Our loss function is composed of the segmentation loss and the normalized coordinate regression loss:
The loss is defined by the mean squared error (MSE) between predicted mask and ground truth mask . is defined between the predicted object NOCS map and their ground truth rendered from the implicit neural network OBJ-NeRF as given below:
The first term is the element-wise loss that measures the average distance between these two coordinates written as:
and are used to penalize the coordinate errors in first and second order. Similar to the loss of , they are defined as:
where and are the gradients for the coordinates along and axis, defines the surface normal vector and measures the cosine similarity of the two vectors.
NeRF-enabled PnP+RANSAC. With the predicted dense 2D-3D correspondences, PnP+RANSAC-based method are typically used for estimating object pose . As shown in Fig. 3, the PnP+RANSAC algorithm iteratively selects a minimum number of correspondences for pose estimation and calculates object pose using PnP method. The pose hypothesis supporting the most inliers is selected as the calculated pose results.
The criteria of pose selection in PnP+RANSAC is inlier ratio , which highly depends on the correspondence quality. To make PnP+RANSAC more stable, we add two extra pose selection criteria which are independent to the correspondences. As shown in Fig. 3, with the calculated pose hypothesis and our implicit object representation, we can render the object mask on the image plane without occlusion. With the mask predicted from our pose regression network, we add the Overlap Recall and Overlap Precision to the score which is defined as: and . The final score is defined as .
Following Plenoctree , we convert the well-trained OBJ-NeRF model to an octree-based data structure, namely octree-NeRF, which stores the average value of sampled in the leaf voxel of octree. The NeRF calculation is simplified by the octree indexing, which significantly accelerates the mask rendering.
Experiments
In this section, we conduct extensive experiments to demonstrate that our proposed weakly-supervised pose estimation method, NeRF-Pose, produces highly-accurate 6D object pose estimation.
Datasets. We conduct our experiments on three publicly available datasets: Linemod (LM), Linemod-Occlusion (LMO), T-Less and HomebrewedDB (HBD). The LM dataset is a standard benchmark for 6D object pose estimation of textureless objects. It offers 13 objects of various sizes in the scenes with large background clutter but with almost no occlusion. The LMO dataset consists of 8 objects from LM dataset but provides more challenging test data with more occlusion. According to our problem settings, we train our model on the real data in LM and strictly follow the training/testing split proposed in . On the LMO dataset, to make a fair comparison to Cai et al. , we train our method with the real training data of LM. Since the HBD dataset has no real training data, we provide an Real data EXTtension of the HBD dataset, named HBD-REXT. It consists of real images of each HBD object presented in BOP challenge. These images are captured by the Azure Kinect camera with only relative poses and segmentation mask provided. We will make the HBD-REXT dataset publicly available soon.
Evaluation Metrics. On the LM and LMO datasets, we report the standard ADD(-S) metrics with the 10% diameter threshold, as it is the most prevalent pose quality metric for these two datasets. The metric measures whether the rotation error is less than and the translation error is below cm. Besides, we also use and , which measure whether the rotation error is less than and the translation error is below cm, respectively. Moreover, following the setup of the BOP challenge, which aims at unifying the evaluation of 6D pose estimation methods, we report the following metrics for the HBD dataset: Visible Surface Discrepancy (VSD) , Maximum Symmetry-Aware Surface Distance (MSSD) , Maximum Symmetry-Aware Projection Distance (MSPD) and Average Recall, which is computed as: .
Since we predict the poses in a canonical orientation, it is necessary to transform them to absolute poses. This is done by transforming them with the offset poses (obtained from the g.t. poses), which represent the transformation from absolute poses to the canonical poses. Note that this transformation is only leveraged for evaluation.
2 Comparisons to the State-of-the-art
LM. We train our method on real images, strictly following the training/testing split from . In order to make a fair comparison to Cai at al. , apart from training using relative poses(Our-weak), we also train our method using pose labels(Our-pose). In Tab. 1, our method outperforms the method of Cai at al. by a large margin () on a ADD(-S) metric. Moreover, our method with weakly-supervised settings performs on-par with the SoTA fully-supervised methods. We also perform an ablation, (Ours-sam) by using SegmentAnything to generate segmentation masks using ground truth bounding boxes as input instead of using ground truth segmentation masks. The mild accuracy drop(3.5%) indicates that our approach can be applied in the real world much more easily compared to other approaches using relative poses from the sensor or SFM and segmentation masks obtained using SAM.
LMO. We compare our method with SoTA methods in terms of ADD(-S) on LMO. Compared to Cai et al. trained with pose labels but without CAD models on real images, our method on the same settings achieves 49.2% on mean ADD score, surpassing it with a large margin of . In our setting, the CAD models are not accessible, so we do not train our method in self-generated synthetic images or pbr images pulished in . It is proved that training with more synthetic images can improve the model performance and pbr images can improve more. So, it is reasonable to claim that we stays comparable with the SoTA fully-supervised methods (GDR-Net : 53.0% and SO-Pose : 54.3%) trained on real and author-generated synthetic images.
In Tab.1 and Tab.2, Our-weak(uses OBJ-NeRF with relative poses) performs better than Our-pose (uses OBJ-NeRF with absolute poses). This is caused by noisy labels in LM, which harm more Our-pose than Our-weak. This is because in Our-weak we optimizes all pose labels by minimizing the reprojection and rendering errors constrained with relative poses. In Our-pose with absolute poses we cannot refine them, since we cannot guarantee to obtain the correct object’s scale.
HBD-REXT. To support weakly-supervised training on real images, we extend the HomebrewedDB training set by capturing more real sequences for all objects used in BOP challenge. We provide about 300 real images for each object with segmentation masks and relative camera poses generated from the markerboard. We train our model on this newly captured data. Since no other method was trained only with relative poses and 2D segmentation, we cannot run other methods using this new weakly-supervised data. In Tab. 3 we report the AR of VSD, MSSD, MSPD metrics on the BOP challenge test set. It illustrates that we stay on-par with the methods trained on synthetic images and fall behind the methods trained on pbr images . Notably, neither CAD models nor ground truth poses are used in the our-weak case, whose results are shown in the last column of Tab. 3.
T-Less. We evaluate our pipeline on T-Less dataset. The T-Less dataset comprises 30 objects with real training images. We train our model using relative camera poses and real training images. In Tab. 4 we report the AR of VSD, MSSD, MSPD metrics on the BOP challenge test set. We achieve closer to benchmark accuracy despite not using a CAD model. It shows that Nerf can learn accurate geometry and render correspondences which are usually extracted from the CAD model. SurfEmb performs better than our approach as their approach is tailored for symmetric objects and also employs an inference pipeline with 2.2s. However, the results compared to other regression-based, Dpod and DpodV2, show that our approach can perform equally better employing NeRF.
3 Ablations
Shape Analysis. The visual results of the objects represented in OBJ-NeRF can be found in Fig. 4(c, g). Compared to their reference CAD models (Fig 4(a)), our OBJ-NeRF keeps both the detailed shapes and the texture information. Accurate shapes obtained in our model promises the accuracy of our consecutive pose estimation. We also observe some minor defects on our predicted shapes, e.g. details of cat eyes and cow legs, which is mainly caused by the imprecise relative pose annotations and imperfect segmentation masks.
View Numbers. To figure out how the number of views used in NeRF reconstruction affects the pose estimation, we train OBJ-NeRF with 4 different numbers of views—32(B1), 64(B2), 128(B3) and 156(Baseline). From the results in Tab. 5, training using 156 views is slightly better than the others. It actually shows small difference (around ) on ADD metric and almost no influence of the number of views to the results. Thus, we use 156 views for training OBJ-NeRF. Moreover, the results also illustrate the number of views does not bring significant influence on the performance, i.e., the OBJ-NeRF can recover the highly precise object model, though trained with fewer views. It is an interesting conclusion that can inspire us to use few-shot labels for reconstruction and pose estimation.
Regression Loss. We study the contribution of our pose regression loss. As shown in Tab. 5 C1-C3, the loss with gradient and normal loss has the best performance, which indicates the efficiency of our utilized loss functions.
Training with occlusions. To further evaluate the influence of occlusions on the performance of OBJ-NeRF, we add random rectangles onto RGB images to mask out certain areas and simulate occlusions and errors in segmentation masks. We test the performance on the same objects as in other ablation studies. The reconstruction computed from the distorted images (Fig.5(d)) has no clear difference to the non-distorted one (Fig.5(e)). However, in Tab 5, training with occlusion(D1) has a slight drop (about on ADD score) on pose estimation performance. We consider it acceptable, as the occlusion increases the difficulty of the reconstruction. Moreover, as indicated by the green ellipses in Fig.5, the g.t. masks and pose labels from the LM and HBD are actually noisy. Due to these imprecision, our reconstructions are missing some details. However, all our experiments demonstrate that such reconstructed implicit object models and small pose errors can be tolerated.
NeRF-enabled PnP+RANSAC. Furthermore, we evaluate the effectiveness of our proposed correspondence solver. We present the results in Tab. 5, which shows about 2% improvement on LM dataset on ADD score (E1). Though on metric, original PnP+RANSAC is slightly better than ours, our method outperforms it on cm metric. On ADD score, our method shows superiority over the original PnP+RANSAC. The main reason is that our NeRF-enabled RANSAC incorporates mask-based Recall and Precision in scoring pose hypothesis, as illustrated in Fig. 3. Since Recall and Precision are more sensitive to translation, our NeRF enabled RANSAC prefers pose hypothesis with less translation error, leading a better performance on both ADD and 2 cm metrics that reflect translation performance. In Tab. 5 E2-E3, we also report the ADD(-S) score on LMO dataset. We observe an average 3% improvement using our proposed NeRF-enabled method. The results show the superiority of our proposed NeRF-enabled PnP-RANSAC algorithm over the original one. LMO dataset contains more occlusion data than LM dataset. The occlusions can make mask scores used in NeRF-PnP-Ransac imprecise. In that case overlap recall and precision can deteriorate. However, since we weight the inlier score the most with vs. for precision and recall scores, and since our inliers are quite accurate we still get better performance than standard PnP-RANSAC as shown in the Tab.5 for the LM-O benchmark.
4 Implementation and Runtime Analysis
We implement our object-centric NeRF network based on the original version of NeRF in, and train the network from scratch. For the 2D detector, we use the standard YOLOv3 detector in stage two. An ImageNet pre-trained ResNet34 network is leveraged as the backbone of our pose regression network. All networks are trained until convergence.
The training is done on a machine with a Titan RTX GPU with 24GB Memory, an Intel(R) i7-8700K CPU and 24GB RAM. During inference, for a single image with resolution, our approach takes about 0.25s for one object, including about 0.03s for YOLOV3 2D detector, 0.01s for pose regression and 0.21s for our pose solver.
Limitations
Though we present NeRF-Pose in a weakly-supervised way, considering the training difference in OBJ-NeRF network and our pose regression net, failing to enable end-to-end optimization sometimes leads to local minima.
As indicated in Tab. 2, when trained on pbr images rendered by Blender with high quality from BOP , GDR and SO-Pose gain about 10% improvement on ADD(-S) metric. Those fully-supervised methods benefit from images that cover more poses and have more realistic occlusion under various light conditions. It inspires us to generate more synthetic training data using our well-trained OBJ-NeRF for better performance.
Conclusion
In this paper, we propose NeRF-Pose, a first-reconstruct-then-regress approach for weakly-supervised object pose estimation. NeRF-Pose first implicitly reconstructs the object as the proposed neural network, namely OBJ-NeRF, from the weak labels and generates the signals to supervise the correspondences predicted from our pose regression network. At inference, a NeRF-enabled PnP+RANSAC algorithm is used to estimate the pose from the predicted correspondences. Finally, A thorough evaluation on LineMod, LineMod-Occlusion, T-Less and Homebrewed DB datasets show our leading performance on the task of weakly-supervised object pose estimation.