TransCG: A Large-Scale Real-World Dataset for Transparent Object Depth Completion and a Grasping Baseline

Hongjie Fang, Hao-Shu Fang, Sheng Xu, Cewu Lu

I Introduction

Transparent materials are widely used in modern industry, and robots inevitably need to process transparent objects no matter in manufacturing, logistics, or household services. Recently, much progress has been made in the field of robot grasping and manipulation . However, many of the advances are not directly applicable in scenes with transparent objects since these methods heavily rely on the depth information collected by the RGB-D cameras, yet ordinary depth sensors usually fail to construct a complete depth image in scenes that include transparent objects. The physical properties of transparent objects would lead to the distortion of light path by reflection and refraction, resulting in noisy depth maps. Therefore, many depth-based algorithms are incapable to handle transparent objects such as plastic bottles and glass containers which can be found everywhere in our daily life.

The geometry estimation of transparent objects remains a challenging task in the computer vision field, though progress has been made by many researchers. Ba et al. utilized a special polarization camera to leverage polarization cues for shape estimation and reached satisfactory results, while Li et al. proposed a two-stage physical-based network to reconstruct the shape of the transparent objects using multi-view images and material prior. However, both methods require specialized hardware, which is not a general setting for robotic manipulation. A more common setting is a robot arm with an RGB-D camera, which is the setting that we mainly focus on.

Under this circumstance, Sajjan et al. adapt the depth completion pipeline to scenes that contain transparent objects and then propose ClearGrasp, which predicts the surface normal and the transparent boundary, followed by the global optimization to solve the depth estimation. A synthetic dataset and a small real-world dataset are also proposed along with the method. Zhu et al. present an end-to-end framework for depth completion using the local implicit depth function, along with a synthetic Omniverse Object dataset. Both synthetic datasets provide images containing transparent objects and their ground-truth depth maps, but the lack of real depth maps degrades the performance of those methods in real-world applications inevitably.

To close the synthetic-to-real gap in the field of grasping concerning transparent objects, we propose TransCG, a large-scale real-world dataset for transparent object depth completion. A novel semi-automatic pipeline is proposed to accelerate the data collection and annotation process. In total, our dataset contains 57,715 RGB-D images of 51 transparent objects and around 200 opaque objects captured from different perspectives of 130 scenes under real-world settings. The 3D mesh model of the transparent objects are also provided in our dataset. The methodology for building our dataset is shown in Fig. 1.

Futhermore, we propose a robust, efficient and effective network Depth Filler Net (DFNet) for depth completion based on our dataset, which allows it to assist 6-DoF grasping methods on transparent object grasping. The quantitative and qualitative results show that our proposed DFNet (1) generalizes better across different scenes compared to existing methods; (2) improves most of the metrics to a great extent; (3) yields the highest inference speed and consumes the least computation overhead. We also apply our network to real-world object grasping for novel transparent objects, and a promising performance is witnessed. The full dataset, source code and pretrained models are released at www.graspnet.net/transcg.

II Related Works

Due to the special optical properties of transparent objects which leads to undetermined and inaccurate results of optical sensors, the datasets concerning of transparent objects is usually difficult to build.

For tasks that do not need the depth information, e.g., transparent object classification, segmentation and pose estimation, there are large-scale datasets such as Trans10K-V2 , Stanford2D-3D and StereOBJ-1M . For datasets that needs accurate sensor information as ground-truth, the common solutions are synthetic datasets, such as ClearGrasp synthetic dataset and Omniverse object dataset . Built by tools like SuperCaustics , the greatest shortcoming to those datasets is that the raw information collected by sensors is unlikely to be obtained in simulation. Xu et al. propose Toronto Transparent Object Depth Dataset (TODD), another real-world depth completion dataset. However, the diversities of transparent objects, translucent objects and cameras in the dataset are limited. Liu et al. build a real-world keypoint estimation dataset for transparent objects, which consists of the ground-truth depth map generated by substituting the transparent object with the identical opaque objects. But the large amount of real-world samples are all captured alone under certain environments, which is rare in real-world applications. The dataset collection approach is also used in ClearGrasp real-world dataset which only has 286 samples since generating the missing information is both time-consuming and labor-consuming.

To overcome the difficulties of building datasets concerning transparent objects, we propose a novel pipeline for transparent dataset construction. Using the pipeline, we build TransCG, a complete large-scale real-world dataset for transparent object depth completion. Details will be introduced in Section III.

II-B Depth Completion and Estimation

Depth completion and estimation has been studied by many researchers for a long time, and can be categorized into three classes: estimating depth directly from an RGB image , estimating depth from an RGB image with sparse depth information and depth completion from an RGB image with inaccurate depth information . Since depth information concerning reflective objects is usually noisy and inaccurate, transparent object depth completion falls into the last category, and it can be further classified into multi-view depth completion and single-view depth completion.

For multi-view depth completion, Ichnowski et al. use NeRF to recover depth for grasping transparent objects. Li et al. present a physical-based network, which uses multi-view images to recover high-quality 3D geometry of transparent objects.

For single-view depth completion, Zhang et al. propose a two-stage depth completion pipeline, which predicts surface normals and occlusion boundaries according to the RGB image, followed by the global optimization to complete the depth map. ClearGrasp makes a few critical modifications to the pipeline and adapts it in depth completion concerning transparent object, but its inference speed is unacceptable in real-time grasping scenarios. Tang et al. utilize self-attentive adversarial network to replace the global optimization in ClearGrasp. Zhu et al. propose a two-stage system consisting of local implicit depth function prediction and depth refinement. Though it outperforms previous works on speed and accuracy, its generalization ability is very limited according to our cross-domain tests in Sec. V-B, which makes it difficult to fit into real-world robotic manipulation settings. A concurrent work combines previous depth completion method with point cloud completion to improve the quality of the refined depth.

To achieve a better generalization and applicability in real-world environment, a tiny but robust model is needed, which requires a large amount of real-world data as support.

II-C 6-DoF Grasping

6-DoF grasping refers to those methods that predict the position and rotation of the gripper in 3D domain. Enabling the robots to grasp objects from various angles, it is the basis of robotic manipulation. Due to the versatility and effectiveness of 6-DoF grasping, it is currently the main stream in the grasping field.

GPD propose a two-stage 6-DoF grasping method, which estimates grasp candidates sampled under empirical constraints. PointNetGPD improves GPD by adapting PointNet in evaluation. Mousavian et al. leverage variational auto-encoder to sample grasp poses, and add refinement process after evaluation for better performance. Qin et al. regress the grasp pose directly from the partial-view point cloud, while Ni et al. regress the grasp pose from features extracted by PointNet++ . Fang et al. propose the GraspNet-1Billion dataset for general object grasping and an end-to-end grasp pose prediction network. Gou et al. incorporate RGB and depth information to improve the performance of 6-DoF grasping. Sundermeyer et al. reduce 6-DoF grasp poses to 4-DoF grasp representations on a known contact point to facilitate the learning problems and improve the grasping quality in cluttered scenes. Wang et al. propose a geometrically-based quality graspness to evaluate the graspable area in cluttered scenes.

All these methods rely heavily on depth image, which makes them unsuitable for transparent object grasping. Thus, utilizing color information to generate high-quality depth map and point cloud to aid the grasping deserves further exploration.

III Dataset

As introduced before, the previous transparent datasets that need accurate ground-truth depth usually use simulation platforms to generate synthetic data, and real-world datasets are usually small in scale due to the workload of annotations. To overcome the difficulties, we propose a novel pipeline to build the dataset efficiently, which utilizes a robot to perform data collection after limited object-level manual annotations.

In brief, we aim to collect a dataset that contains RGBD images using real-world sensors, along with detail annotations for the transparent objects include their depths, masks, 6D poses, normals, etc. To reduce the annotation effort, our methodology for building the dataset is to automatically localize the transparent objects during data collection. To achieve that, we resort to an optical tracker that can accurately localize a target’s 6D pose in real-time from several IR markers attached to it. Besides, we manage to obtain the 3D model of each transparent object in the training set by wrapping them with opaque materials and scanning with a 3D scanner. With these two preliminaries, we can easily obtain the transparent objects’ 6D pose during data collection. The geometry of transparent objects can be restored by leveraging the tracker results and their 3D models during data annotation. Our pipeline is able to build large-scale real-world dataset conveniently and reduce the human workload by a large extent.

Using the pipeline, we build our TransCG dataset, which contains 57,715 RGB-D images captured by two different cameras, along with the refined ground-truth depth images, transparent ground-truth mask and the surface normals, from 130 scenes under various background settings, within a week. We collect 51 common objects in daily life that can lead to inaccurate depth map, including transparent objects, translucent objects, reflective objects and objects with dense tiny holes. Apart from 65 simple isolated scenes that are similar to the scenes in the previous datasets, we also provide 65 challenging cluttered scenes that are closer to the real-world grasping environment, as shown in Fig. 2. More details are presented in supplementary video. The comparisons between our dataset and existing datasets are summarised in Tab. I.

III-B System Setup

To support fast and accurate data collection process, we build a transparent object tracking system, which is the only part that requires human efforts in our dataset building pipeline. The system setup process is illustrated in Fig. 3.

Given a transparent object, we firstly attach a fixer with IR markers to it, as shown in Fig. 3b. A PST optical trackerhttps://www.ps-tech.com/products-pst-base/ is used to record the IR markers and tracks them afterwards. Although we can directly attach IR markers to the object, we found that using a flat fixer can make the tracking more robust. Next, we temporarily wrap the transparent object with some opaque materials that can preserve the shape of objects, and obtain its 3D model with a Shining3D EinScan-SP scannerhttps://www.einscan.com/desktop-3d-scanners/einscan-sp/. With these two steps, we can obtain the transparent object’s 6D pose during data collection, where the details is further explained.

On average, the overall human efforts to process an object is around 1 hour. With the transparent object tracking system we introduced, we can easily recover the ground-truth depth information of the transparent objects afterwards.

III-C Data Collection

To collect a large amount of data automatically, we attach our tracking system to a robot arm that moves along a fixed trajectory containing 240 distinct viewpoints. Transparent objects are randomly selected and placed in the scene. To align with real-world setting, we augment the scene with various table covers and opaque objects. At each viewpoint, the tracking system will capture the RGBD images and the tracking results.

Thus, we can recover the 6D pose of the objects that are not tracked in some viewpoints using the tracking results in the first viewpoint, which are always successfully detected.

After collecting the raw data, we use the collected 6D poses to render the ground-truth depth maps, the transparent masks and ground-truth surface normals.

III-D Dataset Postprocessing

The blurry samples and improperly exposed samples are automatically detected and removed from the dataset using tools like Laplacian operators and histograms provided in OpenCV library . Manual validations are conducted by rendering the object mesh to the scene using the 6D pose and see if they can overlap with the original object. Finally, we generate a metadata containing all valid viewpoints for each scene.

III-D2 Dataset Split

We randomly select 12 objects from different categories and regard all scenes that contain these objects as the testing set, while the other scenes are used as the training set. There are totally 34,191 training samples and 23,524 testing samples.

III-E Discussion

To ensure that the markers have minimal influence to the object, we do not directly place them on the surface of transparent objects. Instead they are placed on the surface of a fixer that is attached to the object. Therefore, we can make sure that the transparent part of the object is retained to the greatest extent. From the cross-dataset experiment conducted in Table III and Table IV, we can see that methods trained on our dataset generalize well to other testing datasets. These experiments demonstrate that the markers have minimal impact on the training of network.

IV Method

IV-B Depth Completion

Inspired by previous literature about depth estimation from sparse sensing, we propose our end-to-end depth completion network Depth Filler Net (DFNet) that predicts full depth map according to RGB information and inaccurate partial depth. Adapting a U-Net architecture with depth of four layers, our network takes dense blocks as backbones and formulates Conv-Dense-Conv-Downsample (CDCD) blocks, Conv-Dense-Conv (CDC) blocks and Conv-Dense-Conv-Upsample (CDCU) blocks. Empirical statistics about depth estimations and completions show that original depth information is critical throughout the networks. Hence we provide the original depth information as an input to every CDCD, CDC and CDCU block. Skip paths are added to retain information in high resolutions. Inspired by AlphaPose , we adapt dense up-sampling convolution (DUC) instead of ordinary deconvolution layers in CDCU blocks.

Our network is trained using the following loss function:

where β\beta is the weight parameter, Ld\mathcal{L}_{d} penalizes depth inaccuracy and Ls\mathcal{L}_{s} is the cosine distance of surface normals computed from predicted depth map and ground-truth depth map , which penalizes unsmoothness. Formally,

where D^\hat{\mathcal{D}} and D∗\mathcal{D}^{*} denotes the predicted depth and the ground-truth depth, Dw\mathcal{D}_{w} and Dh\mathcal{D}_{h} are gradient vectors along width-axis and height-axis of depth map D\mathcal{D} respectively. For both losses, we regard depths out of range [0.3,1.5][0.3,1.5] as invalid pixels and remove them from losses to reduce the impact of outliers.

IV-C Object Grasping

Given an RGB image along with a depth image collected by an RGB-D camera, we first scale the images to an appropriate size and feed into our depth completion model DFNet, which outputs the refined depth in the same resolution as the input. Then, the refined depth is scaled back to the original size, which can be used to construct the scene point cloud using camera intrinsics. After that, the scene point cloud is sent to GraspNet-baseline as the input to the end-to-end grasp pose detection network. Other depth based grasping methods are also applicable. Finally, the grasp pose detection network outputs the grasp candidates, and the grasp will be executed by a parallel-jaw robot.

V Experiments

We compare our method with several representative approaches on our TransCG dataset. ClearGrasp is the first algorithm which leverages deep learning with synthetic training data to estimate depth information concerning transparent objects, and LIDF-Refine utilizes local implicit depth function to solve transparent object depth completion task. We also evaluate the concurrent work TranspareNet in experiments. All baselines are trained in our TransCG dataset using their released source codes and optimal hyper-parameters for fair comparisons.

For our model, the hidden channels in the network is set to 64. In every dense block, the layers LL and the feature channels of each layer kk are set to 5, 12 respectively, as suggested in literature . We use AdamW optimizer with initial learning rate of 10−310^{-3} and multi-step learning rate scheduler which decays the learning rate by 5 after 5, 15, 25, 35 epochs. We train the model for 40 epochs with the batch size of 32. Several data augmentation approaches such as random flipping, rotation, noise adding and chromatic transformations in HLS color space are conducted during training. Concerning loss, we set β=0.001\beta=0.001.

For all methods, we scale the images to 320×240320\times 240 during training and testing. We use 4 NVIDIA GeForce RTX 3090 GPUs for training and one for testing, and the average time to train an epoch is approximately 20 minutes.

The following common metrics of depth completion for transparent objects are used in comparisons. All metrics are calculated on the transparent areas according to transparent masks unless specified.

RMSE: the rooted mean squared error between depth estimates and ground-truth depths.

REL: the mean absolute relative difference.

MAE: the mean absolute error between depth estimates and ground-truth depths.

Threshold δ\delta: the percentage of pixels with predicted depths satisfying max⁡(d/d∗,d∗/d)<δ\max(d/d^{*},d^{*}/d)<\delta, where d,d∗d,d^{*} are corresponding pixels of D,D∗\mathcal{D},\mathcal{D}^{*}, and δ\delta is set to 1.05, 1.10 and 1.25 following .

The quantitative results are reported in Tab. II, and the qualitative results are visualized in Fig. 5. Our method outperforms ClearGrasp and LIDF-Refine on both metrics and quality, and performs better than the concurrent work TranspareNet on some metrics. Qualitatively, it completes the inaccurate region of depth maps more completely (Fig. 5) than . It is also worth noticing that though LIDF-Refine performs well on masked metrics, it may introduce shadows on areas without transparent objects, which may result from the outliers of the original depth map. Moreover, our method has the smallest size, fastest inference time and lowest GPU memory occupation, which allows it to perform depth completion of high quality under limited resources (Tab. II).

V-B Cross-Domain Experiments

Cross-domain experiments are performed to verify the robustness of our proposed depth completion method and the generalization ability of our proposed TransCG dataset.

For cross-domain experiments of different methods, two experiments are conducted, namely (1) test the performance on our TransCG dataset after training on ClearGrasp synthetic dataset and Omniverse object dataset; (2) test the performance on ClearGrasp real-world dataset after training on our TransCG dataset. Results shown in Tab. III reveal that our method is the most robust one among all methods. Also, it is worth noticing that although LIDF-Refine reaches satisfactory results when training domain is similar to the testing domain, the cross-domain testing decreases its performance a lot since its local implicit depth function is environment-dependent. On the contrary, our method is less sensitive to domain changes and is able to achieve satisfactory results under different environment settings.

For cross-domain experiements of different datasets, we select our method as the depth completion model to test the generality of the training dataset on a third-party dataset TOD and ClearGrasp real-world dataset. Results shown in Tab. IV demonstrate that our real-world dataset is more universal compared to the previous synthetic datasets, even though ClearGrasp real-world dataset has a similar environment as ClearGrasp synthetic dataset which is used for training. The performance differences also reflect the shortcomings of synthetic datasets compared to real-world datasets.

V-C Real Robot Experiments

To verify the performance of our method in real-world settings, we conduct real robot object grasping experiments. The object grasping pipeline incorporated with our method is introduced in Sec. IV-C. The experiments are conducted on a UR-5 robot with an Intel RealSense D435 camera and a Robotiq two-finger gripper, as shown in Fig. 6.

We randomly select 8 transparent objects to perform real-robot experiments, 6 of which are completely novel, and the rest 2 objects are the same objects without fixers and markers from our testing set. For every experiment, we randomly put the objects and repeat grasping until an object fails for 3 times. The success rate is defined as #objects#attempts\frac{\#\textrm{objects}}{\#\textrm{attempts}}, and the completion rate is defined as #successfully-grasped objects#objects\frac{\#\textrm{successfully-grasped objects}}{\#\textrm{objects}}. Table V reports the experiment results, which shows the effectiveness and feasibility of our method. More details are presented in supplementary video.

VI Conclusion

In this paper, we propose TransCG, the first large-scale real-world dataset for transparent object depth completion and grasping, built by our novel data collecting pipeline. Our dataset fills the synthetic-to-real gap in the transparent depth completion area and is more general compared to previous synthetic datasets in real-world environments. Moreover, we propose an end-to-end depth completion network DFNet, which is more efficient and robust compared to previous methods according to experiments on various datasets. Real robot experiments of grasping also demonstrate that our method is applicable in real-world settings with novel objects. The compatibility of our method with depth-based manipulation methods allows it to become a default pre-processing step for all downstream tasks concerning transparent objects.

References