Unseen Object Instance Segmentation for Robotic Environments

Christopher Xie, Yu Xiang, Arsalan Mousavian, Dieter Fox

I Introduction

For a robot to work in an unstructured environment, it must have the ability to recognize new objects that have not been seen before. Assuming every object in the environment has been modeled is infeasible and impractical. Recognizing unseen objects is a challenging perception task since the robot needs to learn the concept of “objects” and generalize it to unseen objects. Building such a robust object recognition module is valuable for robots interacting with objects, such as picking up unseen objects or learning to use new tools . A common environment in which manipulation tasks take place is on tabletops. Thus, we approach this by focusing on the problem of Unseen Object Instance Segmentation (UOIS), where the goal is to segment every arbitrary (and potentially unseen) object instance, in tabletop environments.

In order to ensure the generalization capability of the module to recognize unseen objects, we need to learn from data that contains large amounts of various objects. However large-scale datasets with this property do not exist. Since collecting a large dataset with manual annotations is expensive and time-consuming, it is appealing to utilize synthetic data for training, such as using the ShapeNet repository which contains thousands of 3D objects . However, there exists a domain gap between synthetic data and real-world data as many simulators do not provide realistic-looking images. Training directly on such synthetic data only usually does not work well in the real world , and synthesizing photo-realistic images with physics-based rendering can be computationally expensive , making large photorealistic synthetic datasets impractical to obtain.

Consequently, recent efforts in robot perception have been devoted to the problem of Sim-to-Real, where the goal is to transfer capabilities learned in simulation to real-world settings. For instance, some works have used domain adaptation techniques to bridge the gap when unlabeled real data is available . Domain randomization was proposed to diversify the rendering of synthetic data for training. While these techniques attempt to fix the discrepancy between synthetic and real-world RGB, models trained with synthetic depth have been shown to generalize reasonably well for simple settings such as bin-picking . However, in more complex settings, noisy depth sensors can limit the application of such methods and models trained on RGB have been shown to produce accurate masks . An ideal method should combine the generalization capability of training on synthetic depth and the ability to produce sharp masks by utilizing RGB.

In this work, we investigate how to utilize synthetic RGB-D images for UOIS in tabletop environments. We show that simply combining synthetic RGB images and synthetic depth images as inputs does not generalize well to the real world. To tackle this problem, we propose a two-stage network architecture called UOIS-Net that separately leverages the strengths of RGB and depth for UOIS. Our first stage is a Depth Seeding Network (DSN) that utilizes only depth to produce object instance center votes, which are then used to compute rough initial instance masks. We compare multiple architectures for the DSN that produce center votes in 2D and 3D. Training the DSN with depth images allows for better generalization to real-world data. However, the initial masks from the DSN may contain inaccurate object boundaries due to depth senor noise. In this case, exploiting textures in RGB images can significantly help.

Thus, our second stage is a Region Refinement Network (RRN) that takes an initial mask from the DSN and an RGB image as input and outputs a refined mask. Our surprising result is that, conditioned on initial masks, our RRN can be trained on non-photorealistic synthetic RGB images without adopting any of the afore-mentioned Sim-to-Real solutions. We posit that mask refinement is an easier problem than directly using RGB as input to produce masks, mainly because the mask refinement uses a local image patch as input and focuses on a single object. We empirically show robust generalization across many different objects in cluttered real-world data. In fact, our RRN works almost as well as if it were trained on real data. Our framework produces sharp and accurate masks even when the depth images are noisy. We show that it outperforms state-of-the-art methods including Mask R-CNN and PointGroup . Figure 1 illustrates our two-stage framework.

To train our method, we introduce a synthetic dataset of tabletop objects in home environments, which we name Tabletop Object Dataset (TOD). Our dataset consists of indoor scenes of random ShapeNet objects on random ShapeNet tables. We use the PyBullet physics simulator to generate the scenes and render depth and non-photorealistic RGB. Training our proposed method on this dataset results in state-of-the-art results on multiple real-world datasets for UOIS.

We extend our previous work, UOIS-Net-2D , by introducing a novel DSN architecture that reasons in 3D space. As we show in Section VIII, this architecture solves many limitations of producing these center votes in 2D (our previous work), thus providing stronger performance. Additionally, we propose a novel loss function that encourages separation of the center vote clusters and show that this loss function is crucial to achieving strong performance in cluttered environments.

II Related Works

2D semantic segmentation involves assigning pixels in an image to a set of known classes. Deep learning has emerged as the most popular tool for solving this problem . first introduced the concept of using a fully convolutional architecture (FCN). designed an architecture that utilizes dilated convolutions in order to increase the receptive field. further improves performance by introducing decoder architectures on top of the encoders. proposed a multi-path refinement network with long-range residual connections to enable high-resolution predictions. These methods have demonstrated strong performance on datasets such as PASCAL and COCO .

Much work has been devoted to solving the semantic segmentation problem in 3D as well. A common representation of 3D space is voxels; however operating with voxel grids as input can be expensive both computationally and memory-wise. Thus, introduced a submanifold sparse convolutional operator to preserve spatial sparsity of the input. further generalized these sparse convolutions to arbitrary kernel shapes, improving performance. Other 3D methods utilize point clouds. proposed PointNet, a permutation-invariant network architecture to handle point clouds, and extended this to a hierarchical network that recursively applies PointNet in order to obtain multi-resolution features, similar to deep convolutional networks with decreasing resolutions via strides/max pooling.

The advent of RGB-D sensors such as Kinect allowed the research community to utilize of both modalities for semantic segmentation, and drove the creation of datasets such as . investigated a combination of kernel descriptors, support vector machines, and Markov random fields for indoor scene segmentation. Both and leverage RGB-D videos, extracting 2D features from each RGB frame and integrating them with a reconstructed voxel representation of the scene. Deep learning-based approaches include . proposed the HHA encoding of depth images. The authors used this encoding to design an object detection system, which they further exploited to improve semantic segmentation performance. also used the HHA in their pioneering work on FCNs. proposed a depth-aware convolution and pooling mechanism to incorporate geometry into the convolution operators to build a depth-aware receptive field. used the output of a 2D segmentation network to initialize node features of a graph neural network applied on a 3D point cloud which was backprojected from a depth image. In this work, we also leverage RGB-D images, but focus on segmenting each individual object with unknown object class.

II-B Instance-level Object Segmentation

2D object instance segmentation is the problem of segmenting every object instance in an image. Many approaches for this problem involve top-down solutions that combine segmentation with object proposals in the form of bounding boxes , typically produced by a region proposal network (RPN). FCIS utilizes position-sensitive inside/outside score maps for fully end-to-end convolutional instance segmentation. Mask R-CNN , a prominent work in the field, predicts a foreground mask for each object proposal. builds on top of both FCIS and Mask R-CNN and exploits semantic segmentation and direction predictions to assemble foreground masks. proposes a module that iteratively refines segmentation predictions at adaptively selected locations, which can be used in conjunction with Mask R-CNN.

However, when bounding boxes contain multiple objects (e.g. cluttered robot manipulation setups), the true instance mask is ambiguous and these methods struggle. Recently, a few methods have investigated bottom-up methods which assign pixels to object instances . Other methods examine dense sliding-window instance segmentation on 4D tensors , combining top-down and bottom-up methods via blending modules , and alternative mask representations such as contours . Additionally, interactive instance segmentation has shown strong results with few user inputs .

Most of the afore-mentioned algorithms provide instance masks with category-level semantic labels, which do not generalize to unseen objects in novel categories. One approach to adapting these techniques to unseen objects is to employ “class-agnostic” training, which treats all object classes as one foreground category . One family of methods exploits motion cues with class-agnostic training in order to segment arbitrary moving objects . Another family of methods are class-agnostic object proposal algorithms . However, these methods will segment everything and require some post-processing method to select the masks of interest. jointly estimates instance segmentation masks and rigid scene flow, similar to . We also train our proposed method in a class-agnostic fashion, but instead focus our notion of unseen objects in particular environments such as tabletop settings.

In 3D instance segmentation, researchers have recently been investigating architectures to apply on point clouds/voxel grids. introduced the first deep learning method to fuse RGB and geometric information from RGB-D scans. proposed an occupancy term which greatly aids supervoxel clustering, leading to strong results. A few of these methods embrace center voting-based techniques . utilizes metric learning to learn abstract features, and predicts center votes which are post-processed by meanshift clustering. uses the center votes with a simple grouping mechanism to detect 3D bounding boxes of objects from point clouds only. Their follow up work incorporates RGB information by lifting 2D votes and features into 3D. follows a similar architecture but stacks a graph convolutional network to refine proposal features. also performs clustering on votes and semantic features for targeted performance on certain object classes. Our method takes inspiration from these voting-based methods, but is targeted to cluttered robot environments.

II-C Sim-to-Real Perception

Training a model on synthetic RGB and directly applying it to real data typically fails . Many methods employ some level of rendering randomization , including lighting conditions and textures. However, they typically assume specific object instances and/or known object models. Another family of methods employ domain adaptation to bridge the gap between simulated and real images . Algorithms trained on depth have been shown to generalize reasonably well for simple settings . However, noisy depth sensors can limit the application of such methods. Our proposed method is trained purely on (non-photorealistic) synthetic RGB-D data and is accurate even when depth sensors are inaccurate, and can be trained without adapting or randomizing the synthetic RGB.

III Method Overview

Given a single RGB-D image, the goal of our algorithm is to produce object instance segmentation masks for all objects on a tabletop, where the object instances (or even the semantic class) are arbitrary and are not assumed to have been seen during a training phase. These masks do not have any notion of class categorization or semantics. These masks can be employed by robots for interacting with unseen object instances in downstream applications such as grasping and/or manipulation.

Our framework consists of two separate networks that process Depth and RGB separately to produce instance segmentation masks. First, we design a Depth Seeding Network (DSN) that takes a depth image as input and outputs initial object instance segmentation masks. These initial masks can be quite noisy for a number of reasons, thus we design an Initial Mask Processor (IMP) to robustify them with standard image processing techniques. We further refine the processed initial masks using our Region Refinement Network (RRN), which is designed to snap the noisy initial mask edges to object edges in RGB, providing sharp and accurate final instance masks. The full architecture is shown in Figure 2.

Because the DSN incorporates non-differentiable techniques in order to build the initial masks, our DSN and RRN are trained separately as opposed to end-to-end. Both the DSN and RRN can be trained fully in simulation with no fine-tuning on real-world data, allowing our framework to capitalize on large amounts of simulated scenes and objects without resorting to the expensive process of annotating data. Our framework generalizes remarkably well to real-world scenarios despite being trained only on non-photorealistic simulated data, enabling robotic tasks with unseen objects.

IV Depth Seeding Network

We examine two methods of structuring the DSN. First, we investigate building initial masks by predicting centers in 2D pixel space. While this method provides state-of-the-art results, it has some obvious pitfalls (examined in Section VIII) that motivates a novel architecture that builds masks by predicting centers in 3D space.

In order to compute the initial segmentation masks from FF and VV, we design a Hough voting layer similar to . We describe the pseudocode detailed in Algorithm 1. First, we discretize the space of angles [0,2π][0,2\pi] into AA equally spaced bins. For every pixel, we compute the percentage of discretized directions from all other foreground pixels that point to it and use this as a score for how likely the pixel is an object center (lines 3-10). We threshold when a foreground pixel points to it with an inlier threshold and distance threshold. We then threshold the percentages and apply non-maximum suppression (NMS) to select object centers (line 11). Given these object centers, each pixel is assigned to the closest center it points to (line 12), which gives the initial masks as shown in the red box of Figure 2. Note that the inlier, distance, and percentage thresholds provide the Hough voting layer with robustness. For example, if not enough foreground pixels from all directions point towards a potential object center, that center is not selected. This robustifies the algorithm by protecting against false positives. We qualitatively show the efficacy of these design choices in Section VIII-E.

IV-A2 Loss Functions

To train the DSN, we apply two different loss functions on the semantic segmentation FF and the direction prediction VV.

We apply a weighted cosine similarity loss to the direction prediction VV. The cosine similarity is focused on the tabletop object pixels, but we also apply it to the background/tabletop pixels to have them point in a fixed direction to avoid false positives. The loss is given by

IV-B Reasoning in 3D

Reasoning in 2D has some failure cases that can be mitigated by reasoning in 3D. For example, if the center of an object is occluded by another object, the 2D center voting procedure will not detect that object (examples of this can be found in Section VIII-E). Thus, we propose a new architecture to the DSN to better handle these cases and provide stronger results. This formulation requires more sophisticated loss functions. In particular, we introduce a novel separation loss that significantly improves accuracy in cluttered scenes.

We propose to modify the architecture to use dilated convolutions in order to provide the DSN with a higher receptive field. We replace the 6th,8th6^{\textrm{th}},8^{\textrm{th}}, and 10th10^{\textrm{th}} convolution layers with ESP modules . An ESP module is a lightweight module consisting of a reduction operation, a split/transform component that applies convolutions with different dilation rates to get a spatial pyramid, and a merge process that hierarchically fuses the feature maps of the spatial pyramid . The ESP module has less parameters than the convolution layer it replaces, making it more computationally efficient. Details of the implementation can be found in the public code release at the project websitehttps://rse-lab.cs.washington.edu/projects/unseen-object-instance-segmentation/. We show in Section VIII-F that adding this module provides a boost in performance.

To compute initial masks, we perform mean shift clustering in 3D space over our center votes D+V′D+V^{\prime}. Mean shift clustering is an iterative procedure to find the modes of a distribution approximated by a kernel density esimate (KDE). The number of clusters (objects, in our case) is not determined beforehand, but instead by the number of modes in the KDE. We use the Gaussian kernel K(x,y)=exp⁡(1σ2∥x−y∥22)K(x,y)=\exp\left(\frac{1}{\sigma^{2}}\|x-y\|_{2}^{2}\right), which results in Gaussian mean shift (GMS) clustering. σ>0\sigma>0 is a hyperparameter which affects the number of modes (objects) in the KDE. Thus, the choice of σ\sigma is crucial and depends on the relative distance between objects, which is low in clutter. A detailed review of mean shift clustering algorithms can be found in . After clustering, each pixel is assigned to the cluster ID of its center vote to generate the initial masks. The clustering is only applied to the foreground pixels. Note that this method of producing initial masks lacks the thresholds such as ϵit,ϵd,ϵpt\epsilon_{it},\epsilon_{d},\epsilon_{pt} from the 2D DSN Hough voting layer that provide robustness.

IV-B2 Loss Functions

We apply four loss functions on the semantic segmentation FF and center offsets V′V^{\prime} to train the 3D-reasoning version of the Depth Seeding Network.

We apply a Huber loss ρ\rho (Smooth L1 loss) to the center offsets V′V^{\prime} to penalize the distance of the center votes to their corresponding ground truth object centers.

where wijw_{ij} are inverse proportional weights w.r.t. class size, d(⋅,⋅)d(\cdot,\cdot) is Euclidean distance, and [⋅]+=max(⋅,0)[\cdot]_{+}=\textrm{max}(\cdot,0). This loss function influences the KDE modes to be close to its points, and at least δ\delta away from points not belonging to the cluster, encouraging the points X=D+V′X=D+V^{\prime} to be more cluster-like.

We introduce a novel separation loss that encourages the center votes to not necessarily be at the center of an object, as long as it is far away from other object center votes in order to ease the post-processing GMS clustering phase. To do this, we consider the following tensor:

V Initial Mask Processing Module

Computing the initial masks from FF and V/V′V/V^{\prime} often results in noisy masks (see an example of initial masks computed from the 2D DSN using our Hough voting layer in Figure 2). For example, these instance masks often exhibit salt/pepper noise and erroneous holes near the object center (see Section VIII-E for examples). As shown in Section VIII-D, the RRN has trouble refining the masks when they are scattered as such. To robustify the algorithm, we propose to use two simple image processing techniques to clean the masks before refinement.

For a single instance mask, we first apply an opening operation, which consists of mask erosion followed by mask dilation , removing the salt/pepper noise issues. Next we apply a closing operation, which is dilation followed by erosion, which closes up small holes in the mask. Finally, we select the largest connected component and discard all other components. Note that these operations are applied to each instance mask separately. These simple image processing techniques are immensely helpful in robustifying the system.

VI Region Refinement Network

While depth generalizes reasonably well from Sim-to-Real, the initial masks (after IMP) are still subject to many errors due to noisy depth sensors. The RRN is designed to snap the initial mask edges to the object edges in RGB to provide accurate and sharp instance masks.

VI-B Mask Augmentation

Recall that the DSN and RRN are trained separately. In order to train the RRN, we need examples of perturbed instance masks. While we could train the RRN with the outputs of the DSN, we found that they are typically too clean on our synthetic dataset and we achieved better results by perturbing the ground truth masks instead. This problem can be seen as a data augmentation task where we augment the mask into something that resembles an initial mask (after the IMP). We detail the different augmentation techniques used below:

Translation/rotation: We translate the mask by sampling a displacement vector proportionally to the mask size. Rotation angles are sampled uniformly in [−10°,10°][-10\degree,10\degree].

Adding/cutting: For this augmentation, we choose a random part of the mask near the edge, and either remove it (cut) or copy it outside of the mask (add). This reflects the setting when the initial mask egregiously overflows from the object, or is only covering part of it.

Morphological operations: We randomly choose multiple iterations of either erosion or dilation of the mask. The erosion/dilation kernel size is set to be a percentage of the mask size, where the percentage is sampled from a beta distribution. This reflects inaccurate boundaries in the initial mask, e.g. due to noisy depth sensors.

Random ellipses: We sample the number of ellipses to add or remove in the mask from a Poisson distribution. For each ellipse, we sample both radii from a gamma distribution and a random rotation angle. This augmentation requires the RRN to learn to remove irrelevant blots outside of the object and close up small holes within it.

VII Tabletop Object Dataset

Many desired robot environment settings (e.g. kitchens, cabinets) lack large scale training data to train deep networks. To our knowledge, there is no large scale dataset for unseen tabletop objects. To remedy this, we generate our own synthetic dataset which we name the Tabletop Object Dataset (TOD). This dataset is comprised of 40k synthetic scenes of cluttered ShapeNet objects on a (ShapeNet) tabletop in SUNCG home environments . We only use ShapeNet tables that have convex tabletops and filter the ShapeNet object classes to roughly 25 classes of objects that could potentially be on a table. Example classes include: jar, mug, helmet, and pillow.

Each scene in the dataset is of a random room chosen from a random SUNCG house loaded without any furniture. We sample a ShapeNet table and scale it such that its height is in the range [0.75m, 1m], and place it in the room so that it is not colliding with walls or other fixtures in the room. Next, we randomly sample anywhere between 5 and 25 objects to put on the table, rescaling them such that they are not larger than 14min⁡{th,tl}\frac{1}{4}\min\left\{t_{h},t_{l}\right\} where th,tlt_{h},t_{l} are the height and length of the table, respectively. The objects are either randomly placed on the table, on top of another object (stacked), or generated at a random height and orientation above the table. We use PyBullet to simulate physics until the objects come to rest and remove any objects that fell off the table. Next, we generate seven views: one view is of only background, another is of just the table in the room, and the rest are taken with random camera viewpoints with the tabletop objects in view. The viewpoints are sampled at a height of between .5m and 1.2m above the table and rotated randomly with an angle in [−12°,12°][-12\degree,12\degree]. The images are generated at a resolution of 640×480640\times 480 with vertical field-of-view of 45 degrees. The segmentation has a tabletop (table plane only) label and instance labels for each object.

We show some example images of our dataset in Figure 3. The rightmost two examples show that some of our scenes are heavily cluttered. Note that the RGB looks non-photorealistic. In particular, PyBullet is unable to load textures of some ShapeNet objects (see gray objects in leftmost two images). PyBullet was built for reinforcement learning, not computer vision, thus its rendering capabilities are insufficient for photorealistic tasks . Despite this, our RRN can learn to snap masks to object boundaries from this synthetic dataset.

VIII Experiments

We evaluate our method on real datasets against the state-of-the-art (SOTA) methods Mask R-CNN and PointGroup . We denote our method as UOIS-Net-2D or UOIS-Net-3D, depending on whether the DSN reasons in either 2D or 3D.

All images have resolution H=480,W=640H=480,W=640.

We train all 2D DSN models for 100k iterations with stochastic gradient descent (SGD) using a fixed learning rate of 1e-2. We use a batch size of 8 and set λfg=λdir=1\lambda_{fg}=\lambda_{dir}=1. For the Hough voting layer, we use discretize the angles into A=100A=100 bins, and set ϵit=0.9,ϵd=20,ϵpt=0.5\epsilon_{it}=0.9,\epsilon_{d}=20,\epsilon_{pt}=0.5. We also process every 10th pixel instead of every pixel in line 4 of Algorithm 1 for computational efficiency.

During DSN training, we augment depth maps with multiplicative gamma noise similar to , and add Gaussian Process noise to the backprojected point clouds.

VIII-A2 RRN

All RRNs are trained with SGD for 100k iterations with a fixed learning rate of 1e-2 and batch size of 16. Inputs to the RRN are padded by 25% of the initial mask’s bounding box size in each dimension. Our entire pipeline runs at approximately 3-5 frames per second on a single NVIDIA RTX 2080ti.

VIII-A3 Baselines

For the baselines in , we use the results graciously provided by the authors. We follow the official Detectron schedule when training Mask R-CNN on TOD, and train for 100k iterations using SGD with a batch size of 8. We train PointGroup for 300k iterations with a batch size of 4 on TOD. We remove clustering of the semantic scores since our problem only has one meaningful semantic class, foreground, which makes no sense to cluster.

Note that the training procedure for 2D DSNs and Mask R-CNN (reported in ) and differs slightly from the training procedure for 3D DSNs. We replicated these experiments using the same conditions as 3D DSNs (150k iterations with Adam/SGD), but observed very similar results to . Thus, we report numbers from .

VIII-B Datasets

We evaluate quantitatively and qualitatively on two real-world datasets: OCID and OSD , which have 2346 images of semi-automatically constructed labels and 111 manually labeled images, respectively. OSD is a small dataset that was manually annotated, so the annotation quality is high. OCID, which is much larger, uses a semi-automatic process of annotating the labels. It incrementally builds up the scene by adding one object at a time, and computes labels by calculating difference in depth. However, this process is subject to depth sensor noise, so while the majority of the instance label is accurate, the label boundaries are noisy. Additionally, OCID contains images with objects on a tabletop, and images with objects on a floor. Despite our method being trained in synthetic tabletop settings, it generalizes to floor settings as well.

Lastly, we use the Google Open Images Dataset v5 (OID) in an experiment to test the Sim-to-Real gap of the RRN. OID contains approximately 9 million “in-the-wild” images with image-level annotations, bounding boxes, and segmentations. In particular, it contains 2.8 million segmentations for 350 object classes. We filtered the 350 object object classes down to 156 classes that could potentially be on a tabletop, resulting in roughly 220k instance masks on real RGB images.

VIII-C Metrics

We use the precision/recall/F-measure (P/R/F) metrics as defined in . These metrics promote methods that segment the desired objects and penalize methods that provide false positives. Specifically, the F-measure is computed between all pairs of predicted and ground truth masks, which are matched via the Hungarian method. Given a matching, the final P/R/F is computed by

where sis_{i} denotes the set of pixels belonging to predicted object ii, g(si)g(s_{i}) is the set of pixels of the matched ground truth object of sis_{i}, and gjg_{j} is the set of pixels for ground truth object jj. We denote this as Overlap P/R/F. See for more details.

Segmentations with sharp or fuzzy boundaries will obtain similar Overlap P/R/F scores, which will not reflect the efficacy of the RRN. To remedy this, we introduce a Boundary P/R/F measure to complement the Overlap P/R/F. Using the same Hungarian matching used to compute Overlap P/R/F, we compute Boundary P/R/F by

where we overload notation and denote si,gjs_{i},g_{j} to be the set of pixels belonging to the boundaries of predicted object ii and ground truth object jj, respectively. D[⋅]D[\cdot] denotes the dilation operation, which allows for some slack in the prediction. Note that these metrics are sensitive to the amount of allowed slack. However, without this, these numbers can look deceivingly low as no slack requires exact pixel boundary matching which can lead to issues with either noisy manual annotations or noisy semi-automatic annotations . For the dilation operation, we use a circular kernel where the diameter of the kernel depends on the size of the image. Roughly, this metric is a combination of the F\mathcal{F}-measure in along with the Overlap P/R/F as defined in .

We report all P/R/F measures in the range $(P/R/F(P/R/F\times 100$).

VIII-D 2D Quantitative Results

Comparison to baselines. We compare to baselines shown in , which include GCUT , SCUT , LCCP , and V4R . In , these methods were only evaluated on the ARID20 and YCB10 subsets of OCID, so we compare our results these subsets as well. These baselines are designed to provide over-segmentations (i.e., they segment the whole scene instead of just the objects of interest). To allow a more fair comparison, we set all predicted masks smaller than 500 pixels to background, and set the largest mask to table label (which is not considered in our metrics). Results are shown in Table I. Because the baselines aim to over-segment the scene, the precision is in general low while the recall is high. LCCP is designed to segment convex objects (most objects in OCID are convex), but its predicted boundaries are noisy due to utilizing depth only. Recall that OCID labels suffer from depth sensor noise, which explains why LCCP’s boundary recall is quite high (see Section VIII-E for a visual example). Both SCUT and V4R utilize models trained on real data, putting them at an advantage with respect to UOIS-Net-2D. V4R was trained on OSD which has an extremely similar data distribution to OCID, giving V4R a substantial advantage and making it the strongest baseline in . However, our UOIS-Net-2D (line 5 in Table I), despite never having seen any real data, significantly outperforms these baselines on F-measure.

Comparison to SOTA. In Table II, we compare UOIS-Net-2D to two SOTA methods, Mask R-CNN and PointGroup , both trained on RGB-D from TOD. While Mask R-CNN is a general method for detection that can be applied to different types of input modes, the more recent PointGroup requires point clouds and utilizes a sparse convolutional backbone. We see that UOIS-Net-2D (line 6) outperforms Mask R-CNN (line 3) on OSD and slightly on OCID. On the other hand, PointGroup (line 4) provides comparable performance to UOIS-Net-2D. Note, however, that PointGroup reasons in 3D with center voting while UOIS-Net-2D does not. In that sense, Mask R-CNN is more similar to UOIS-Net-2D.

Note that the performance of UOIS-Net-2D is similar to Mask R-CNN and PointGroup on OCID in terms of boundary F-measure. This result is misleading: it turns out that using the RRN to refine the initial masks produced by UOIS-Net results in degraded quantitative performance on OCID, while the qualitative results are better. This is due to OCID’s noisy label boundaries. An illustration of this can be found in the first example (row) of Figure 4. Table III shows the performance of UOIS-Net without applying RRN (DSN and IMP only) on OCID and OSD. This method utilizes only depth and predicts segmentation boundaries that are aligned with the sensor noise. In this setting, UOIS-Net-2D outperforms both SOTA methods, with a significant gain in boundary F-measure and a minor gain in overlap F-measure. In fact, UOIS-Net-3D also gains performance on OCID when removing the RRN which further demonstrates the issue with OCID labels.. On the other hand, OSD has manually annotated labels and this issue is not present, and our methods’ performances deteriorate in the absence of the RRN. In particular, UOIS-Net-2D drops almost 20% relatively without it.

Effect of input mode. To evaluate how different input modes affect results, we train Mask R-CNN on RGB, depth, and RGB-D and compare it to UOIS-Net-2D in Table II. Training Mask R-CNN on synthetic RGB only poorly generalizes from Sim-to-Real. Training on depth drastically boosts this generalization, which is in agreement with . When training on RGB-D, we posit that Mask R-CNN relies heavily on depth as adding RGB to depth results in little change. However, UOIS-Net-2D exploits RGB and depth separately, leading to better results on OSD while being trained on the exact same synthetic dataset. Furthermore, when our DSN is trained directly on RGB-D (line 4, Table II), we see a drop in performance, suggesting that training directly on (non-photorealistic) RGB is not the best way of utilizing synthetic data.

Degradation of training on non-photorealistic simulated RGB. To quantify how much non-photorealistic RGB degrades performance, we train an RRN on real data. This serves as an approximate upper bound on how well the synthetically-trained RRN can perform. We use the instance masks from OID and show results in Table IV. Both models share the same DSN and IMP. The Overlap measures are roughly the same, while the RRN trained on OID has slightly better performance on the Boundary measures. This suggests that while there is still a gap, our method is surprisingly not too far off, considering that we train with non-photorealistic synthetic RGB. We conclude that mask refinement with RGB is an easier task to transfer from Sim-to-Real than directly segmenting from RGB.

Ablation studies. We report ablation studies on OSD to evaluate each component of our proposed method in Table V (left). We omit the Overlap P/R/F results since they follow similar trends to Boundary P/R/F. Running the RRN on the raw masks output by DSN without the IMP module actually hurts performance as the RRN cannot refine such noisy masks. Adding the open/close morphological transform and/or the closest connect component results in much stronger results, showing that the IMP is crucial in robustifying our proposed method. In these settings, applying the RRN significantly boosts Boundary P/R/F. In fact, Table V (right) shows that applying the RRN to the Mask R-CNN results effectively boosts the Boundary P/R/F for all input modes on OSD, showing the efficacy of the RRN. Note that even with this refinement, the Mask R-CNN results are outperformed by our method.

VIII-E 2D Qualitative Results

Comparison to baselines and SOTA. First, we show qualitative results on OCID of baseline methods, Mask R-CNN (trained on RGB-D), PointGroup, and UOIS-Net-2D in Figure 4. It is clear that the baseline methods suffer from over-segmentation issues; they segment the table and background into multiple pieces. This is especially the case for methods that utilize RGB as an input (GCUT and SCUT); the objects are often over-segmented as well. In example (row) 1 of Figure 4, the ground truth label for the keyboard is riddled with holes. Methods that operate only on depth (LCCP and V4R) mirror these noisy boundaries, leading to inflated boundary P/R/F measures. Despite this, our quantitative results still outperform these baselines.

The main failure mode for Mask R-CNN is that it tends to undersegment objects. This is the typical failure mode of top-down instance segmentation algorithms in clutter. A close inspection of Figure 4 shows that Mask R-CNN frequently segments multiple objects as one. Examples 1 and 4 show undersegmentation of many neighboring small objects, and example 2 shows undersegmentation of larger objects. Since PointGroup requires point clouds, it is susceptible to degradation from depth sensor noise as well. Additionally, some masks (e.g. example 4) can be quite noisy, which we hypothesize is due to the rudimentary breadth-first search clustering algorithm.

On the other hand, our method utilizes depth and RGB separately to provide sharp and accurate masks. Because UOIS-Net-2D leverages RGB after depth, it can fix the issues with depth sensors shown in example 1. Additionally, our method can segment complicated structures such as stacks (example 2 and 4) and cluttered environments (example 4).

RRN Refinements. In Figure 5, we qualitatively show the effect of the RRN. The top row shows the masks before refinement (after IMP), and the bottom row shows the refined masks. These images were taken in our lab with an Intel RealSense D435 RGB-D camera to demonstrate the robustness of our method to camera viewpoint variations and distracting backgrounds, as OCID and OSD have relatively simple backgrounds. Due to noise in the depth sensor, it is impossible to get sharp and accurate predictions from depth alone without using RGB. Our RRN can provide sharp masks even when the boundaries of objects are occluding other objects (examples 2 and 5). Our RRN is able to fix squiggly mask boundaries and patch up large missing chunks in the initial mask.

Robustness. In Figure 6 (top), we demonstrate how our method is robust to errors in the pipeline. In row 1, the DSN produces a false positive foreground region in the top right of the image. However, there is not enough evidence in the center directions (not enough discretized angles are present), thus the Hough voting layer suppresses this potential object. In row 2, there are many spurious center direction predictions, however these are not considered as potential object centers by the Hough voting layer because they are not detected by the foreground mask. Additionally, one can see the many holes and spurious mask components in column 5. As seen in the ablation studies in Section VIII-D, applying the RRN here actually degrades the performance. The IMP cleans this image, making the refinement task easier for the RRN.

Failure Modes. We show some failure modes of our DSN in Figure 6 (bottom) on the OSD dataset. In row 1, we see an undetected object (green book) because the center of the mask is occluded. In the center direction image, there is no convergent point for this object which would be considered as the object center with all discretized angles pointing to it. In row 2, there is a false positive region detected by the DSN in the top right corner, which cannot be undone by the RRN. Lastly, in row 3, the DSN cannot correctly segment an object whose mask has been split in two by an occluding object.

Our RRN also has some common failure modes. Column 4 in Figure 5 shows that the RRN can fail when the object has a complex texture. Column 6 demonstrates that if there is not enough padding to the initial mask to see the entire object, the RRN cannot segment the entire object (lego block). In column 2, the RRN has a tough time fixing the segmentation mask when the DSN has incorrectly undersegmented a few objects together (the two cups to the left of the eraser).

VIII-F 3D Quantitative Results

Comparison to baselines. We compare the full UOIS-Net-3D with an RRN trained on TOD to all previous baselines from and our UOIS-Net-2D in Table I. Discussion of all baselines and UOIS-Net-2D is given in Section VIII-D. Compared to UOIS-Net-2D, the 3D version substantially increases the recall in both the overlap and boundary metrics, leading to a relative increase of 4.5% in overlap F-measure and 5.5% in boundary F-measures. In fact, UOIS-Net-3D, despite no convexity assumptions, gets quite close to the overlap recall of LCCP, which segments objects based on convexity. Note that only the DSN structure is changed when moving from 2D to 3D, thus this extra recall is mainly due to the 3D center voting structure and better initial masks.

Comparison to SOTA. In Table II, we show comparisons of UOIS-Net-3D with SOTA methods Mask R-CNN and PointGroup , and the previous UOIS-Net-2D on OCID and OSD. Again, the trend of performance increase from UOIS-Net-2D is similar to before, where a significant increase in recall leads to a boost in F-measure. On OCID, UOIS-Net-3D receives a relative increase of 5.8% in overlap F-measure and a 6.7% in boundary F-measure. Compared to the best performing Mask R-CNN, our performance provides a relative increase of 8.1% on overlap F-measure and 6.0% on boundary F-measure, while being trained on the exact same dataset (TOD). Additionally, we outperform PointGroup by 7.9% and 6.3% on overlap and boundary F-measures, respectively. Note that UOIS-Net-3D performs center voting in a similar fashion to PointGroup, however we believe our novel loss functions target cluttered environments more effectively. On OSD overlap F-measure, we provide a relative increase of 4.3% over UOIS-Net-2D, 12.4% over Mask R-CNN, and 5.7% over PointGroup. For boundary F-measure, we show 8.5% over UOIS-Net-2D, 32.3% over Mask R-CNN, and 8.9% over PointGroup. Note that while our recall jumps modestly in comparison to UOIS-Net-2D, the precision increase is more sizable.

RRN ablation. We refer readers to Table IV to show the full power of our method. When using an RRN trained on OID , our full UOIS-Net-3D further increases the boundary F-measures, leading to 79.9 points on OCID and 77.3 points on OSD, which is 7.8% and 9.2% higher than the full UOIS-Net-2D on OCID and OSD, respectively. In the rest of this section, all UOIS-Net-3D’s will use an RRN trained on OID in order to show the full performance of our method.

ESP module ablation. We evaluate the significance of using the ESP module to obtain a higher receptive field by testing both UOIS-Net-3D and UOIS-Net-2D (which normally does not include the ESP module) with and without the ESP modules embedded into the DSN in Table VII. We immediately see that obtaining a higher receptive field is beneficial to the performance. For UOIS-Net-3D, we see relative gains on OCID of 3.0% and 3.5% on overlap and boundary F-measures, respectively. On OSD, we see similar performance. For UOIS-Net-2D, we see a large relative gain on OCID of 4.9% and 5% on overlap and boundary F-measures, respectively. Note that the DSN with ESP modules has slightly fewer parameters than without them, thus these experiments highlight the usefulness of a higher receptive field.

VIII-G 3D Qualitative Results

2D vs. 3D comparison. In Figure 8, we qualitatively compare the predictions of UOIS-Net-2D and UOIS-Net-3D to understand how reasoning in 3D can solve the 2D issues mentioned beforehand (in Section VIII-E). The first row of Figure 6 (bottom) exhibits the 2D issue of false negative detection when the 2D center of the object is occluded. However, this is easily corrected when voting for 3D centers, as seen in columns 2 and 5 of Figure 8 In columns 1, 3, and 4, we see that UOIS-Net-3D detects and segments more small, thin objects such as pens and bananas. The 2D Hough voting procedure typically fails to detect enough discretized directions for objects with such shape. This issue is ameliorated when reasoning in 3D, as there are no approximations made with discretized directions. Lastly, in columns 3, 4, and 6, we show that the 3D DSN architecture performs better due to a higher receptive field provided by the ESP modules.

Failure modes. In Figure 10, we demonstrate common failure modes of UOIS-Net-3D. Row 1 shows that when objects are close together, they may under-segmented into a single segment. The center votes shows that the three small fruits have center votes that are too close to separate in the post-processing clustering step. Another case of under-segmentation is shown in row 3, where multiple objects (cereal boxes) with flat surfaces are lined up such that the depth map shows a large flat surface. It is difficult to correctly separate these boxes from depth alone; appearance is key in providing the correct segmentation in such situations. Row 4 shows that highly nonconvex objects such as power drills can be over-segmented into pieces. The center votes image clearly shows that the DSN believes this object to be two objects. Lastly, row 2 indicates that the 2D issue of attempting to segment a mask that is split into multiple pieces by an occluding object is still present when reasoning in 3D.

VIII-H Quantifying Generalization from Sim to Real

In order to quantify generalization from simulation to the real-world, we show performance of our methods and SOTA methods Mask R-CNN and PointGroup on the TOD test set. This dataset of 20k images was generated using instances that are not present in the training set. In Table VIII, we show that both Mask R-CNN and PointGroup outperform UOIS-Net on all metrics on this test set. However, as seen in Table II, UOIS-Net outperforms Mask R-CNN and PointGroup on real-world data, indicating that our methods handle the distribution shift better, which is ultimately what we care about.

VIII-I Application in Grasping Unknown Objects

We use our model to demonstrate manipulation of unknown objects in a cluttered environment using a Franka robot with panda gripper and wrist-mounted RGB-D camera. The task is to collect objects from a table and put them in a bin. Object instances are segmented using our method and the point cloud of the closest object to the camera is fed to 6-DOF GraspNet to generate diverse grasps, with other objects considered obstacles. Other objects are represented as obstacles by sampling fixed number of points using farthest point sampling from their corresponding point cloud. The grasp that has the maximum score and has a feasible path is chosen to execute. Figure 11 shows the instance segmentation at different stages of the task and also the execution of the robot grasps. Video of the experiments can be found at the project websitehttps://rse-lab.cs.washington.edu/projects/unseen-object-instance-segmentation/. Our method segments the objects correctly most of the time but fails sometimes, such as the over-segmentation of the drill in the scene. Our method considers the top of the drill as one object and the handle as an obstacle. This is because the object is highly nonconvex as discussed in the previous section. We conducted the experiment to collect 51 objects from 9 different scenes. Each object is considered to be successfully grasped if the robot can pick it up in maximum 2 attempts. Otherwise, we count that object as failure case and remove it from the scene manually and proceed to other objects. In our trials, the robot successfully grasped 41/51 objects (80.3%80.3\% success rate). The failures stem from either imperfections in segmentation or inaccurate generated grasps. Note that neither our method nor 6-DOF GraspNet are trained on real data.

IX Conclusion

We proposed a deep network, UOIS-Net, that separately leverages RGB and depth to provide sharp and accurate masks for unseen object instance segmentation. Our two-stage framework produces rough initial masks using only depth by regressing center votes in either 2D or 3D, then refines those masks with RGB. Surprisingly, our RRN can be trained on non-photorealistic RGB and generalize quite well to real world images. We demonstrated the efficacy of our approach on multiple tabletop environment datasets and showed that our model can provide strong results for unseen object instance segmentation. Finally, we also addressed the weaknesses of our method which will serve as the base of future work.

References