Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance

Phuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan, Anh Tran, Cuong Pham, Khoi Nguyen

Introduction

This paper addresses the challenging problem of open-vocabulary 3D point cloud instance segmentation (OV-3DIS). Given a 3D scene represented by a point cloud, we seek to obtain a set of binary instance masks of any classes of interest, which may not exist during the training phase. This problem arises to overcome the inherent constraints of the conventional fully supervised 3D instance segmentation (3DIS) approaches , which are bound by a closed-set framework – restricting recognition to a predefined set of object classes that are determined by the training datasets. This task has a wide range of applications in robotics and VR systems. This capability can empower robots or agents to identify and localize objects of any kind in a 3D environment using textual descriptions that detail names, appearances, functionalities, and more.

There are a few studies addressing the OV-3DIS so far . Most recently, proposes the use of a pre-trained 3DIS model instance proposals network to capture the geometrical structure of 3D point cloud scenes and generate high-quality instance masks. However, this approach faces challenges in recognizing rare objects due to their incomplete appearance in the 3D point cloud scene and the limited detection capabilities of pre-trained 3D models for such infrequent classes. Another approach involves leveraging 2D off-the-shelf open-vocabulary understanding models to easily capture novel classes. Nevertheless, translating these 2D proposals from images to 3D point cloud scenes is challenging. This is because of the fact that 2D proposals capture only the visible portions of 3D objects and may also include irrelevant regions, such as the background. These two approaches are summarized in Fig. 1.

In this work, we introduce Open3DIS, a method for OV-3DIS that extends the understanding capability beyond predefined concept sets. Given an RGB-D sequence of images and the corresponding 3D reconstructed point cloud scene, Open3DIS addresses the limitations of existing approaches. It complements two sources of 3D instance proposals by employing a 3D instance network and a 2D-guide-3D Instance Proposal Module to achieve sufficient 3D object binary instance masks. The module (our key contribution) extracts geometrically coherent regions from the point cloud under the guidance of 2D predicted masks across multiple frames and aggregates them into higher-quality 3D proposals. Later, Pointwise Feature Extraction aggregates CLIP features for each instance in a multi-scale manner across multiple views, constructing instance-aware point cloud features for open-vocabulary instance segmentation.

To assess the open-vocabulary capability of Open3DIS, we conduct experiments on the ScanNet200 , S3DIS , and Replica datasets. Open3DIS achieves state-of-the-art results in OV-3DIS, surpassing prior works by a significant margin. Especially, Open3DIS delivers a noteworthy performance improvement of ∼1.5{\sim}1.5 times compared to the leading method on the large-scale dataset ScanNet200.

In summary, the contributions of our work are as follows:

We present the “2D-guided 3D Proposal Module” creating precise 3D proposals by clustering cohesive point cloud regions using aggregated 2D instance masks from multi-view RGB-D images.

We introduce a novel pointwise feature extraction method for open-vocabulary 3D object proposals.

Open3DIS achieves state-of-the-art results on ScanNet200, S3DIS, and Replica datasets, exhibiting comparable performance to fully supervised methods.

Related Work

Open-vocabulary 2D scene understanding methods aim to recognize both base and novel classes in testing where the base classes are seen during training while the novel classes are not. Based on the types of recognition tasks, we can categorize them into open-vocabulary object detection (OVOD) , open-vocabulary semantic segmentation (OVSS) , and open-vocabulary instance segmentation (OVIS) . A typical approach for handling the novel classes is to leverage a pre-trained visual-text embedding model, such as CLIP or ALIGN as a joint text-image embedding where base and novel classes co-exist, in order to transfer the models’ capabilities on base classes to novel classes. However, these methods cannot trivially extend to 3D point clouds because 3D point clouds are unordered and imbalanced in density, and the variance in appearance and shape is much larger than that of 2D images.

Fully-supervised 3D Instance Segmentation (F-3DIS) aims to segment 3D point cloud into instances of training classes. Methods of F-3DIS can be categorized into three main groups: box-based , cluster-based , and dynamic convolution-based techniques. Box-based methods detect and segment the foreground region inside each 3D proposal box to get instance masks. Cluster-based methods employ the predicted object centroid to group points to clusters or construct a tree or graph structure and subsequently dissect these into subtrees or subgraphs . For the third group, Mask3D and ISBNet , proposed using dynamic convolution whose kernels, representative of different object instances, are convoluted with pointwise features to derive instance masks. In this paper, we use ISBNet as a 3D network, yet with necessary adaptations to output 3D class-agnostic proposals.

Open-vocabulary 3D semantic segmentation (OV-3DSS) and object detection (OV-3DOD) enable the semantic understanding of 3D scenes in an open-vocabulary manner, including affordances, materials, activities, and properties within unseen environments. This capability is highlighted in recent work for OV-3DSS and for OV-3DOD. Nevertheless, these methods cannot precisely locate and distinguish 3D objects with 3D instance masks, and thus cannot fully describe 3D object shapes.

Open-vocabulary 3D instance segmentation (OV-3DIS) concerns segmenting both seen and unseen classes (during training) of a 3D point cloud into instances. Methods of OV-3DIS can be split into 3 groups: open-vocabulary semantic segmentation-based, text description and 3D proposal contrastive learning based, and 2D open-vocabulary powered approaches. The first group includes OpenScene and Clip3D utilize clustering techniques such as DBScan on OV-3DSS results to generate 3D instance proposals. However, their quality relies on clustering accuracy and can lead to unreliable results for unseen classes. On the other hand, the second group comprising PLA , RegionPLC , and Lowis3D focuses on training the 3D instance proposal network along with a contrastive open-vocabulary between the predicted proposals and their corresponding text captions. However, when growing the number of classes, these methods struggle to handle and may degrade their ability to distinguish diverse object classes. For the final group, OpenMask3D uses a pre-trained 3DIS model to produce class-agnostic 3D proposals and classifies them based on their 2D mask projection CLIP score. However, the pre-trained 3DIS model faces challenges in identifying small or rare object categories with uncommon geometric structures. Meanwhile, OVIR-3D and SAM3D leverage pretrained 2D open-vocabulary models to produce 2D instance masks and back-project them onto the associated 3D point cloud. However, the 2D segmentation masks are not well-aligned with objects giving rise to low-quality 3D proposals where background points can be included in foreground objects. Nonetheless, the advantage of this group over other groups is in their leverage of 2D pretrained model on large-scale datasets such as CLIP or SAM which can be scaled to hundreds of classes as in Scannet200 . Following the final group, Open3DIS generates high-quality 3D instance proposals by combining 3D masks from a 3DIS network with proposals produced by grouping geometrically coherent regions (superpoints) with the guidance of 2D instance masks. This complements the class-agnostic 3D instance proposals from 3D networks. Our method excels at capturing rare objects while preserving their 3D geometrical structures, achieving state-of-the-art performance in the OV-3DIS domain.

Method

Our approach processes a 3D point cloud and an RGB-D sequence, producing a set of 3D binary masks indicating object instances in the scene. We assume known camera parameters for each frame. Our architecture is depicted in Fig. 2. Similarly to prior work , we employ a 3DIS network module to extract object proposals directly from the 3D point cloud. This module leverages 3D convolution and attention mechanisms, capturing spatial and structural relations for robust 3D object instance detection. Despite its advantages, sparse point clouds, sampling artifacts, and noise can lead to missed objects, especially for small objects e.g., the tissue box in Fig. 1.

Our approach integrates a novel 2D-guide-3D instance proposal module, leveraging 2D instance segmentation networks trained on large image datasets to better capture smaller objects in individual images. However, resulting 2D masks may only capture parts of actual 3D object instances due to occlusions (Fig. 2 - \Circled[]2). To address this, we propose a strategy that constructs 3D object instance proposals by hierarchically aggregating and merging point cloud regions from back-projected 2D masks of the same object. To enhance the robustness and geometric homogeneity, we use “superpoints” during the merging process. This yields complete object instances, complementing those extracted by 3DIS networks.

Detailed analysis in Tab. 1 exhibits the significant enhancement in recall rate, especially for rare classes, when integrating 2D and 3D proposals.

To enable open-vocabulary classification, we additionally employ a point-wise feature extraction module to construct a dense feature map across the 3D point cloud. In the following sections, we explain our modules in more detail, starting with the 2D-guide-3D Instance Proposal Module which constitutes our main contribution.

Superpoints. In a pre-processing step, we utilize the method of to group points into geometrically homogeneous regions, termed superpoints (Fig. 2 - \Circled[]1). This yields a set of UU superpoints {qu}u=1U∈{0,1}U×N\{{\bf q}_{u}\}_{u=1}^{U}\in\{0,1\}^{U\times N}, where qu{\bf q}_{u} is a binary mask of points. Superpoints enhance processing efficiency in the later stages of our pipeline and contribute to well-formed candidate object instances.

Per-frame superpoint merging. For all input frames, we utilize a pretrained 2D instance segmenter, employing Grounding-DINO and SAM . The network outputs a set of 2D masks (Fig. 2 - \Circled[]2). For each 2D mask with index mm (unique across all frames), we calculate the IoU ou,mo_{u,m} with each superpoint qu{\bf q}_{u} when projecting all points of qu{\bf q}_{u} onto the image plane of mask mm using the known camera matrix, excluding points outside the camera’s field of view, and determining image pixels containing projected points. A superpoint is considered to have sufficient overlap with a 2D mask if the IoU is higher than a threshold ou,m>τiouo_{u,m}>\tau_{iou}.

3D object proposal formation. To create 3D object proposals, one option is to utilize the point cloud regions obtained from the merging procedure across individual frames. However, this results in fragmented proposals, capturing only parts of object instances, as the regions correspond to 2D masks from single views (Fig. 2 - \Circled[]2). To address this, we merge point cloud regions from different frames in a bottom-up manner, creating more complete and coherent 3D object masks. Agglomerative clustering combines region sets from pairs of frames until no compatible pairs remain. The resulting set includes merged and standalone regions, which can be matched with other region sets from subsequent frames. In the following paragraphs, we discuss three crucial design choices in this process: (a) the matching score between region pairs, (b) the matching process between sets of regions, and (c) the order of frames or region sets used in matching and merging.

Matching score. For a pair of point cloud regions (ri,rj)({\bf r}_{i},{\bf r}_{j}), we define a matching score based on (a) feature similarity and (b) overlap degree. Their feature-based similarity s′s^{\prime} is measured through cosine similarity between the regions’ feature vectors fi3D{\bf f}_{i}^{\text{3D}}, or si,j′=cos(fi3D,fj3D)s^{\prime}_{i,j}=\text{cos}({\bf f}_{i}^{\text{3D}},{\bf f}_{j}^{\text{3D}}), which are in turn computed as the average of their point features. While this measures if the regions belong to the same object’s shape, it may yield high similarity for duplicate instances with the same geometry. To address this, we also consider the degree of overlap, expressed as the IoU oi,j′=IoU(ri,rj)o^{\prime}_{i,j}=\text{IoU}({\bf r}_{i},{\bf r}_{j}) between the two regions ri,rj{\bf r}_{i},{\bf r}_{j}, which is expected to be high for overlapping regions of the same instance. Two regions are considered matching if their feature-based similarity and IoU score satisfy si,j′>τsims^{\prime}_{i,j}>\tau_{sim} and oi,j′>τiouo^{\prime}_{i,j}>\tau_{iou} (same thresholds used during per-frame superpoint merging). Our approach, incorporating matching scores based on point cloud deep features and geometric structures, results in more coherent and well-defined point cloud regions compared to other strategies (see Tab. 7).

Agglomerative clustering process. To merge region sets {ri}i=1I\{{\bf r}_{i}\}_{i=1}^{I} and {rj}j=1J\{{\bf r}_{j}\}_{j=1}^{J} from different frames into a unified set {rl}l=1L\{{\bf r}_{l}\}_{l=1}^{L}, where L≤I+JL\leq I+J, we employ Agglomerative clustering . We begin by concatenating them into a single “active set” {rl}l=1I+J\{{\bf r}_{l}\}_{l=1}^{I+J}. We compute the each entry ci,jc_{i,j} of the binary cost matrix C{\bf C} of size (I+J)×(I+J)(I+J)\times(I+J) as:

The agglomerative clustering procedure iteratively merges regions within the “active set” according to the cost matrix C{\bf C} and continues to update this matrix until no further merges are possible - indicated by the absence of any positive elements in C{\bf C}.

Merging order. We explored two merging strategies: a sequential order, where region sets are merged between consecutive frames, and the resulting set is further merged with the next frame, and a hierarchical order, which involves merging region sets between non-consecutive frames in separate passes. The hierarchical approach forms a binary tree, with each level merging sets from consecutive pairs of the previous level (see Fig. 3). Details and performance analysis are presented in the Experiments section.

2 3D Instance Segmentation Network

Network design. This network directly processes 3D point clouds to generate 3D object instance masks. We employ established 3D instance segmentation networks like Mask3D and ISBNet as our backbone. For each object candidate, the kernel computed from sampled points and their neighbors is convolved with point-wise features to predict the binary mask. In our open-vocabulary scenario, we exclude semantic labeling heads, focusing solely on the binary instance mask head. The output consists of K2K_{2} binary masks in a K2×NK_{2}{\times}N binary matrix M2{\bf M}_{2} (see Fig. 2 - \Circled[]4).

Combining object instance proposals. We simply append the proposals of set M2{\bf M}_{2} to M1{\bf M}_{1} to form the final set of KK proposals M{\bf M} with the size of K×NK{\times}N. Note that we apply NMS here to remove near-duplicate proposals with the overlapping IoU threshold τdup\tau_{dup}.

3 Pointwise Feature Extraction

In the final stage of our pipeline, we compute a feature vector for each 3D object proposal from our combined proposal set. This per-proposal feature vector serves various instance-based tasks, such as comparison with text prompts in the CLIP space . Unlike prior open-vocabulary instance segmentation methods , which use a top-λ\lambda frame/view approach, we employ a more “3D-aware” pooling strategy. This strategy accumulates feature vectors on the point cloud, considering the frequency of each point’s visibility in each view (see Fig. 4). Our rationale is that points more frequently visible in the top-λ\lambda views should contribute more to the proposal’s feature vector.

where ∗* is the element-wise multiplication (broadcasting if necessary) and ∥x∥2\|x\|_{2} is the L2 normalized vector of xx.

The final classification score between a text query ρ\rho and a 3D mask mk3D{\bf m}^{\text{3D}}_{k} is the average cosine similarity score between its CLIP text embedding eρ{\bf e}_{\rho} and all points within the mask, particularly:

where ∣mk3D∣|{\bf m}^{\text{3D}}_{k}| is the number of points in the kk-th mask.

Experiments

Datasets. We mainly conduct our experiments on the challenging dataset ScanNet200 , comprising 1,201 training and 312 validation scenes with 198 object categories. This dataset is well-suited for evaluating real-world open-vocabulary scenarios with a long-tail distribution. Additionally, we conduct experiments on Replica (48 classes) and S3DIS (13 classes) for comparison with prior methods . Replica has 8 evaluation scenes, while S3DIS includes 271 scenes across 6 areas, with Area 5 used for evaluation. We follow the categorization approach from for S3DIS. Notably, we omit experiments on ScanNetV2 due to its relative ease compared to ScanNet200 and identical input point clouds.

Evaluation metrics. We evaluate using standard AP metrics at IoU thresholds of 50% and 25%. Additionally, we calculate mAP across IoU thresholds from 50% to 95% in 5% increments. For ScanNet200, we report category group-specific APhead{}_{\text{head}}, APcom{}_{\text{com}}, and APtail{}_{\text{tail}}.

Implementation Details. To process ScanNet200 and S3DIS scans efficiently, we downsampled the RGB-D frames by a factor of 10. Our approach utilizes the Grounded-SAM frameworkhttps://github.com/IDEA-Research/Grounded-Segment-Anything, integrating Grounding-DINO and Segment Anything . We employ the dataset class names as text prompts for generating 2D instance masks, followed by NMS with τdup=0.5\tau_{dup}=0.5 to handle overlapping instances. Our PyTorch implementation includes a 2D-guide-3D Instance Proposal Module, generating superpoints from . In Pointwise Feature Extraction, each proposal is projected into all viewpoints, and we select the top λ=5\lambda{=}5 views with the largest projected points. For CLIP, we use the ViT-L/14 variant trained on OpenAI’s WIT dataset .

2 Comparison to prior work

Setting 1: ScanNet200. The quantitative evaluation of the ScanNet200 dataset is summarized in Tab. 2. Following , we utilize the class-agnostic 3D proposal network trained on the ScanNet200 training set, then test the OV-3DIS on the validation set. Employing our 2D-Guided-3D Instance Proposal Module, Open3DIS achieves 18.218.2 and 19.219.2 in AP and APtail{}_{\text{tail}}. We outperform OVIR-3D and OpenMask3D by margins of +5.2+5.2 and +2.8+2.8 in AP, and surpass all other methods, even the fully-supervised approaches in the APtail{}_{\text{tail}} metric. This emphasizes the effectiveness of our 2D-Guide-3D Instance Proposal Module, which is effective in crafting precise 3D instance masks independently of any 3D models. Combining with class-agnostic 3D proposals from ISBNet boosts our performance to 23.723.7, 29.429.4, and 32.832.8 in AP, AP50, and AP25 — reflecting a 1.5x1.5{\bf x} enhancement in AP compared to prior methods. Impressively, our method competes closely with fully supervised techniques, attaining approximately 96%96\% and 88%88\% of the AP scores of ISBNet and Mask3D, and excelling in the APcom{}_{\text{com}} and APtail{}_{\text{tail}}. This performance underscores the advantages of merging 2D and 3D proposals and demonstrates our model’s adeptness at segmenting rare objects.

To assess the generalizability of our approach, we conducted an additional experiment where the class-agnostic 3D proposal network is substituted with the one trained solely on the ScanNet20 dataset. We then categorized the ScanNet200 instance classes into two groups: the base group, consisting of 51 classes with semantics similar to ScanNet20 categories, and the novel group of the remaining classes. We report the APnovel{}_{\text{novel}}, APbase{}_{\text{base}}, and AP in Tab. 3. Our proposed Open3DIS achieves superior performance compared to PLA , OpenMask3D , with large margins in both novel and base classes. Notably, PLA , trained with contrastive learning techniques, falls in a setting with hundreds of novel categories.

Setting 2: Replica. We further evaluate the zero-shot capability of our method on the Replica dataset, with results detailed in Tab. 4. Considering that several Replica categories share semantic similarities with ScanNet200 classes, to maintain a truly zero-shot scenario, we omitted the class-agnostic 3D proposal network for this dataset (using proposals from 2D only). Under this constraint, our approach still outperforms OpenMask3D and OVIR-3D by margins of +5.0+5.0 and +7.0+7.0 in AP, respectively.

Setting 3: S3DIS. In line with the setting of PLA , we trained a fully-supervised 3DIS model on the base classes of the S3DIS dataset, followed by testing the model on both base and novel classes. The results are shown in Tab. 5, where we report the performance in terms of AP50B{}^{B}_{50} and AP50N{}^{N}_{50}, representing the AP50 for the base and novel categories, respectively. Open3DIS significantly outperforms existing methods in AP50N{}^{N}_{50}, achieving more than double their scores. This remarkable performance underscores the efficacy of our approach in dealing with unseen categories, with the support of the 2D foundation model.

Our qualitative results with arbitrary text queries. We visualize the qualitative results of text-driven 3D instance segmentation in Fig. 5. Our model successfully segments instances based on different kinds of input text prompts, involving object categories that are not present in the labels, object’s functionality, object’s branch, and other properties.

3 Ablation study

To validate the design choices of our proposed method, we have carried out a series of ablation studies on the validation set of the ScanNet200 dataset.

Study on different kinds of features for open-vocabulary classification is presented in Tab. 6. In the first three rows (setting A1-A3), we employ the pointwise feature map extracted by OpenScene to perform classification on our 3D proposals. Of these, the fusion approach, which directly projects CLIP features from 2D images onto the 3D point cloud, yields the highest results, 17.517.5 in AP. In setting B, we adopt a strategy akin to , extracting features for each mask by projecting the 3D proposals onto the top-λ\lambda views, which attains an AP of 22.222.2. Surpassing these, our Pointwise Feature Extraction (setting C) achieves the best AP score of 23.723.7, substantiating our design choice.

Study on the 2D-guide-3D Instance Proposal Module is in Tab. 7. Our proposed approach (row 1), utilizing superpoints to merge 3D points into regions and filter outliers based on cosine similarity in feature space, achieves an AP of 18.2. Disabling this filtering notably reduces AP by 2.3. Comparatively, a more basic method (row 3) relying on Euclidean distance to eliminate outlier superpoints yields an AP of 16.0, showing the lesser effectiveness of Euclidean distance for noise filtering. Our baseline (last row), grouping 3D points solely based on 2D masks, significantly decreases AP to 12.0, underscoring the necessity of superpoint merging for effective 3D proposal creation.

We study different merging configurations, including merging strategy and merging order in Tab. 8. We compare the proposed Agglomerative clustering and the Hungarian matching. Specifically, we first establish a partial matching between two sets of regions, then matched pairs are merged into new refined regions, and unmatched ones remain the same. Using Hungarian matching yields inferior results relative to agglomerative clustering, with a drop of ∼2.0{\sim}2.0 in AP. Adopting the sequential merging order leads to a slight decrease by ∼1.0{\sim}1.0 in AP in performance. The best results are achieved when agglomerative clustering is paired with the hierarchical merging order.

Ablation Study on Segmenters. Our comparative analysis of various class-agnostic 3D segmenters and open-vocabulary 2D segmenters is presented in Tab. 9 and 10. The findings reveal that utilizing either ISBNet or Mask3D leads to similar levels of performance, achieving an AP of 23.723.7. Incorporating 2D instance masks from SEEM , Detic or ODISE leads to a slight decrease in AP by ∼1.4{\sim}1.4, which we attribute to the less refined outputs produced by these models.

Ablation study on different values of visibility threshold and similarity threshold. We report the performance of our version using only proposals from the 2D-G-3DIP with different values of the visibility threshold and similarity threshold in Tab. 11 and 12. Using τiou=0.9\tau_{iou}{=}0.9 and τsim=0.9\tau_{sim}{=}0.9 yields optimal results.

Study on different values of viewpoints is illustrated in Tab. 13. Relying only on the viewpoint with the highest number of projected points reduces the AP score to 21.221.2. Conversely, raising the number of views to 10 or more also yields worse results, likely due to the presence of inferior, occluded 2D masks. λ=5\lambda{=}5 reports the best performance.

Discussion

Limitations. Our Class-agnostic 3D Proposal and 2D-guided-3D Instance Proposal Module currently operate independently, with their outputs being combined to obtain the final 3D proposal set. A better-integrating strategy, where these modules enhance each other’s performance in a synergistic fashion, would be an interesting future direction.

Conclusions. This paper has introduced Open3DIS, the pioneering method that combines 3D class-agnostic proposals with 2D open-vocabulary segmentation to tackle open-vocabulary 3D instance segmentation. Our 2D-Guided-3D Instance Proposal Module generates top-tier 3D proposals by integrating 2D instance masks from RGB-D images. We significantly outperform existing OV-3DIS approaches across benchmarks (around 50% on ScanNet200, 80% on S3DIS, and 40% on Replica) and exhibit comparable performance to fully supervised approaches like ISBNet or Mask3D. Additionally, our method excels in segmenting objects in 3D spaces based on diverse textual descriptions, unlocking new capabilities for machine comprehension and interaction within intricate 3D environments.

References