SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation
Mutian Xu, Xingyilang Yin, Lingteng Qiu, Yang Liu, Xin Tong, Xiaoguang Han
Introduction
3D indoor scene segmentation is a fundamental task in 3D visual understanding and plays a vital role in diverse applications such as augmented reality, room navigation, and robotics. The objective is to predict 3D object masks from input 3D scenes which are often represented by meshes, point clouds, and posed RGB images.
Traditional methods for this task lack the zero-shot capability. They often struggle to accurately segment newly introduced object categories that were not encountered during training . Despite recent efforts that harness vision foundation models to enhance zero-shot 3D scene understanding, they necessitate either 3D pretrained networks or training on domain-specific data. As a result, directly applying them to novel 3D scenes remains challenging in terms of generalization.
In the field of 2D image segmentation, the Segment Anything Model (SAM) recently brought the breakthrough. Trained on an extensive SA-1B dataset , SAM can “segment any unfamiliar images” without the need for further training. Having witnessed the great power of SAM, and recognizing that a 3D scene is essentially a combination of multiple 2D image frames, a curious question arises: Given 3D scene point clouds associated with posed 2D frames, is it possible to apply SAM directly to 2D frames for segmenting any 3D scenes without additional training?
We attempt to investigate this question by revisiting the fundamental design principle of SAM, with a primary focus on the concept of prompt. The distinguishing characteristic of SAM lies in its promptability as a segmentation model, allowing it to accept various types of input prompts such as pixel coordinates, bounding boxes, masks, or free-form text. These prompts serve to specify where or what is to be segmented within an image. Leveraging this capability, the final stage of SAM conducts a fully automatic segmentation process, referred to as automatic-SAM in our paper. In the output masks of automatic-SAM, each 2D object is segmented by a single prompt (see Fig. 2 (d)). Extending this underlying design rationale of SAM to 3D scene segmentation, our goal is to ensure the consistency of corresponding 2D segmentations for the same 3D object across different frames by aligning them with a single prompt.
Yang et al. recently embarked on a similar inquiry to ours. They developed a project called SAM3D to utilize automatic-SAM on individual 2D frames, generating pixel prompts and corresponding segmentations in a fully automatic manner. However, this method assigns frame-specific pixel prompts that lack alignment across frames, causing substantial inconsistencies in the 2D masks for the same 3D object across frames (see Fig. 2 (a)). Consequently, it yields inferior 3D segmentation results, rendering it unsuitable for general-purpose applications.
To achieve prompt alignment across frames, one possible solution is to employ automatic-SAM on an initial frame to generate 2D-pixel prompts, which can then be propagated to subsequent frames, analogous to SAM-PT for video tracking. Nevertheless, videos processed by SAM-PT exhibit foreground objects that consistently appear across frames, whereas in 3D scenes, objects are often scattered in different frames. Consequently, the prompts initially generated on 2D frames of 3D scenes cannot be propagated to cover newly emerged objects in other frames, leading to incomplete segmentation of the entire scene (see Fig. 2 (c)).
In this paper, to tackle the aforementioned challenges, we propose to locate 3D points in input scenes as SAM prompts. These 3D prompts are then projected onto 2D frames to get pixel prompts. In this way, a 3D point serves as the natural prompt to align the pixel prompts projected from this 3D point across different frames, making both pixel prompts and their SAM-predicted masks for the same 3D object exhibit consistency across frames (see Fig. 2 (b)). Moreover, as the located 3D points sufficiently cover the whole 3D scene, their corresponding segmentation masks also encompass the entirety of the 3D scene.
Building upon this key idea, we present a straightforward yet effective framework called SAMPro3D to maximize the potential of SAM for zero-shot 3D indoor scene segmentation. Given a 3D scene point cloud and the associated posed 2D frames, our SAMPro3D begins by initializing 3D prompts and using SAM to generate their corresponding 2D masks across frames.
Next, since some 3D prompts may generate low-quality and redundant masks that degrade final results, we introduce an algorithm to filter 3D prompts according to their corresponding masks’ quality on all 2D frames. This mechanism “collects the feedback” from all 2D views, and prioritizes the retention of prompts that “make all frames happy”, while still maintaining prompt consistency.
Additionally, we find that sometimes the 2D masks aligned by a single 3D prompt may only segment part of the object due to the limited coverage of 2D frames (see Fig. 4). Thus we designed a simple strategy to consolidate 3D prompts that exhibit certain overlaps in their generated masks into one single prompt, as they are likely segmenting the same 3D object. This information integration on prompts brings a more comprehensive object segmentation.
Finally, we project all input points of the scene onto each segmented frame and accumulate their predictions across frames to derive the final 3D segmentation. Notably, our approach does not require additional training or 3D pretrained networks on domain-specific data, enabling us to retain the zero-shot ability of SAM. We conduct comprehensive experiments to validate the efficacy of our approach and gain insights into our design choices. Rich qualitative and quantitative results show that our method consistently achieves higher quality and more diverse segmentations than previous zero-shot or fully supervised approaches, and in many cases even surpasses human-level annotations.
Furthermore, we incorporate HQ-SAM and Mobile-SAM into our pipeline, showing that enhancements on 2D images seamlessly translate into improved 3D results.
Related Work
The field of 3D scene understanding has been dominated by closed set methods which primarily focus on training deep neural networks on domain-specific datasets . The first line of research focuses on improving representation learning from human-annotated 3D labels , for solving different 3D scene understanding tasks . Another stream of works aims to construct semi-/weak-/self- supervision signals from 3D data , so as to minimize the need for 3D annotations. In addition, some methods leverage 2D supervision to assist the model training for 3D scene understanding .
However, the aforementioned methods all rely on training with domain-specific 3D or 2D data, limiting their zero-shot ability to understand new scenes that have never been seen during training. Instead, our framework seeks to straightforwardly harness the inherent zero-shot ability of SAM for segmenting 3D scenes, thereby eliminating the need for additional model training.
The studies on zero-shot 3D scene understanding are very limited , and they still involve training with supervised 3D labels for a predefined set. Recently, a bunch of 2D vision foundation models have shown their remarkable zero-shot recognition abilities . This encourages the researchers from the 3D vision community to leverage them for 3D scene parsing . For example, all use CLIP to extract pixel-wise features and align them with 3D space to realize language-guided segmentation of 3D scenes.
Despite recent advancements in open-set methods, applying them directly to new 3D data remains challenging, since they still necessitate model tuning , 3D-2D distillation , or 3D pretrained networks on domain-specific 3D data to bridge the gap between 2D and 3D modalities. In contrast, our method just uses SAM on RGB frames to get segmentation masks without requiring any of the aforementioned factors. This preserves the zero-shot ability of SAM and enables direct deployment for segmenting any 3D scenes.
The Segment Anything Model (SAM) has brought about a revolution in the field of 2D image segmentation. Trained on an astonishing SA-1B dataset, SAM has acquired extensive knowledge that empowers it to effectively segment unfamiliar images without further training. Appreciating the superior power of SAM, several recent works are striving to lift SAM into 3D visual tasks. Cen et al. use SAM to segment target objects in NeRF via one-shot manual prompting. Zhang et al. define hand-crafted grid prompts on Bird’s Eye View images and perform SAM for 3D object detection in the context of autonomous driving.
For 3D indoor scene segmentation, Yang et al. recently employed automatic-SAM on individual 2D image frames to generate pixel prompts and segmentation masks automatically. However, this method assigns pixel prompts specific to each frame that lack alignment across frames, causing segmentation inconsistencies across frames and subpar 3D segmentation results.
Different from them, our key idea is to locate SAM prompts in 3D space so that the pixel prompts derived from different frames but projected by the same 3D prompt become harmonized in the 3D space, leading to the frame consistency of prompts and their masks, and bringing high-quality 3D segmentation.
Method
Given 3D point clouds of indoor scenes and associated 2D posed frames, we aim to directly apply the Segment Anything Model (SAM) to 2D frames for zero-shot 3D indoor scene segmentation. An overview of our framework is presented in Fig. 3, highlighting how the prompt is processed. Our core idea lies in 3D Prompt Proposal (Sec. 3.1), where 3D prompts are located in input scenes, which are then projected onto 2D frames to generate pixel prompts for feeding SAM. Next, we introduce a 2D-Guided Prompt Filter Module (Sec. 3.2), which filters the 3D prompts based on their segmentation performance across all 2D frames. We further propose a Prompt Consolidation Strategy (Sec. 3.3) to consolidate different prompts if they are segmenting the same object. Finally, we project all input points onto each segmented 2D frame and aggregate their mask predictions across frames to obtain the final 3D segmentation (Sec. 3.4).
1 3D Prompt Proposal
SAM is a promptable segmentation model that can accept various inputs such as pixel coordinates, bounding boxes or masks and predict the segmentation area associated with each prompt. In our framework, we feed all pixel coordinates calculated before to prompt SAM and obtain the 2D segmentation masks on all frames. As depicted in Fig. 2 and described in Sec. 1, through locating prompts in 3D space, the pixel prompts originating from distinct frames but projected by the identical 3D prompt will be aligned in 3D space, bringing frame-consistency on pixel prompts along with their SAM-predicted masks. Notably, this step is the only step that we need to perform SAM. In the later stages of our pipeline, our attention is directed towards filtering and consolidating the initial 3D prompts along with their initial segmentation masks, while preserving their original segmentation areas and patterns intact.
2 2D-Guided Prompt Filter
During the previous prompt initialization, some prompts may generate low-quality and redundant masks that will degrade the final results To handle this issue, we introduce a mechanism to “collect the feedback from all frames”.
As outlined in Algorithm 1, we first adopt the strategy proposed in automatic-SAM to filter the prompts on each individual frame. Basically, this strategy eliminates the prompts whose corresponding masks have low confidence scores or have large overlaps with others. If a 3D prompt has a valid pixel projection within a frame, its counter increments. If the prompt successfully survives the filtering stage in that frame, its score accumulates. After evaluating all frames, we compute the probability of retaining a 3D prompt and keep the prompt when its probability exceeds a predefined threshold .
This algorithm enables us to “make all frames are happy” by considering the feedback from all 2D views. It prioritizes high-quality prompts while maintaining prompt consistency across frames, ultimately enhancing the 3D segmentation result. We also tried other filtering strategies and ablated in Sec. 4.2.2. More details are provided in the supplementary material.
3 Prompt Consolidation
Sometimes the 2D masks aligned by a single 3D prompt may only segment part of the object due to the limited coverage of 2D frames. For example in Fig. 4 (a), the floor is segmented by the masks belonging to several 3D prompts.
To address this, we designed a Prompt Consolidation strategy (see Fig. 4 (b)). The strategy involves examining the generated masks of different 3D prompts and identifying a certain overlap between them. In such cases, we consider these prompts as likely segmenting the same object and consolidate them into a single pseudo prompt. This process facilitates the integration of information across prompts, leading to a more comprehensive segmentation of the object.
4 3D Scene Segmentation
After previous procedures, we have obtained the final set of 3D prompts and their 2D segmentation masks across frames. In addition, we have also ensured that each 3D object is segmented by a single prompt, allowing the prompt ID to naturally serve as the object ID.
With the ultimate goal of segmenting all points within the 3D scene, we continue by projecting all input points of the scene onto each segmented frame and compute their predictions using the following steps: For each individual input point in the scene, if it is projected within the mask area segmented by a prompt in frame , we assign its prediction within that frame as the prompt ID . We accumulate the predictions of across all frames and assign its final prediction ID based on the prompt ID that has been assigned to it the most number of times. By repeating this process for all input points, we can achieve a complete 3D segmentation of the input scene.
Experiments
We utilize the ScanNet200 dataset which provides extensive annotations for 200 categories based on RGB-D data from ScanNet . This dataset presents a conspicuously challenging scenario for zero-shot 3D indoor scene segmentation. We evaluate our framework on its validation set containing 312 scenes. To expedite processing, we resize each RGB image frame to a resolution of 240320, which has proven to be sufficient for our method to generate high-quality 3D segmentation results. We employ the ViT-H SAM model, which is the default public model for SAM. The entire framework is executed on a single NVIDIA A100 GPU.
1 Main Results
As highlighted in the paper, our method directly applies SAM to RGB-frames, without further training, distillation or 3D pretrained networks on domain-specific data. We first compare our method with SAM3D proposed by Yang et al. which also eliminates the need for the aforementioned factors. Yang et al. also present an ensembled version of SAM3D where it leverages a graph-based image segmentation algorithm to ensemble and refine their original result. Here, in order to fairly compare the performance of the original algorithm itself, we present its original results. We also compare our method with Mask3D , a state-of-the-art fully supervised method trained on the ScanNet200 training dataset with comprehensive 3D annotations. It’s worth mentioning that OpenMask3D , which focuses on open-vocabulary 3D instance segmentation, relies on Mask3D for generating 3D instance masks in its initial stage. Therefore, the quality and diversity of its 3D segmentation results heavily depend on Mask3D. In our evaluation, we directly present the results obtained by Mask3D. Additionally, Mask3D does not treat floor and wall as instances, resulting in the absence of these two labels in its results. Besides, the original annotations of ScanNet200 are also illustrated.
In Fig. 5, we showcase qualitative comparisons, where our method demonstrates impressive 3D scene segmentation results across a variety of scenes (e.g., bedroom, office, bathroom) and objects, from the holistic view to the focused perspective. Notably, our method consistently outperforms SAM3D in terms of both segmentation quality and diversity across nearly all cases. Moreover, when compared to Mask3D, our method achieves competitive or even superior segmentation accuracy and diversity in a lot of cases. It is also worth highlighting that our segmentation results are not only comparable in accuracy but also exhibit greater diversity compared to human annotations in many cases. More results are provided in the supplementary material.
1.2 Quantitative Results
Traditional closed-set methods (also including OpenMask3D which relies on Mask3D for generating 3D object proposals) are typically trained or finetuned to get 3D object proposals around ScanNet200’s annotated objects . It is suitable to evaluate these methods using mean Average Precision (mAP), which is calculated by examining their 3D segmentations on ScanNet200’s annotated objects. However, since our results are predicted without any guidance from the annotations of ScanNet200, it is infeasible to use mAP to evaluate our method.
To handle this, we devised a new scheme for evaluating the accuracy of segmentation boundaries on human annotations. In this scheme, we begin by selecting an object from the annotations of the ScanNet200 validation set. We then traverse our segmentation outputs to identify all objects that intersect with and denote the intersection area (i.e., number of points) as . If the ratio exceeds a certain threshold , we consider the object to have a probability of belonging to . We then group all such objects into a “grouped” object denoted as . Next, we calculate the Intersection over Union (IoU) between and . This process is repeated for all annotated objects in the ScanNet200 validation set to obtain the final mean IoU (mIoU).
As listed in Tab. 1, we compare the mIoU of our method with SAM3D and Mask3D . A higher mIoU value indicates a better alignment between the predicted and annotated boundaries. Our method substantially surpasses SAM3D and also performs slightly better than Mask3D. This explicitly validates the segmentation accuracy of our method, reinforcing the consistent conclusions drawn from the earlier qualitative results.
We further conducted a subjective user study to assess the segmentation quality and diversity of our method. We selected a set of 20 reference results randomly, with most of them adjusted to a focused view for clearer discrimination. For the study, we invited 50 subjects to participate through an online questionnaire. The majority of the subjects had no prior experience with 3D scene understanding, and none of them had seen our results before. During the study, we presented each subject with five images for each case. included qualitative results from SAM3D , Mask3D , and ScanNet200’s annotations , arranged side by side in random order. Additionally, we included an input image as a reference for comparison. During the evaluation, we instructed the subjects to rank the four results based on two criteria: segmentation accuracy and diversity. Segmentation accuracy aimed to assess the clarity of the segmented boundaries, while diversity focused on determining the extent of whether “segmented anything”. Lastly, the result ranked first by the subjects was assigned a score of 4, while the last-ranked result received a score of only 1.
The mean scores for accuracy (mAcc) and diversity (mDiv) are presented in Tab. 2, demonstrating that our method surpasses both SAM3D and Mask3D by a significant margin. Notably, even when compared to ScanNet200’s annotations, our method achieves slightly higher scores in both quality and diversity. This user study further confirms the efficacy and capability of our method.
2 Ablation Studies and Analysis
We provide comprehensive ablation studies to explore a clearer picture of our method.
In this experiment, we investigate the impact of 2D-Guided Prompt Filter and Prompt Consolidation.
As depicted in Fig. 6 and outlined in Tab. 1, removing the 2D-Guided Prompt Filter leads to a degradation in segmentation accuracy. Additionally, it introduces inaccuracies in the examination of overlap areas during the Prompt Consolidation process, thereby adding complexity to prompt consolidation. This highlights the significance of prompt filtering for achieving segmentation accuracy.
On the other hand, even when Prompt Consolidation is omitted, the boundary accuracy remains consistent, as evidenced by the segmentation mIoU. However, as demonstrated in the qualitative result, this omission leads to fragmented segmentation of certain objects, particularly when the object requires segmentation or coverage from prompts in frames with significant view changes, such as the floor. This observation confirms the advantages of Prompt Consolidation in achieving comprehensive 3D segmentation.
2.2 Design Choices
This is a common question when sampling prompts and defining in the prompt filter. Our method ensures that each 3D prompt segments a single 3D object, bringing stable results when adequate prompts effectively cover the entire scene. We find a principle that maintaining an average of 100 to 150 final kept prompts per scene yields the best results.
Specifically, our method demonstrates consistent and satisfactory performance when an appropriate number of prompts is maintained (Tab. 1: 0.8 and 1.2 on the number of prompts used in the best result). However, doubling the number of final kept prompts introduces more redundant prompts, leading to a degradation in accuracy (Tab. 1: 2). Conversely, using only a few prompts results in incomplete segmentation results (Fig. 6: 0.1).
During the 2D-Guided Prompt Filter, we employ a “hard” voting approach, where a prompt is either kept or discarded based on a binary decision (yes or no). In our experiments, we also explored alternative filtering methods. One approach involved using the actual scores obtained during the filtering process on individual frames and averaging these scores across all frames. We compared the mean score with , to determine whether a prompt should be retained. This method can be referred to as “soft” voting. Additionally, we experimented with another filtering scheme where prompts are kept based on their frequency of being retained across all frames, referred to as the top-k voting scheme. As shown in Tab. 1, both the “soft” and top-k voting schemes yield consistently competitive results. This demonstrates the stability of our method in utilizing different filtering ways.
2.3 Efficiency
Our pipeline exhibits good efficiency, with the majority of computational time and memory usage allocated to the inference process of SAM across the RGB frames. In terms of memory usage, a single GPU with approximately 8000MB is sufficient to run SAM along with our entire pipeline.
Regarding computational time, we provide a breakdown of the time consumed by each step in our framework in Tab. 3. Additionally, a comparison with SAM3D is presented. Similar to SAM3D, our pipeline sequentially processes all frames, allowing room for speed improvement through parallel computation across frames.
In SAM3D, each 2D mask must be projected into 3D masks, which are then iteratively merged based on k-nearest-neighbor search across adjacent frames until achieving the final 3D segmentation of the entire scene. This iterative process increases their time cost.
3 Integrating HQ-SAM and Mobile-SAM
We have successfully integrated HQ-SAM and Mobile-SAM into our framework. As depicted in Fig. 7 and Tab. 1, their impact on 2D images seamlessly translate to the improved performance in our method. These experimental results validate a core concept of our method that “What we can do in 2D = What we can do in 3D”. This crucial insight highlights the importance of considering the holistic system rather than solely focusing on pure 3D data representation in future research endeavors.
Conclusion
We have proposed SAMPro3D, a novel framework for segmenting any 3D indoor scenes by utilizing SAM on 2D frames. Our approach leverages 3D points as natural prompts to align pixel prompts across frames, ensuring consistency in both pixel prompts and their SAM-predicted masks for the same 3D object. Additionally, we propose an algorithm to filter low-quality 3D prompts, enhancing segmentation accuracy. We also design a strategy to consolidate prompts with overlapping masks into a single prompt, achieving more comprehensive segmentation. Our method does not need additional training on domain-specific data, preserving the zero-shot capability of SAM.
As illustrated in Fig. 5, our method may not achieve perfect segmentation for all objects, primarily due to limitations in SAM’s performance. Fortunately, as depicted in Fig. 7, improvements made to SAM can seamlessly enhance the performance of our framework. Furthermore, by leveraging Mobile-SAM and parallel processing of 2D frames on our framework, it is possible to achieve real-time harmonization of 3D scene segmentation and reconstruction in practical applications.
References
Appendix A More Qualitative Results
We present more qualitative comparisons in Fig. I. Following the main paper, we compare our method with SAM3D , Mask3D , and the original annotations of ScanNet200 , on the ScanNet200 validation set. Note that Mask3D does not treat floor and wall as instances, resulting in the absence of these two labels in its results
Consequently, our method consistently achieves remarkable 3D scene segmentation results across diverse scenes and objects, from holistic views to focused perspectives. Notably, our approach significantly outperforms SAM3D in terms of segmentation quality and diversity across nearly all scenarios. Furthermore, when compared to Mask3D, our method demonstrates competitive or superior segmentation quality and diversity in numerous cases. Importantly, our results not only match the quality of human annotations but also exhibit greater diversity in many cases.
Appendix B An Augmented 2D Propagation
In Sec. 1 and Fig. 2 (c) of the main paper, we discuss achieving prompt consistency by using automatic-SAM on an initial frame to generate pixel prompts which can be propagated to subsequent frames, similar to SAM-PT for video tracking. However, videos in SAM-PT exhibit consistently appearing foreground objects, whereas in 3D scenes, objects are scattered across frames. As a result, prompts generated on 2D frames of 3D scenes cannot cover newly emerged objects in other frames, leading to incomplete scene segmentation.
In this section, we evaluate an alternative scheme. Instead of performing automatic-SAM only once on an initial frame, we check if any areas lack segmentation masks in a frame, indicating the presence of newly emerged objects. In such cases, we reapply automatic-SAM to make prompts cover these objects.
As depicted in Fig. II, although the augmented version of 2D propagation improves the completeness of 3D segmentation results, it still falls short in terms of both segmentation quality and diversity. Its deficiency is further highlighted by the comparison of segmentation mIoU in Tab. I. The primary reason behind this inferior performance is that the augmented scheme only aligns pixel prompts within a limited range, from the frame where automatic-SAM is applied to the next time reapplying it. Consequently, the mask consistency is restricted to a few frames. In contrast, our 3D prompts globally align pixel prompts across all frames, resulting in comprehensive frame-consistent pixel prompts and 2D masks, as well as superior 3D segmentation results.
Appendix C Frame Gaps
In the context of performing SAM on 2D image frames, an alternative approach is to skip frames with a certain gap. Fig. III illustrates the qualitative results obtained by skipping frames with different gap numbers. Fig. IV depicts the quantitative results of segmentation mIoU with ScanNet200’s annotations and time cost.
The results indicate that the segmentation quality remains satisfactory with a gap of 5, while there is a slight degradation when using a gap of 10 or 20. Besides, according to Fig. III, our framework stably maintains good segmentation diversity across different gap settings. It’s also worth mentioning that our method consistently outperforms SAM3D in view of both segmentation accuracy and efficiency under different frame gaps, as indicated in Fig. IV. To summarize, increasing the number of frame gaps reduces the number of frames on which SAM is applied, resulting in lower time costs. Therefore, one can choose a suitable frame gap that strikes a balance between segmentation quality and efficiency.
Appendix D Results on Matterport3D
We also applied our framework to the Matterport3D dataset , which contains 194,400 RGB-D images of 90 building-scale indoor scenes. In this dataset, the RGB frames exhibit more view changes compared to ScanNet , which presents additional challenges when performing segmentation solely on 2D frames.
As demonstrated in Fig. V, our method consistently performs well in various scenes of the Matterport3D dataset, from holistic to focused view. This further proves the generalization ability of our method on novel 3D scenes.
Appendix E Implementation Details
When performing SAM segmentation on all 2D frames, we use the PyTorch code provided in the original SAM repository .
Following this code, we begin by converting all the projected pixel coordinates into batched torch tensors represented as an array (BxNx2), where each element corresponds to the (u, v) coordinates in pixels. Next, we pass these tensors to the SAM predictor to process all the pixel prompts in parallel. For this process, we consider all pixel prompts as foreground pixels and assign them a label of 1.
We employ the ViT-H SAM model, which is the default public model of SAM. We resize each RGB image frame to a resolution of 240320. Through experimentation, we found that this resolution is adequate for SAM to deliver satisfactory results. Therefore, we choose this resolution to optimize efficiency in our pipeline.
As mentioned in Sec. 3.2 of the main paper, the first step of 2D-Guided Prompt Filter is to filter prompts on each individual frame. In detail, this filtering process on individual frames involved three steps which are borrowed from the strategy proposed in automatic-SAM . Firstly, we retained only confident masks by applying a threshold of 70.0 to the model’s predicted IoU scores. Secondly, we focused on stable masks by comparing pairs of binary masks derived from the same soft mask. We kept the prediction (i.e., binary mask resulting from thresholding logits at 0) if the IoU between its pair of -1 and +1 thresholded masks was 60.0 or higher. Finally, to eliminate duplicate masks, we employed standard greedy box-based non-maximal suppression (NMS) and cut the masks with a box IoU smaller than 80.0.
As for in 2D-Guided Prompt Filter, we set it to 0.4. We found that flexible values ranging from 0.3 to 0.6 consistently yield decent results, in accordance with the ablation study on the number of prompts in Sec. 4.2.2 of the main paper.
During the Prompt Consolidation, we examine the mask areas of each 3D prompt in 3D space. We denote the intersection area (i.e., number of points) of each of the two prompts as , and the mask area of a single prompt as . If , we consider these two prompts segmenting the same 3D object and consolidate them into one single pseudo prompt. We found that adjusting this overlap threshold to stably performs well.