Segment Any 3D Gaussians

Jiazhong Cen, Jiemin Fang, Chen Yang, Lingxi Xie, Xiaopeng Zhang, Wei Shen, Qi Tian

Introduction

Interactive 3D segmentation in radiance fields has attracted a lot of attention from researchers, due to its potential applications in various domains like scene manipulation, automatic labeling, and virtual reality. Previous methods predominantly involve lifting 2D visual features into 3D space by training feature fields to imitate multi-view 2D features extracted by self-supervised visual models . Then the 3D feature similarities are used to measure whether two points belong to the same object. Such approaches are fast due to their simple segmentation pipeline, but as a price, the segmentation granularity may be coarse since they lack the mechanism for parsing the information embedded in the features (e.g., a segmentation decoder). In contrast, another paradigm proposes to lift the 2D segmentation foundation model to 3D by projecting the multi-view fine-grained 2D segmentation results onto 3D mask grids directly. Though this approach can yield precise segmentation results, its substantial time overhead restricts interactivity due to the need for multiple executions of the foundation model and volume rendering. Specifically, for complex scenes with multiple objects requiring segmentation, this computational cost becomes unaffordable.

The above discussion reveals the dilemma of currently existing paradigms in achieving both efficiency and accuracy, pointing out two factors that limit the performance of existing paradigms. First, implicit radiance fields employed by previous approaches hinder efficient segmentation: the 3D space must be traversed to retrieve a 3D object. Second, the utilization of the 2D segmentation decoder brings high segmentation quality but low efficiency.

Accordingly, we revisit this task starting from the recent breakthrough in radiance fields: 3D Gaussian Splatting (3DGS) has become a game changer because of its ability in high-quality and real-time rendering. It adopts a set of 3D colored Gaussians to represent the 3D scene. The mean of these Gaussians denotes their position in the 3D space thus 3DGS can be seen as a kind of point cloud, which helps bypass the extensive processing of vast, often empty, 3D spaces and provides abundant explicit 3D prior. With this point cloud-like structure, 3DGS not only realizes efficient rendering but also becomes as an ideal candidate for segmentation tasks.

On the basis of 3DGS, we propose to distill the fine-grained segmentation ability of a 2D segmentation foundation model (i.e., the Segment Anything Model) into the 3D Gaussians. This strategy marks a departure from previous methods that focuses on lifting 2D visual features to 3D and enables fine-grained 3D segmentation. Moreover, it avoids the time-consuming multiple forwarding of the 2D segmentation model during inference. The distillation is achieved by training 3D features for Gaussians based on automatically extracted masks with the Segment Anything Model (SAM) . During inference, a set of queries are generated with input prompts, which, are then used to retrieve the expected Gaussians through efficient feature matching.

Named as Segment Any 3D GAussians (SAGA), our approach can achieve fine-grained 3D segmentation in milliseconds and support various kinds of prompts including points, scribbles and masks. Evaluation on existing benchmarks demonstrates the segmentation quality of SAGA is on par with previous state-of-the-art.

As the first attempt of interactive segmentation in 3D Gaussians, SAGA is versatile, accommodating a range of prompt types, including masks, points, and scribbles. Our evaluation on existing benchmarks demonstrates that SAGA performs on par with the state-of-the-art. Notably, the training of Gaussian features typically concludes within merely 5-10 minutes. Subsequently, the segmentation of most target objects can be completed in milliseconds, achieving nearly 1000×1000\times acceleration.

Related Work

Inspired by natural language processing and recent computer vision progress, Kirillov et al. proposed the task of promptable segmentation. The goal of this task is to return segmentation masks given input prompts that specify the segmentation target in an image. To solve this problem, they present the Segment Anything Model (SAM), a revolutionary segmentation foundation model. An analogous model to SAM is SEEM , which also exhibits impressive open-vocabulary segmentation capabilities. Before them, the most closely related task to promptable 2D segmentation is the interactive image segmentation, which have been explored by many studies .

Lifting 2D Vision Foundation Models to 3D

Recently, 2D vision foundation models have experienced robust growth. In contrast, 3D vision foundation models have not seen similar development, primarily due to the scarcity of data. Acquiring and annotating 3D data is significantly more challenging than its 2D counterpart. To tackle this problem, researchers attempted to lift 2D foundation models to 3D . A noteworthy attempt is LERF , which trains a feature field of the Vision-Language Model (i.e., CLIP ) together with the radiance field. Such paradigm helps locating objects in radiance fields based on language prompts but falls short in precise 3D segmentation, especially when faced with multiple objects of similar semantics. The remaining methods mainly focus on point clouds. By associating the 3D point cloud with 2D multi-view images with the help of camera poses, the extracted features by 2D foundation models can be projected to the 3D point cloud. Such integration is similar to LERF but incurs a higher data acquisition cost compared to radiance field-based methods.

D Segmentation in Radiance Fields

Inspired by the success of radiance fields , numerous studies have explored 3D segmentation within them. Zhi et al. proposes Semantic-NeRF, which demonstrates the potential of Neural Radiance Field (NeRF) in semantic propagation and refinement. NVOS introduces an interactive approach to select 3D objects from NeRF by training a lightweight multi-layer perception (MLP) using custom-designed 3D features. By using 2D self-supervised models, e.g. N3F , DFF , LERF and ISRF , aim to lift 2D visual features to 3D by training additional feature fields that can output 2D feature maps imitating the original 2D features in different views. NeRF-SOS distills the 2D feature similarities into 3D features with a correspondence distillation loss . In these 2D visual feature-based approaches, 3D segmentation can be achieved by comparing the 3D features embedded in the feature field, which appears to be efficient. However, since the information embedded in the high-dimensional visual features cannot be fully exploited when relying solely on Euclidean or cosine distances, the segmentation quality of such methods is limited. There are also some other instance segmentation and semantic segmentation approaches combined with radiance fields.

Two most closely related approach to our SAGA is ISRF and SA3D . The former follows the paradigm of training a feature field to imitate multi-view 2D 2D visual features. Thus it struggles with distinguishing different objects (especially parts of object) with similar semantics. The latter iteratively queries SAM to get 2D segmentation results and projecting them onto mask grids for 3D segmentation. Though good segmentation quality, its complex segmentation pipeline leads to high time consumption and inhibits the interaction with users. Compared with them, SAGA can handle multi-granularity 3D segmentation within milliseconds and achieve a better trade-off between the segmentation quality and efficiency.

Methodology

As a recent advancement of radiance fields, 3DGS uses trainable 3D Gaussians to represent the 3D scene and proposes an efficient differentiable rasterization algorithm for rendering and training. Given a training dataset I\mathcal{I} of multi-view 2D images with camera poses, 3DGS learns a set of 3D colored Gaussians G={g1,g2,...,gN}\mathcal{G}=\{\mathbf{g}_{1},\mathbf{g}_{2},...,\mathbf{g}_{N}\}, where NN denotes the number of 3D Gaussians in the scene. The mean of each Gaussian represents its position in the 3D space and the covariance represents the scale. Thus 3DGS can be regarded as a kind of point cloud. Given a specific camera pose, 3DGS projects the 3D Gaussians to 2D and then computes the color C\mathbf{C} of a pixel by blending a set of ordered Gaussians N\mathcal{N} overlapping the pixel:

where ci\mathbf{c}_{i} is the color of each Gaussian and αi\alpha_{i} is given by evaluating a 2D Gaussian with covariance Σ\Sigma multiplied with a learned per-Gaussian opacity. From Eq. 1 we can learn the linearity of the rasterization process: the color of a rendered pixel is the weighted sum of the involved Gaussians. In our framework, such characteristic ensures the alignment of 3D features with the 2D rendered features.

Segment Anything Model (SAM)

SAM takes an image I\mathbf{I} and a set of prompts P\mathcal{P} as input, and outputs the corresponding 2D segmentation mask M\mathbf{M}, i.e.,

2 Overall Pipeline

3 Training Features for Gaussians

Given a training image I\mathbf{I} with its specific camera pose vv, we first render the corresponding feature map according to the pre-trained 3DGS model G\mathcal{G}. Similar to Eq. 1, the rendered feature FI,pr\mathbf{F}^{r}_{\mathbf{I},p} of a pixel pp is computed as:

where N\mathcal{N} is the ordered set of Gaussians overlapping the pixel. During the training phase, we freeze all other attributes of the 3D Gaussians G\mathcal{G} (e.g., mean, covariance and opacity) except the newly attached features.

The automatically extracted 2D masks MI\mathcal{M}_{\mathbf{I}} via SAM are complex and confusing (i.e., a point in the 3D space may be segmented as different objects / parts on different views). Such ambiguous supervision signal poses a great challenge to training 3D features from scratch. To tackle this problem, we propose to use the features generated by SAM for guidance. As shown in Fig. 2, we first adopt an MLP φ\varphi to project the SAM features to the same low-dimensional space as the 3D features:

where σ\sigma denotes the element-wise sigmoid function. The SAM-guidance loss is defined as the binary cross entropy between the segmentation result PM\mathbf{P}_{\mathbf{M}} and the corresponding SAM extracted mask M\mathbf{M}:

Correspondence Loss

In practice, we find the learned features with the SAM-guidance loss are not compact enough, which degrades the segmentation quality of various kinds of prompts (refer to the ablation study in Sec. 4 for more details). Inspired by previous contrastive correspondence distillation methods , we introduce the correspondence loss to tackle the problem.

As mentioned before, for each image I\mathbf{I} with height HH and width WW in the training set I\mathcal{I}, a set of masks MI\mathcal{M}_{\mathbf{I}} are extracted with SAM. Considering two pixels p1,p2p_{1},p_{2} in I\mathbf{I}, they may belong to many masks in MI\mathcal{M}_{\mathbf{I}}. Let MIp1,MIp2\mathcal{M}_{\mathbf{I}}^{p_{1}},\mathcal{M}_{\mathbf{I}}^{p_{2}} denote the masks that p1,p2p_{1},p_{2} belong to respectively. Intuitively, if the intersection over union of the two sets is larger, the two pixels should share more similar features. Thus the mask correspondence KI(p1,p2)\mathbf{K}_{\mathbf{I}}(p_{1},p_{2}) is defined as:

The feature correspondence SI(p1,p2)\mathbf{S}_{\mathbf{I}}(p_{1},p_{2}) between two pixels p1,p2p_{1},p_{2} is defined as the cosine similarity between their rendered features:

then the correspondence loss is defined as:

If two pixels never belong to the same segment, we reduce their feature similarity by setting the 0-valued entries in KI\mathbf{K}_{\mathbf{I}} to −1-1.

With the two components of the SAM-guidance loss (Eq. 7) and the correspondence loss (Eq. 10), the final loss of SAGA is:

where λ\lambda is a hyper-parameter for balancing the two loss terms (set to 1 in default).

4 Inference

Though the training is performed on the rendered feature maps, the linearity of the rasterization operation (shown in Eq. 3) ensures that the features in the 3D space are aligned with the rendered features on the image plane. Thus, the segmentation of the 3D Gaussians can be achieved with 2D-rendered features. This characteristic endows SAGA with the compatibility with different kinds of prompts including points, scribbles and masks. Moreover, we introduce an efficient post-processing algorithm (Sec. 3.5) based on the 3D prior provided by 3DGS.

With a rendered feature map Fvr\mathbf{F}^{r}_{v} for a specific view vv, we generate queries for positive points and negative points by directly retrieving their corresponding features on Fvr\mathbf{F}^{r}_{v}. Let Qvp\mathcal{Q}^{p}_{v} and Qvn\mathcal{Q}^{n}_{v} denote the NpN_{p} positive queries and NnN_{n} negative queries respectively. For a 3D Gaussian g\mathbf{g}, its positive score SgpS^{p}_{\mathbf{g}} is defined as the maximum cosine similarity between its feature fg\mathbf{f}_{\mathbf{g}} and the positive queries Qvp\mathcal{Q}^{p}_{v}, i.e., max⁡{<fg,Qp>∣Qp∈Qvp}\max\{<\mathbf{f}_{\mathbf{g}},\mathbf{Q}^{p}>|\mathbf{Q}^{p}\in\mathcal{Q}^{p}_{v}\}. Similarly, the negative score SgnS^{n}_{\mathbf{g}} is defined as max⁡{<fg,Qn>∣Qn∈Qvn}\max\{<\mathbf{f}_{\mathbf{g}},\mathbf{Q}^{n}>|\mathbf{Q}^{n}\in\mathcal{Q}^{n}_{v}\}. The 3D Gaussian belongs to the target Gt\mathcal{G}^{t} only if Sgp>SgnS^{p}_{\mathbf{g}}>S^{n}_{\mathbf{g}}.

To further filter out noisy Gaussians, an adaptive threshold τ\tau is set to the positive score, i.e., g∈Gt\mathbf{g}\in\mathcal{G}^{t} only if Sgp>τS^{p}_{\mathbf{g}}>\tau. τ\tau is set as the mean of the maximum positive scores. Note that such filtering may cause many false negatives, but can be solved by the post-processing introduced in Sec. 3.5.

Mask And Scribble Prompts

Simply treating the dense prompts as multiple points will lead to unaffordable GPU memory overhead. Thus we employ the K-means algorithm to extract some positive queries Qvp\mathcal{Q}^{p}_{v} and negative queries Qvn\mathcal{Q}^{n}_{v} from the dense prompts. The number of clusters of K-means is set to 5 empirically, but is adjustable according to the complexity of the target object.

SAM-based Prompt

The previous prompts are obtained from rendered feature maps. With the SAM-guidance loss, we can directly use the low-dimensional SAM features Fv′\mathbf{F}^{\prime}_{v} for generating queries. The input prompts are first fed into SAM for generating accurate 2D segmentation result Mvref\mathbf{M}^{\text{ref}}_{v}. With this 2D mask, we first obtain a query Qmask\mathbf{Q}^{\text{mask}} with the masked average pooling and use this query to segment the 2D rendered feature map Fvr\mathbf{F}^{r}_{v} to get a temporary 2D segmentation mask Mvtemp\mathbf{M}^{\text{temp}}_{v}, which is then compared with Mvref\mathbf{M}^{\text{ref}}_{v}. If the intersecting region of Mvtemp\mathbf{M}^{\text{temp}}_{v} and Mvref\mathbf{M}^{\text{ref}}_{v} occupies a large proportion (90%, by default) of Mvref\mathbf{M}^{\text{ref}}_{v}, Qvmask\mathbf{Q}^{\text{mask}}_{v} is accepted as the query. Otherwise, we use the K-means algorithm to extract another set of queries Qvkmeans\mathcal{Q}^{\text{kmeans}}_{v} from the low-dimensional SAM features Fv′\mathbf{F}^{\prime}_{v} within the mask. We adopt such strategy because that the segmentation target may contain many components, which cannot be captured by simply applying the masked average pooling.

After obtaining the query set QvSAM={Qvmask}\mathcal{Q}^{\text{SAM}}_{v}=\{\mathbf{Q}^{\text{mask}}_{v}\} or QvSAM=Qvkmeans\mathcal{Q}^{\text{SAM}}_{v}=\mathcal{Q}^{\text{kmeans}}_{v}, the subsequent process is almost the same as the former prompt approaches. We use the point product instead of the cosine similarity as the metric for segmentation to align with the SAM-guidance loss. For a 3D Gaussian g\mathbf{g}, its positive score SgpS^{p}_{\mathbf{g}} is defined as the maximum point product computed with these queries:

The 3D Gaussian g\mathbf{g} belongs to the segmentation target Gt\mathcal{G}^{t} if its positive score is greater than another adaptive threshold τSAM\tau^{\text{SAM}}, which is the sum of the mean and the standard deviation of all scores SG={Sgp∣g∈G}\mathcal{S}_{\mathcal{G}}=\{S^{p}_{\mathbf{g}}|\mathbf{g}\in\mathcal{G}\}.

5 3D Prior Based Post-processing

The initial segmentation Gt\mathcal{G}^{t} of the 3D Gaussians exhibits two primary problems: (i) the presence of superfluous noisy Gaussians and (ii) the omission of certain Gaussians integral to the target object. To tackle the problem, we utilize traditional point cloud segmentation techniques , including statistical filtering and region growing. For segmentation based on point and scribble prompts, statistical filtering is employed to filter out noisy Gaussians. For mask prompts and SAM-based prompts, the 2D mask is projected onto Gt\mathcal{G}^{t} to get a set of validated Gaussians and projected onto G\mathcal{G} to exclude unwanted Gaussians. The resulting validated Gaussians serve as the seed for the region-growing algorithm. Finally, a ball query-based region growing method is applied to retrieve all required Gaussians of the target from the original model G\mathcal{G}.

The distance between two Gaussians can indicate whether they belong to the same target. Statistical filtering begins by employing the K-Nearest Neighbors (KNN) algorithm to calculate the average distance of the nearest ∣Gt∣\sqrt{|\mathcal{G}^{t}|} Gaussians for each Gaussian within the segmentation result Gt\mathcal{G}^{t}. Subsequently, we compute the mean (μ\mu) and standard deviation (σ\sigma) of these average distances across all Gaussians in Gt\mathcal{G}^{t}. We then remove Gaussians with an average distance exceeding μ+σ\mu+\sigma to get Gt′\mathcal{G}^{t\prime}.

Region Growing Based Filtering

The 2D mask from mask prompt or SAM-based prompt can serve as a prior for accurately localizing the target. Initially, we project the mask onto the segmented Gaussians Gt\mathcal{G}^{t}, yielding a subset of validated Gaussians, denoted as Gc\mathcal{G}^{c}. Subsequently, for each Gaussian g\mathbf{g} within Gc\mathcal{G}^{c}, we compute its Euclidean distance dgd_{\mathbf{g}} to its closest neighbor in the same subset:

where D(⋅,⋅)D(\cdot,\cdot) denotes the Euclidean distance. Then we iteratively incorporate neighboring Gaussians in Gt\mathcal{G}^{t} whose distances are less than the maximum nearest neighbor distance observed in the set Gc\mathcal{G}^{c}, formalized as max⁡{dgGc∣g∈Gc}\max\{d_{\mathbf{g}}^{\mathcal{G}^{c}}|\mathbf{g}\in\mathcal{G}^{c}\}. After the region growing converging, where no new Gaussians in Gf\mathcal{G}^{f} meet the criteria, we get the filtered segmentation result Gt′\mathcal{G}^{t\prime}.

Note that though the point prompt and scribble prompt can also roughly locate the target, region growing based on them is time-consuming. Thus we only apply the region growing based filtering when a mask is available.

Ball Query Based Growing

The filtered segmentation output Gt′\mathcal{G}^{t\prime} may not contain all Gaussians belong to the target. To address this problem, we utilize a ball query algorithm to retrieve all required Gaussians from all Gaussians G\mathcal{G}. Concretely, this is achieved by checking spherical neighborhoods with a radius rr, centered at each Gaussian in Gt′\mathcal{G}^{t\prime}. Gaussians that are located within these spherical boundaries in G\mathcal{G} are then aggregated into the final segmentation result Gs\mathcal{G}^{s}. The radius rr is set to be the maximum nearest neighbor distance in Gt′\mathcal{G}^{t\prime}, i.e., r=max⁡{dgGt′∣g∈Gt′}r=\max\{d_{\mathbf{g}}^{\mathcal{G}^{t\prime}}|\mathbf{g}\in\mathcal{G}^{t\prime}\}.

Experiments

For quantitative experiments, we use the Neural Volumetric Object Selection (NVOS) , SPIn-NeRF datasets. The NVOS dataset is based on the LLFF dataset , which includes several forward-facing scenes. For each scene, the NVOS dataset provides a reference view with scribbles and a target view with 2D segmentation masks annotated. Similarly, the SPIn-NeRF dataset also annotates some data manually based on widely-used NeRF datasets . Futhermore, we also use SA3D to annotate some objects in the LERF-figurines scene to demonstrate the better trade-off of efficiency and segmentation quality achieved by SAGA. For qualitative analysis, we use the LLFF dataset, the MIP-360 dataset , the T&T dataset and the LERF dataset .

2 Quantitative Results

We follow SA3D to process the scribbles provided by the NVOS dataset to meet the requirements of SAM. As shown in Table 1, SAGA is on par with previous SOTA SA3D and significantly outperforms previous feature imitation-based approach (ISRF and SGISRF), which demonstrates its fine-grained segmentation ability.

SPIn-NeRF

We follow SPIn-NeRF to conduct label propagation for evaluation, which specifies a view with its 2D ground-truth mask and propagate this mask to other views to check the mask accuracy. This operation can be seen as a kind of mask prompt. The results are shown in Table 2. MVSeg adopts the video segmentation approach to segment the multi-view images and SA3D automatically queries 2D segmentation foundation model for rendered images on the training views. Both of them need to forward a 2D segmentation model for many times. Remarkably, SAGA shows comparable performance with them in nearly one-thousandth of the time. Note that the slight degradation is caused by the sub-optimal geometry learned by 3DGS. Please refer to Sec. 4.3 for more details.

Comparison with SA3D

To further demonstrate the effectiveness of SAGA, we compare the segmentation time consumption and the quality with SA3D. We run SA3D based on the LERF-figurines scene to get a set of annotations for many objects. Subsequently we use SAGA to segment the same objects and check the IoU and time cost for each object. The results are shown in Table 3, We also provide visualization results for comparison with SA3D, please refer to Sec. 4.3 for more details. It is noteworthy to mention that limited by the huge GPU memory cost of SA3D, the training resolution of SAGA is much higher. This indicates that SAGA can get 3D assets with higher quality in much less time. Even considering the training time (about 10 minutes per scene), the average segmentation time for each object of SAGA is much less than SA3D.

3 Qualitative Results

We begin by establishing that SAGA attains a segmentation accuracy on par with the prior SOTA, SA3D, while significantly reducing time cost. Subsequently, we demonstrate the enhanced performance of SAGA over ISRF, in both part and object segmentation tasks. Results are shown in Fig. 3.

The first row shows the segmentation results of SA3D and SAGA on the LERF-figurines scene, with segmentation times annotated in the lower right of each segmented object. The second row compares SAGA with ISRF, which trains a feature field by imitating the 2D features extracted by a self-supervised vision transformer (e.g., DINO ). ISRF struggles to differentiate between objects of similar semantics, like parts of the T-Rex skeleton. In contrast, SAGA distills the knowledge embedded in the SAM decoder into the feature field, thereby adeptly managing such complexities. Additional segmentation results for the MIP-360-counter and T&T-truck scenes are presented in the third row. It’s important to note the noise present at the periphery of the segmented targets. This is attributed to the inherent properties of 3D Gaussians, where a certain Gaussian intersect multiple objects, particularly at the boundaries where different objects meet.

In Table 2, SAGA exhibits sub-optimal performance compared to the previous state-of-the-art methods. This is because of a segmentation failure of the LLFF-room scene, which reveals a limitation of SAGA. We show the mean of the colored Gaussians in Fig. 4, which can be seen as a kind of point cloud. SAGA is susceptible to inadequate geometric reconstruction of the 3DGS model. As marked by the red boxes, the Gaussians of the table is notably sparse, where the Gaussians representing the table surface are floating beneath the actual surface. Even worse, the Gaussians from the chair are in close proximity to those of the table. These issues not only impede the learning of discriminative 3D features but also compromise the efficacy of the post-processing. We believe that enhancing the geometric fidelity of the 3DGS model can ameliorate this issue.

4 Ablation Study

Our loss function comprises two key components: 1) SAM-guidance loss and 2) Correspondence loss. We demonstrate their efficacy quantitatively and qualitatively. As shown in Table 4, the absence of SAM-guidance loss significantly hinders the performance of SAGA in complex scenes, such as LLFF-fern, due to the ambiguous nature of the segmentation targets. Furthermore, as indicated in Fig. 5, excluding the correspondence loss leads to less compact 3D features. This affects the effectiveness of various kinds of prompts and makes the SAM-based prompt as the sole effective approach.

Post-processing

As shown in Fig. 6, without the post-processing there are some noisy Gaussians in the segmentation result and the segmentation target (the flowers) seems translucent due to the missing Gaussians.

Computation Consumption

We analyse the time cost of SAGA based on the T&T-truck scene and the LERF-figurines scene . The segmentation target for the former is the truck and for the latter is the green apple on the table. Both of them can be found in Fig. 3. As shown in Table 5, for large targets, the primary computation lies in post-processing. In contrast, for the smaller target, the time cost of Gaussians retrieving becomes the main consumption, which depends on the complexity of the scene.

Limitation

SAGA requires training features for 3D Gaussians, which makes it more suitable for scenes with multiple objects to be segmented than object-centric scenes. Besides, the primary limitations of SAGA stem from 3DGS and SAM, which can be summarized as follows:

The Gaussians learned by 3DGS are ambiguous without any constraint on geometry. A single Gaussian might correspond to multiple objects, complicating the task of accurately segmenting individual objects through feature matching. We believe this issue can be alleviated by future progress in the 3DGS representation.

The masks automatically extracted by SAM tend to exhibit a certain level of noise as a byproduct of the multi-granularity characteristic. This can be alleviated by adjusting the hyper-parameters involved in automatic mask extraction.

Additionally, it’s important to note that the post-processing step in SAGA is semantic-agnostic, which may bring some false positive points into the segmentation result. We leave this issue as future work.

Conclusion

In this paper, we introduce SAGA, a novel interactive 3D segmentation method. As the first attempt of interactive segmentation in 3D Gaussians, SAGA effectively distills knowledge from the Segment Anything Model (SAM) into 3D Gaussians using two carefully designed losses. After training, SAGA allows for rapid, millisecond-level 3D segmentation across various input types like points, scribbles, and masks. Extensive experiments are conducted to demonstrate the efficiency and effectiveness of SAGA.

References