OpenScene: 3D Scene Understanding with Open Vocabularies

Songyou Peng, Kyle Genova, Chiyu "Max" Jiang, Andrea Tagliasacchi, Marc Pollefeys, Thomas Funkhouser

Introduction

3D scene understanding is a fundamental task in computer vision. Given a 3D mesh or point cloud with a set of posed RGB images, the goal is to infer the semantics, affordances, functions, and physical properties of every 3D point. For example, given the house shown in Figure 1, we would like to predict which surfaces are part of a fan (semantics), made of metal (materials), within a kitchen (room types), where a person can sit (affordances), where a person can work (functions), and which surfaces are soft (physical properties). Answers to these queries can help a robot interact intelligently with the scene or help a person understand it through interactive query and visualization.

Achieving this broad scene understanding goal is challenging due to the diversity of possible queries. Traditional 3D scene understanding systems are trained with supervision from benchmark datasets designed for specific tasks (e.g., 3D semantic segmentation for a closed set of 20 classes ). They are each designed to answer one type of query (is this point on a chair, table, or bed?), but provide little assistance for related queries where training data is scarce (e.g., segmenting rare objects) or other queries with no 3D supervision (e.g., estimating material properties).

In this paper, we investigate how to use pre-trained text-image embedding models (e.g., CLIP ) to assist in 3D scene understanding. These models have been trained from large datasets of captioned images to co-embed visual and language concepts in a shared feature space. Recent work has shown that these models can be used to increase the flexibility and generalizability of 2D image semantic segmentation . However, nobody has investigated how to use them to improve the diversity of queries possible for 3D scene understanding.

We present OpenScene, a simple yet effective zero-shot approach for open-vocabulary 3D scene understanding. Our key idea is to compute dense features for 3D points that are co-embedded with text strings and image pixels in the CLIP feature space (Fig. 2). To achieve this, we establish associations between 3D points and pixels from posed images in the 3D scene, and train a 3D network to embed points using CLIP pixel features as supervision. This approach brings 3D points in alignment with pixels in the feature space, which in turn are aligned with text features, and thus enables open vocabulary queries on the 3D points.

Our 3D point embedding algorithm includes both 2D and 3D convolutions. We first back-project the 3D position of the point into every image and aggregate the features from the associated pixels using multi-view fusion. Next, we train a sparse 3D convolutional network to perform feature extraction from only the 3D point cloud geometry with a loss that minimizes differences to the aggregated pixel features. Finally, we ensemble the features produced by the 2D fusion and the 3D network into a single feature for each 3D point. This hybrid 2D-3D feature strategy enables the algorithm to take advantage of salient patterns in both 2D images and 3D geometry, and thus is more robust and descriptive than features from either domain alone.

Once we have computed features for every 3D point, we can perform a variety of 3D scene understanding queries. Since the CLIP model is trained with natural language captions, it captures concepts beyond object class labels, including affordances, materials, attributes, and functions (Fig. 1). For example, computing the similarity of 3D features with the embedding for “soft” produces the result shown in the bottom-left image of Fig. 1, which highlights the couches, beds, and comfy chairs as the best matches.

Since our approach is zero-shot (i.e. no use of labeled data for the target task), it does not perform as well as fully-supervised approaches on the limited set of tasks for which there is sufficient training data in traditional benchmarks (e.g., 3D semantic segmentation with 20 classes). However, it does achieve significantly stronger performance on other tasks. For example, it beats a fully-supervised approach on indoor 3D semantic segmentation with 40, 80, or 160 classes. It also performs better than other zero-shot baselines, and can be used without any retraining on novel datasets even if they have different label sets. It works for indoor RGBD scans as well as outdoor driving captures.

Overall, our contributions are summarized as follows:

We introduce open vocabulary 3D scene understanding tasks where arbitrary text queries are used for semantic segmentation, affordance estimation, room type classification, 3D object search, and 3D scene exploration.

We propose OpenScene, a zero-shot method for extracting 3D dense features from an open vocabulary embedding space using multi-view fusion and 3D convolution.

We demonstrate that the extracted features can be used for 3D semantic segmentation with performance better than fully supervised methods for rare classes.

Related Work

This paper draws on a large literature of previous work on 3D scene understanding, multi-modal embedding, and zero-shot learning.

Closed-set 3D Scene Understanding. There is a long history of work on 3D scene understanding for vision and robotics applications. Most prior work focuses on training models with ground-truth 3D labels . These works have yielded network architectures and training protocols that have significantly pushed the boundary of several 3D scene understanding benchmarks, including 3D object classification , 3D object detection and localization , 3D semantic and instance segmentation , 3D affordance prediction , and so on. The most closely related work to ours of this type is Rozenberszki et al. , since they use the CLIP embedding to pre-train a model for 3D semantic segmentation. However, they only use the text embedding for point encoder pretraining, and then train the point decoder with 3D GT annotations afterwards. Their focus is on using the CLIP embedding to achieve better supervised 3D semantic segmentation, rather than open-vocabulary queries.

Another line of research performs 3D scene understanding experiments with only 2D ground truth supervisions . For example, generates pseudo 3D annotation by backprojecting and fusing the 2D predicted labels, from which they learn the 3D segmentation task. However, their 2D network is trained with ground truth 2D labels. A couple of works pretrain the 3D segmentation network using point-pixel pairs via contrastive learning between 2D and 3D features. We also utilize 2D image features as our pseudo-supervision when training the 3D network and no 2D labels are needed.

All these approaches have mainly been applied with small predefined labelsets containing common object categories. They do not work as well when the number of object categories increases, as tail classes have few training examples. In contrast, we are able to segment with arbitrary labelsets without any re-training, and we show strong ability of understanding different contents, ranging from rare object types to even materials or physical properties, which is impossible for previous methods.

Open-Vocabulary 2D Scene Understanding. The recent advances of large visual language models have enabled a remarkable level of robustness in zero-shot 2D scene understanding tasks, including recognizing long-tail objects in images. However, the learned embeddings are often at the image level, thus not applicable for dense prediction tasks requiring pixel-level information. Many recent efforts attempt to correlate the dense image features with the embedding from large language models. In this way, given an image at test time, users can define arbitrary text labels to classify, detect, or segment the image.

More recently, Ha and Song take a step forward and perform open-vocabulary partial scene understanding and completion given a single RGB-D frame as input. This method is limited to small partial scenes and requires ground truth training data for supervision. In contrast, in this work, we solely rely on pretrained open-vocabulary 2D models and perform a series of 3D scene-level understanding tasks, without the need for any ground truth training data in 2D or 3D. Moreover,in the absence of 2D images, our method can perform 3D-only open-vocabulary scene understanding tasks based on a 3D point network distilled from an open-vocabulary 2D image model through 3D fusion.

Zero-shot Learning for 3D Point Clouds. While there have been a number of studies on zero-shot learning for 2D images, their application to 3D is still recent and scarce. A handful of works attempt to address the 3D point classification and generation tasks. More recently, investigated zero-shot learning for semantic segmentation for 3D point clouds. They train with supervision of 3D ground truth labels for a predefined set of seen classes and then evaluate on new unseen classes. However, these methods are still limited to the closed-set segmentation setting and still require GT training data for the majority of the 3D dataset. Our method does not require any labeled 3D data for training, and it handles a broad range of queries supported by a large language model.

Method

An overview of our approach is illustrated in Fig. 3. We first compute per-pixel features for every image using a model pre-trained for open-vocabulary 2D semantic segmentation. We then aggregate the pixel features from multiple views onto every 3D point to form a per-point fused feature vector Sec. 3.1. We next distill a 3D network to reproduce the fused features using only the 3D point cloud as input Sec. 3.2. Next, we ensemble the fused 2D features and distilled 3D features into a single per-point feature Sec. 3.3 and use it to answer open-vocabulary queries Sec. 3.4.

The first step in our approach is to extract dense per-pixel embeddings for each RGB image from a 2D visual-language segmentation model, and then back-project them onto the 3D surface points of a scene.

2 3D Distillation

The feature cloud F2D\mathbf{F}^{\text{2D}} can be directly used for language-driven 3D scene understanding when images are present. Nevertheless, such fused features could lead to noisy segmentation due to potentially inconsistent 2D predictions. Moreover, some tasks only provide 3D point clouds or meshes. Therefore, we can distill such 2D visual-language knowledge into a 3D point network that only takes 3D point positions as input.

Specifically, given an input point cloud P\mathbf{P}, we seek to learn an encoder that outputs per-point embeddings:

where F3D={f13D,⋯ ,fM3D}\mathbf{F}^{\text{3D}}=\{\mathbf{f}_{1}^{\text{3D}},\cdots,\mathbf{f}_{M}^{\text{3D}}\}. To enforce the output of the network F3D\mathbf{F}^{\text{3D}} to be consistent with the fused features F2D\mathbf{F}^{\text{2D}}, we use a cosine similarity loss:

We use MinkowskiNet18A as our 3D backbone E3D\mathcal{E}^{\text{3D}}, and change the dimension of outputs to CC.

Since the open-vocabulary image embeddings from are co-embedded with CLIP features, the output of our distilled 3D model naturally lives in the same embedding space as CLIP. Therefore, even without any 2D observations, such text-3D co-embeddings F3D\mathbf{F}^{\text{3D}} allow 3D scene-level understanding given arbitrary text prompts. We show such results in the ablation study in Sec. 4.2.

3 2D-3D Feature Ensemble

Although one can already perform open-vocabulary queries with the 2D fused features F2D\mathbf{F}^{\text{2D}} or 3D distilled features F3D\mathbf{F}^{\text{3D}}, here we introduce a 2D-3D ensemble method to obtain a hybrid feature to yield better performance.

The inspiration comes from the observation that 2D fused features specialize in predicting small objects (e.g. a mug on the table) or ones with ambiguous geometry (e.g., a painting on the wall), while 3D features yield good predictions for objects with distinctive shapes (e.g. walls and floors). We aim to combine the best of both.

Once having the similarity scores wrt. every text prompt tn\mathbf{t}_{n}, we can use the max value s2D=max⁡n(sn2D)\mathbf{s}^{\text{2D}}=\max_{n}(\mathbf{s}^{\text{2D}}_{n}) and s3D=max⁡n(sn3D)\mathbf{s}^{\text{3D}}=\max_{n}(\mathbf{s}^{\text{3D}}_{n}) among all NN prompts as the ensemble scores for both features. Our final 2D-3D ensemble feature f2D3D\mathbf{f}^{\text{2D3D}} is simply the feature with the highest ensemble score.

4 Inference

With any per-point feature described in the previous subsections (f2D\mathbf{f}^{\text{2D}}, f3D\mathbf{f}^{\text{3D}}, or f2D3D\mathbf{f}^{\text{2D3D}}) and CLIP features from an arbitrary set of text prompts, we can estimate their similarities by simply calculating the cosine similarity score between them. We use this similarity score for all of our scene understanding tasks. For example, for the zero-shot 3D semantic segmentation using 2D-3D ensemble features, the final segmentation for each 3D point is computed point-wise by argmax ⁡n{cos(f2D3D,tn)}\operatorname*{argmax~{}}_{n}\{\text{cos}(\mathbf{f}^{\text{2D3D}},\mathbf{t}_{n})\}.

Experiments

We ran a series of experiments to test how well the proposed methods work for a variety of 3D scene understanding tasks. We start by evaluating on traditional closed-set 3D semantic segmentation benchmarks (in order to be able to compare to previous work), and later demonstrate the more novel and exciting open-vocabulary applications in the next section.

Datasets. To test our method in a variety of settings, we evaluate on three popular public benchmarks: ScanNet , Matterport3D , and nuScenes Lidarseg . These three datasets span a broad gamut of situations – the first two provide RGBD images and 3D meshes of indoor scenes, and the last provides Lidar scans of outdoor scenes. We use all three datasets to compare to alternative methods. Moreover, Matterport3D is a complex dataset with highly detailed scenes, and thus provides the opportunity to stress open-vocabulary queries.

Comparison on zero-shot 3D semantic segmentation. We first compare our approach to the most closely related work on zero-shot 3D semantic segmentation: MSeg Voting and 3DGenz . MSeg Voting predicts a semantic segmentation for each posed image using MSeg with mapping to the corresponding label sets. For each 3D point, we perform majority voting of the logits from multi-view images. 3DGenZ divides the 20 classes of the ScanNet dataset into 16 seen and 4 unseen classes, and trains a network utilizing the ground truth supervision on the seen classes to generate features for both sets.

Following the experimental setup in , we report the mIoU and mAcc values on their 4 unseen classes in Table 1. Our results on those classes is significantly better than (7.7% vs 62.8% mIoU), even though 3DGenz utilizes ground truth data for 16 seen classes and ours does not. We also outperform MSeg Voting. In this case, the difference is mainly because our method (regress CLIP features and then classify) naturally models the similarities and differences between classes, where as the MSeg Voting approach (classify and then vote) treats every class as equally distinct from all other classes (a couch and a love seat are just as different as a couch and an airplane in their model).

Comparison on 3D semantic segmentation benchmarks. In Table 2 we compare our approach with both fully-supervised and zero-shot methods on all classes of the nuScenes validation set, ScanNet validation set, and Matterport3D test set. Again, we outperform the zero-shot baseline (MSeg Voting) on both mIoU and mAcc metrics all three datasets. Although we have noticeable gap to the state-of-the-art fully-supervised approaches, our zero-shot method is surprisingly competitive with fully-supervised approaches from a few years ago . Among all 3 datasets our approach has the smallest gap (only -11.6 mIoU and -8.0 mAcc) to the SOTA fully-supervised approach on Matterport3D. We conjecture the reason is that Matterport3D is more diverse, which makes the fully-supervised training harder.

Visual comparisons of semantic segmentations are shown in Fig. 4. They show that some of the predictions marked wrong in our results are actually either incorrect or ambiguous ground truth annotations. For example, in the first row in Fig. 4, we successfully segment the picture on the wall, while the GT misses it. And in the nuScenes results, the truck composed of a trailer and the truck head is segmented correctly, but the annotation is not fine-grained enough to separate the parts.

Impact of increasing the number of object classes. Besides the standard benchmarks with only a small set of classes, we also show comparisons as the number of object classes increases. We evaluate on the most frequent KK classesK=21K=21 was from original Matterport3D benchmark. For K=40,80,160K=40,80,160 we use most frequent KK classes of the NYU label set provided with the benchmark. of Matterport3D, where K=21,40,80,160K=21,40,80,160. For the baseline, we train a separate MinkowskiNet for each KK. However, for ours we use the same model for all KK, since it is class agnostic.

As shown in Table 3 (a), when trained on only 21 classes, the fully-supervised method performs much better due to the rich 3D labels in the most common classes (wall, floor, chair, etc.). However, with the increase of the number of classes, our zero-shot approach overtakes the fully-supervised approach, especially when KK gets large. The reason is demonstrated in Table 3 (b), where we show the mean accuracy for groups of 20 classes ranked by frequency. Fully-supervised suffers severely in segmenting tail classes because there are only a few instances available in the training dataset. In contrast, we are more robust to such rare objects since we do not rely upon any 3D labeled data.

2 Ablation Studies & Analysis

Does it matter which 2D features are used? We tested our method with features extracted from both OpenSeg and LSeg . In most experiments, we found the accuracy and generalizability of OpenSeg features to be better than LSeg (Table 1, Table 2, and Table 4), so we use OpenSeg for all experiments unless explicitly stated otherwise.

Is our 2D-3D ensemble method effective? In Table 4, we ablate the performance for predicting features on 3D points including only image feature fusion (Sec. 3.1), only running the distilled MinkowskiNet (Sec. 3.2), and our full 2D-3D ensemble model (Sec. 3.3). As can be seen, on all scenarios (different datasets, metrics, and 2D features), our proposed 2D-3D ensemble model performs the best. This suggests that leveraging patterns in both 2D and 3D domains makes the ensemble features more robust and descriptive.

What features does our 2D-3D ensemble method use most? Here we study how our ensemble model selects among the 2D and 3D features, and investigate how it changes with increasing numbers of classes in the label set. As shown in Table 5, we find that the majority of predictions (∼70%\sim 70\%) select the 3D features, corroborating the value of our 3D distillation model. However, the percentage of predictions coming from 2D features increases with the number of classes, suggesting that the 2D features are more important for long-tailed classes, which tend to be smaller in both size and number of training examples.

Applications

This section investigates new 3D scene understanding applications enabled by our approach. Since the feature vectors estimated for every 3D point are co-embedded with text and images, it becomes possible to extract information about a scene using arbitrary text and image queries. The following are just a few example applications.Please note that all of these applications are zero-shot – i.e., none of them leverage any labeled data from any 3D scene understanding dataset.

Open-vocabulary 3D object search. We first investigate whether it is possible to query a 3D scene database to find examples based on their names – e.g., “find a teddy bear in the Matterport3D test set.” To do so, we ask a user to enter an arbitrary text string as a query, encode the CLIP embedding vector for the query, and then compute the cosine-similarity of that query vector with the features of every 3D point in the Matteport3D test set (containing 18 buildings with 406 indoor and outdoor regions) to discover the best matches. In our implementation, we return at most one match per region (i.e., room, as defined in the dataset) to ensure diversity of the retrieval results.

Fig. 5 shows a few example top-1 results. Most other specific text queries yield nearly perfect results. To evaluate that observation quantitatively, we chose a sampling of 10 raw categories from the ground truth set of Matterport3D, retrieved the best matching 3D points from the test set, and then visually verified the correctness of the top matches. For each query, Table 6 reports the numbers of instances in the test set (# Test) along with the number of instances found with 100% precision before the first mistake in the ranked list. The results are very encouraging. In all of these queries, only two ground truth instances were missed (two telephones). On the other hand, 26 instances were found among these top matches that were not correctly labeled in the ground truth, including 13 telephones. Overall, these results suggest that our open-vocabulary retrieval application identifies these relatively rare classes at least as well as the manually labeling process did. See the supplemental material for the full set of results.

Image-based 3D object detection. We next investigate whether it is possible to query a 3D scene database to retrieve examples based on similarities to a given input image – e.g., “find points in a Matterport3D building that match this image.” Given a set of query images, we encode them with CLIP image encoder, compute cosine-similarities to 2D-3D ensemble features for 3D points, and then threshold to produce a 3D object detection and mask, see Fig. 7. Note that the pool table and dining table are identified correctly, even though both are types of “tables.”

Open-vocabulary 3D scene understanding and exploration. Finally, we investigate whether it is possible to query a 3D scene to understand properties that extend beyond category labels. Since the CLIP embedding space is trained with a massive corpus of text, it can represent far more than category labels – it can encode physical properties, surface materials, human affordances, potential functions, room types, and so on. We hypothesize that we can use the co-embedding our 3D points with the CLIP features to discover these types of information about a scene.

Fig. 6 shows some example results for querying about physical properties, surface materials, and potential sites of activities. From these examples, we find that the OpenScene is indeed able to relate words to relevant areas of the scene – e.g., the beds, couches, and stuffed chairs match “Comfy,” the oven and fireplace match “Hot,” and the piano keyboard matches “Play.” This diversity of 3D scene understanding would be difficult to achieve with fully supervised methods without massive 3D labeling efforts. In the authors’ opinion, this is the most interesting result of the paper.

Limitations and Future Work

This paper introduces a task-agnostic method to co-embed 3D points in a feature space with text and image pixels and demonstrates its utility for zero-shot, open-vocabulary 3D scene understanding. It achieves state-of-the-art for zero-shot 3D semantic segmentation on standard benchmarks, outperforms supervised approaches in 3D semantic segmentation with many class labels, and enables new open-vocabulary applications where arbitrary text and image queries can be used to query 3D scenes, all without using any labeled 3D data. These results suggest a new direction for 3D scene understanding, where foundation models trained from massive multi-modal datasets guide 3D scene understanding systems rather than training them only with small labeled 3D datasets.

There are several limitations of our work and still much to do to realize the full potential of the proposed approach. First, the inference algorithm could probably take better advantage of pixel features when images are present at test time using earlier fusion (we tried this with limited success). Second, the experiments could be expanded to investigate the limits of open-vocabulary 3D scene understanding on a wider variety of tasks. We evaluated extensively on closed-set 3D semantic segmentation, but provide only qualitative results for other tasks since 3D benchmarks with ground truth are scarce. In future work, it will be interesting to design experiments to quantify the success of open vocabulary queries for tasks where ground truth is not available.

Acknowledgement. We sincerely thank Golnaz Ghiasi for providing guidance on using OpenSeg model. Our appreciation extends to Huizhong Chen, Yin Cui, Tom Duerig, Dan Gnanapragasam, Xiuye Gu, Leonidas Guibas, Nilesh Kulkarni, Abhijit Kundu, Hao-Ning Wu, Louis Yang, Guandao Yang, Xiaoshuai Zhang, Howard Zhou, and Zihan Zhu for helpful discussion. We are also grateful to Charles R. Qi and Paul-Edouard Sarlin for their proofreading.

References

A Implementation Details

3D Distillation. We implement our pipeline in PyTorch . To distill E3D\mathcal{E}^{\text{3D}}, we use Adam as the optimizer with an initial learning rate of 1e−41e{-}4 and train for 100100 epochs. For MinkowskiNet we use a voxel size of 2cm for ScanNet and Matterport3D experiments, and 5cm for nuScenes. For indoor datasets, we input all points of a scene to the 3D backbone to have the full contexts, but for the distillation loss (Eq. 2) in the paper we only supervise with 20K uniformly sampled point features at every iteration due to the memory constraints. For nuScenes, we input all Lidar points within the half-second segments, and only train with point features at the last time stamp. We use a batch size of 8 for ScanNet and Matterport3D with a single NVIDIA A100 (40G). For nuScenes, we use a batch size of 16 with 4 A100 GPUs. It takes around 24 hours to train, and 0.1 seconds for inference. Moreover, for all dataset we only take in the 3D point position as input to the MinkowskiNet during distillation.

More Details of Feature Fusion. For Matterport3D and nuScenes, we use all images of each scene for fusion, while for ScanNet, we sample 1 out of every 20 video frames.

As for the occlusion test, for dataset like ScanNet and Matterport3D where the depth map is provided for each RGB image, we do occlusion test to guarantee that a pixel is only paired with a visible surface point. For every surface point, we first find its corresponding pixel in an image, and we can obtain the distance between that pixel and 3D point. The 3D points and pixel are only paired when the difference between the distance and the depth value of that pixel is smaller than a threshold σ\sigma. The threshold σ\sigma is proportional to the depth value DD. We use σ=0.2D\sigma=0.2D for ScanNet due to the highly noisy depths and σ=0.02D\sigma=0.02D for Matterport. For pixels with “invalid” regions of the depth map, we do not project their features to 3D points.

For nuScenes Lidar points, since no depth images are provided, no occlusion test is conducted, and we only use the synchronized images and the corresponding Lidar points on the last timestamp of a 0.5 second segment.

Our nuScenes Evaluation. Unlike ScanNet and Matterport where each 3D surface point usually having multiple corresponding images, in nuScenes most Lidar points only have one corresponding view, maximum two views at the same time stamp. Therefore, we directly project the single pixel feature to most Lidar points, and use average pooling when there are 2 views. There are 16 classes in the nuScenes lidarseg benchmark, and some of these class names are ambiguous. Since our method can take in arbitrary text prompts, we can pre-define some non-ambiguous classes names for each class, and then map the predictions from these non-ambiguous classes back to the 16 classes. The pre-defined classes names are listed in Table A.

MSeg Voting. MSeg supports a unified taxonomy of 194 classes. We use their official image semantic segmentation codehttps://github.com/mseg-dataset/mseg-semantic and their pretrained MSeg-3m-1080p model. MSeg already provided the mapping from some of 194 classes to 20 ScanNet classes, so we directly use the mapping. For Matterport3D, we simply add the mapping from “ceiling” in the MSeg labelset. As for nuScenes, we manually define the mapping from MSeg to nuScenes 16 labelsets. However, for the “construction vehicles”, “traffic cone”, and ’other flat’, there is no mapping at all, so we set them to unknown.

As for the majority voting for MSeg multi-view predictions, what we do is the following. Given a surface point and its corresponding multi-view MSeg semantic segmentation, we take the class in the majority of views as the voting results for this point. If there are two classes and only two views, we directly flip the coin to decide point labels.

Simple Prompt Engineering. Given a set of text prompts, we use a simple prompt engineering trick before extract CLIP text features. For each object class “XX” (except for “other”) we modify the text prompts to “a XX in a scene”, for instance “a chair in a scene”. With such a simple modification, we observe +2.3 mIoU performance boost with our LSeg ensemble model for ScanNet evaluation. We apply the trick for all our benchmark comparison experiments.

B Additional Analysis

Can we transfer to another dataset with different labelsets?

Here we investigate the ability of our trained models to handle domain transfer between 3D segmentation benchmarks with different labelsets. We train on one dataset (e.g., ScanNet20) and then test on another (e.g., Matterport40) without any retraining (Table B). Since our trained model is task agnostic (it predicts only CLIP features), it does not over-fit to the classes of the training set, and thus can transfer to other datasets with different classes directly. Doing the same using a fully-supervised approach would require a sophisticated domain-transfer algorithm (e.g., ).

We ablate different multi-view feature fusion strategies in Table C. Random means that having multiple features corresponding to one surface point, we randomly assign one feature to the 3D point. For Median, we take the feature that has the smallest Euclidean distance in the feature space to all other features. As can be seen, the simple average pooling yields the best results, and we use it for all our experiments.

Visualization of our 2D-3D ensemble model. In Fig. A we study our ensemble model on how to select 2D and 3D features for prediction on a Matterport3D house based on different labelsets. First, we can notice that our ensemble model uses 3D features for those large areas like floors and walls, while 2D features are preferrable for smaller objects. Second, when comparing the feature selections using 21 and 160 classes, we can see that when the number of classes increases, our ensemble model selects more 2D features for the segmentation. The possible reason is that 2D image features can better understand those fine-grained concept than purely from 3D point clouds. For example, on the bottom-right there is a pool table there. When using 21 class labels, it is segmented as a table, so 3D features are preferrable. When using 160 class labels for 2D-3D ensemble, it is much easier to understand the concept of “pool table” using 2D images than 3D point clouds.

Definition of Zero-Shot Learning (ZSL). The terminology for ZSL is ambiguous in the literature. In a theoretical ZSL system, there should be no training data of any kind from seen classes. However, almost all real-world ZSL systems utilize general-purpose feature extractors pretrained for proxy tasks on large datasets. For example, 3DGenZ proposes a ZSL variant utilizing image features pretrained on ImageNet (see section 4.5 in their paper), while OpenSeg , LSeg , CLIP , and ALIGN propose ZSL methods trained on alt-text, as we also do. The authors of those papers all describe their methods as ZSL. We followed the same terminology as the latter one.

C Full Results of Open-vocabulary Object Retrieval

This section details the experiments in the main paper on open-vocabulary search in a 3D scenes database (Figure 4, Table 6). We employed our 2D-3D ensemble method to produce features for the Matterport3D test set’s mesh vertices, using NYU160 as the labelset for ensembling. For each given search query, the vertices were sorted by cosine similarity to the CLIP text query embedding, resulting in a ranked retrieval list with one match per region.

When selecting the queries to use in the experiment, we limited our selections to raw category names provided with the ground truth of the dataset. This allows us to reason about how many examples exist in the test set so that we can know how many matches to expect. Among those candidate queries, we used two strategies to select ones to test: 1) we chose several of the most specific categories (e.g., “yellow egg-shaped vase”) in order to test the method on the most difficult cases, and 2) we chose all raw categories with 15 ground truth examples in the entire Matterport3D dataset that had clear definitions (which excepted “lounge chair,” “side table,” and “office table”) in order to avoid bias in the selection of queries (i.e., no cherry-picking).

Figures B-E display ranked retrieval lists, with the best match first and expected matches in parentheses. A red wireframe sphere highlights the top match. Green/red-bordered images show correct/incorrect matches, with incorrect ones ranked lower than the last ground truth instance. Gray-bordered images demonstrate near misses, not expected as matches due to their rank exceeding labeled ground truth examples.

We can see that the algorithm is able to retrieve very specific objects from the database with great precision. For example, when queried with “yellow egg-shaped vase,” its top match is indeed a match (which was not labeled in the ground truth), and the following retrieval results are tan vase, a pumpkin, and a white egg-shaped vase with gold decorations. Similarly, when queried with “teddy bear,” it retrieves two teddy bears (neither labeled in the ground truth), a stuffed monkey, and a stuffed lion among the top four matches. Among all the queries in all of the experiments, The only false positive occurred with “telephone” where a bowl of stones ranked 15th, while two ground truth instances ranked 25th and 29th. In this case, 29 of the top 30 matches were correct (20 are shown in Figure D).

These results suggest that the open-vocabulary features computed with our 2D-3D ensemble algorithm are very effective at retrieving object types with specific names. Further experiments are required to understand the limitations.

D More Results of Open-vocabulary 3D Scene Exploration

The main paper demonstrates that open-vocabulary queries can be used to explore the content of 3D scenes via text queries (Figures 1 and 6 of the main paper). This section provides more examples demonstrating the power of open-vocabulary exploration. Please see the supplemental video for a live demo.

For each query, the user types a text string (e.g., “glass”), it gets encoded with the pre-trained CLIP text encoder, we compute the cosine similarity with features we’ve computed for every 3D vertex, and then color the vertices by similarity (yellow is high, green is middle, blue is low).

Figures F-K show results for a broad range of queries, including ones that describe object categories in Fig. F, room types in Fig. G, activities in Fig. H, colors in Fig. I, materials in Fig. J, and abstract concepts in Fig. K.

Please note the power of using language models learned via CLIP to reason about scene attributes and abstract concepts that would be difficult to label in a supervised setting. For example, searching for “store” highlights 3D points mainly on closets and cabinets (middle-right of Fig. H), and searching for “cluttered” yields points in a particularly busy closet (top-right of Fig. K). These examples demonstrate the power of the proposed approach for scene understanding, which goes far beyond semantic segmentation.

Quantitative results on 3DSSG dataset . To evaluate 3D scene exploration performance, we conduct an experiment on the 3DSSG dataset that has annotations in object-level material estimation. In Table D, we compare material class predictions for the 3DSSG test set using variants of our approach trained on ScanNet and a fully-supervised MinkowskiNet. Findings align with the paper: 1) 2D-3D ensembling is our best variant, 2) it underperforms fully-supervised methods for classes with abundant examples, and 3) it excels for classes with fewer examples.