CLIP-FO3D: Learning Free Open-world 3D Scene Representations from 2D Dense CLIP

Junbo Zhang, Runpei Dong, Kaisheng Ma

Introduction

3D scene understanding aims to distinguish objects’ semantics, identify their locations, and infer the geometric attributes from 3D scene data. It has a wide range of applications in virtual reality , robot navigation and autonomous driving . However, training traditional 3D scene understanding models requires a large number of human annotations, which are laborious to collect. Besides, the human annotations used in the current 3D scene understanding benchmark only contain close-set semantic information of the objects (e.g., 20 classes in ScanNet ). These make it difficult for 3D scene understanding systems to recognize open-vocabulary categories and to infer open-world semantics.

Large-scale vision-language foundation models (e.g., CLIP ) capture rich visual and language features. They only require image-text pairs mined from the Internet for unsupervised pre-training, and have demonstrated superior ability in zero-shot and open-vocabulary reasoning for classification and dense prediction tasks . However, due to the difficulty of collecting 3D-text pairs and the complexity of scene data, is extremely challenging to build analogous 3D foundation models. Moreover, transferring CLIP’s capabilities to 3D scene understanding models is still an understudied problem.

In this paper, we propose directly transferring CLIP’s feature space to 3D scene understanding models without any supervision (e.g., 2D/3D annotations or vision-language grounding annotations). Once the transfer is complete, our model can complete the open-vocabulary 3D scene understanding task without additional annotation or training. In this way, we intend to preserve the rich information inherited from CLIP model to the maximum. It is because that training with any human annotations restricts the feature space to a limited set of vocabularies and loses CLIP’s powerful open-world properties.

Recent methods in 3D vision that directly distill CLIP’s open-vocabulary knowledge to 3D models focus only on the object-level data . By contrast, it is much more difficult to directly transfer CLIP’s feature space to 3D models on scene-level data. Because CLIP’s vision encoder only extract image-level global feature, while 3D scene understanding requires dense point-level features. Although MaskCLIP in 2D vision has taken the first step towards extracting pixel-level features from CLIP’s final feature map, it has difficult adapting to 3D scene’s contents since they contain much more objects and more complex structures. Besides, it has an inherent defect in precisely locating and segmenting objects’ boundaries due to the limited resolution of CLIP’s final feature map.

To address the abovementioned problems, we present a new method to extract free pixel-level CLIP features without breaking CLIP’s feature space. We first crop the input images at multi-scale to accommodate various object sizes. To preserve object-level semantics, we divide each cropped sample into semantically relevant regions and extract a local feature for each region. In the ViT encoder, we add several local classification tokens, which only aggregate information from local patches within a region. By simply forwarding CLIP’s encoder, we extract a pixel-level feature map of each 3D scene’s RGB view.

After extracting pixel-level CLIP features, we adopt the feature projection scheme in 3DMV to project multi-view image features to the point cloud. The resulting point features are aligned with CLIP’s feature space and are used as off-line training targets. Then we train the 3D understanding model with feature distillation, minimizing the distance between learned point features and the target features. This way, we obtain CLIP-FO3D, which extracts free and open-world 3D scene representations aligned with CLIP.

Our CLIP-FO3D can perform annotation-free open-vocabulary 3D semantic segmentation without additional training processes. Since CLIP’s vision features are well aligned with text features, we can take the text embeddings of each class name’s prompts as classification weights to perform semantic segmentation. CLIP-FO3D performs remarkably on standard close-set ScanNet and S3DIS segmentation benchmark. In addition, to examine CLIP-FO3D’s open-vocabulary capability inherited from CLIP, we extend the label set of the standard ScanNet dataset with NYU labels . We demonstrate that CLIP-FO3D has remarkable segmentation results on NYU-40 classes and other long-tailed categories beyond the NYU label set.

Besides, given that collecting 3D point clouds and annotating are laborious, CLIP-FO3D can also be regarded as an unsupervised cross-modal pre-training framework to benefit data efficiency. CLIP-FO3D achieves great performance in traditional benchmarks where limited annotations are provided, such as zero-shot and data-efficient learning.

Most importantly, CLIP-FO3D encodes rich open-world knowledge inherited from CLIP. It understands not only object concepts, but also text queries with open world semantics, broadening the application of 3D scene understanding.

Our contributions can be summarized as follows:

We propose directly transferring CLIP’s feature space to 3D scene understanding models without any supervision, which preserves CLIP’s open-world properties to the maximum.

We present a method to extract pixel-level CLIP features, and a feature distillation method to align 3D point representations with CLIP’s feature space.

Our model achieves promising annotation-free 3D semantic segmentation performance on large vocabularies, and shows remarkable open-world properties.

As an unsupervised pre-training method, our model outperforms previous state-of-the-art methods in zero-shot and data-efficient learning.

Related Work

Open-vocabulary dense prediction aims to recognize and localize objects with open-set semantics. Pioneering works in 2D vision mainly utilize large-scale image-caption annotations as weak supervision source to enlarge the vocabulary set. Other works distill vision-language foundation models’ (e.g., CLIP) knowledge to transfer their open-vocabulary capability. However, they all require some form of human annotations, such as box/mask proposals and pixel semantic annotations. Open-world knowledge encoded in CLIP is forgotten when fine-tuning with human annotations. In contrast, MaskCLIP proposes directly utilizing CLIP for dense prediction tasks without training. However, it struggles to handle 3D scene’s RGB views with more complex contents, as shown in Section 4. We present a new method to extract free pixel-level CLIP features from 3D scene’s RGB views by only modifying CLIP’s inputs and forwarding process without fine-tuning.

2 Zero-shot 3D Visual Recognition

Zero-shot learning is relatively under-explored in 3D. Most works focus on recognition tasks on object-level data . Recent works adopt CLIP to perform zero-shot 3D classification using rendered views. For 3D scene understanding, Michele et al. and Chen et al. study zero-shot semantic segmentation with generative model and word embedding prototypes. Lu et al. study zero-shot 3D object detection with pseudo-labels generated with a 2D classifier. These methods require point cloud labels on seen categories for supervised training, which still results in a model with close-set vocabularies. More recent works apply CLIP to 3D scenes for open-vocabulary tasks. However, they require human annotations for training, such as 2D pixel labels and 3D point labels . PLA utilizes an image-captioning model to generate captions for 3D content, bringing superior results on open-vocabulary segmentation by aligning 3D representations with text embeddings. Unlike existing works, we propose directly transferring CLIP’s knowledge to 3D scene understanding models with no annotations.

3 3D Representation Learning

Inspired by the success of self-supervised representation learning in 2D vision, 3D pre-training methods achieve better fine-tuning performance and efficiency in various down-stream tasks by leveraging contrastive learning , masked auto-encoder , or both .

Beyond self-supervised pre-training, cross-modal learning methods propose to distill knowledge in images/text and pre-trained models to 3D representations . For scene understanding tasks, recent works enrich 3D point representations by utilizing 2D/text annotations , pixel-point alignments , neural rendering and pre-trained models . Our method can be regarded as an unsupervised cross-modal 3D representation learning method. We demonstrate that distilling CLIP’s richly-structured vision knowledge to 3D models can benefit 3D scene understanding when limited annotations are available, outperforming previous self-supervised and cross-modal pre-training methods.

Method

This section describes the process of transferring CLIP’s feature space to 3D scene understanding models. We first extract pixel-level CLIP features from the 3D scene’s RGB views by modifying CLIP’s inputs and forwarding process, introduced in Section 3.1. We then get target point features from pixel features and train the 3D model by feature distillation, introduced in Section 3.2.

Although CLIP only aligns image-level global features with text embeddings, it should inherently encode local and dense semantics. As demonstrated in MaskCLIP , CLIP must divide image-level semantics into local segments, and properly align each segment’s semantics with independent concepts in the text. MaskCLIP proposes to discard the global pooling layer and extract dense features from the final feature map with reformulated 1×11\times 1 convolutions.

However, adapting MaskCLIP to 3D scene contents brings poor dense prediction results, as shown in Section 4. On the one hand, CLIP’s feature map has a much lower resolution than the input images (downsampled by 16216^{2} in ViT/16 and 32232^{2} in ResNet-50 ). Although MaskCLIP is reasonably capable of recognizing salient semantics in the images, it struggles to segment the numerous objects in 3D scenes at different scales. On the other hand, the pixel features in CLIP’s final feature map contain much global semantics regarding the entire image. They may contain multiple objects’ semantics since each pixel’s feature aggregates information from all other pixels in forwarding. However, ideally, we hope to extract object-level features of different objects in a 3D scene. We propose increasing the resolutions and extracting local features from CLIP to address the aforementioned problems, described as follows.

Multi-scale region extraction. Firstly, the input view is cropped at multi-scales to adapt to the recognition of objects of various sizes in 3D scenes, as shown in Figure 2 (a). For each cropped sample, we divide into many local regions. This effectively improves the feature resolution. Specifically, we divide the image sample into super-pixels with SLIC as in . The super-pixel roughly covers an object or object’s part, resulting in locally visually similar regions. After processing the input image, we extract an embedding vector for each super-pixel with modified CLIP’s ViT encoder, which is introduced below.

Extracting local features from CLIP. We hope the super-pixel feature aggregates information from local patches rather than from the entire image like the global classification token in ViT’s encoder. This process is shown in Figure 2 (b). Given a cropped image sample IcI_{\text{c}}, we segment it into NN super-pixels: Ic=S1∪S2∪⋯∪SNI_{\text{c}}=\mathcal{S}_{1}\cup\mathcal{S}_{2}\cup\cdots\cup\mathcal{S}_{N}, where Si∩Sj=∅, ∀ i≠j\mathcal{S}_{i}\cap\mathcal{S}_{j}=\varnothing,~{}\forall~{}i\neq j. We then upsample the exact cropped image to CLIP’s input size and obtain Ic′I_{\text{c}}^{{}^{\prime}}. In ViT, Ic′I_{\text{c}}^{{}^{\prime}} is first reshaped into MM flattened patches: Ic′=P1∪P2∪⋯∪PMI_{\text{c}}^{{}^{\prime}}=\mathcal{P}_{1}\cup\mathcal{P}_{2}\cup\cdots\cup\mathcal{P}_{M}, where M=142M=14^{2} in ViT-B/16. We then assign each patch in Ic′I_{\text{c}}^{{}^{\prime}} to a specific super-pixel in IcI_{\text{c}} based on their spatial locations by interpolation, since the number of super-pixels NN is always fewer than the number of patches MM. We denote the patch Pi\mathcal{P}_{i} being assigned to the super-pixel Sj\mathcal{S}_{j} as Pi∼Sj\mathcal{P}_{i}\sim\mathcal{S}_{j}.

To represent each super-pixel’s feature, we add NN local classification tokens beyond the original global classification token in CLIP’s forwarding process. The local tokens share some similarities with the group tokens in , but are not learnable and have different updating mechanism. The local tokens are initialized the same as the global one and are updated by the same self-attention mechanism and pre-trained weights in ViT. The only difference in updating the local tokens during inference is how attention scores are computed. Recall that in ViT’s forwarding process, the attention score of the classification token is calculated as:

where CC is a constant scaling factor and Emb(⋅)\text{Emb}(\cdot) denotes the linear layers encoding the query, key, and value embeddings. xgx^{\text{g}} is the global classification token and xix_{i} represents the input feature of patch Pi\mathcal{P}_{i}.

In our method, the attention score of each local classification token is computed from local patches as:

Note that the original tokens in ViT are not affected during inference. By forwarding CLIP with additional tokens, we obtain a local feature for each super-pixel that is aligned with CLIP’s feature space. We then apply the same local token feature to all pixels within a super-pixel to preserve object-level semantics. In this way, we obtain a feature map for each cropped image of the same size as CLIP’s input.

Multi-scale feature fusion. After extracting local feature maps for all the cropped samples, we resize them back to the cropping sizes and stitch them back to the original input image, as shown in Figure 2 (c). Since one pixel may belong to different super-pixels from different cropped samples, we average all features as the final pixel feature. We extract pixel features for each 3D scene’s RGB view. While the whole process is slow and difficult to apply to real-time inference, we only do the above process once for each view and use the pixel-level features as offline training targets.

2 Feature Distillation with 2D Teacher

After extracting offline target point features for each scene, we train the 3D scene understanding model to learn from these targets by feature distillation. Denoting the learned point features for a scene as Flearn={f3D,i}i=1Np\mathcal{F}_{\text{learn}}=\{f_{\text{3D},i}\}_{i=1}^{N_{p}}, the loss function of feature distillation is:

The distance D(⋅,⋅)\mathcal{D}(\cdot,\cdot) is the negative cosine similarity:

where ∥⋅∥2\|\cdot\|_{2} is L2-norm\mathcal{L}_{2}\text{-norm}. We use cosine distance because CLIP-driven classification relies on the cosine distance between vision and text embeddings.

The whole training process only requires the 3D dataset and the pre-trained CLIP vision encoder, without any form of supervision (e.g., 2D/3D annotations or vision-language grounding annotations). Since the learned point features are consistent with CLIP’s feature space, the 3D model can perform open-vocabulary semantic segmentation and open-world reasoning once the feature distillation is finished.

Experiments

Dataset. We train CLIP-FO3D on ScanNet’s training set . We use the RGB-D images and the 3D scene meshes for training, and no labels are used. Specifically, we sub-sample RGB-D images from the raw ScanNet videos every ten frames for each scene. We evaluate our method on ScanNet and S3DIS datasets. ScanNet’s validation set and S3DIS’s “Area 5” are used for all the evaluation experiments. Notice that raw categories’ names and the mapping to NYU40 label set are provided in the ScanNet dataset, which we use to enlarge the vocabulary.

Implementation Details. We use CLIP’s ViT-B/16 encoder as the 2D backbone and use the corresponding text encoder for generating text embeddings. For all the experiments, we adopt MinkowskiNet16 as 3D scene understanding backbone. We remove the final classifier and change the output feature dimension to 512 to match CLIP’s feature dimension. We train CLIP-FO3D using SGD optimizer with a learning rate of 0.8 and a batch size of 4. The model is trained for 80K steps, and the learning rate is decreased by 0.99 for every 1,000 steps. The fine-tuning experiments on CLIP-FO3D are trained with a batch size of 4 for 40K steps. We set the initial learning rate of the pre-trained CLIP-FO3D networks to 0.001, with polynomial decay with power 0.9, because we find that a small learning rate leads to better results when fine-tuning CLIP’s feature space. The initial learning rate of the classifier in data-efficient learning is set to 0.1 following the original setting in . For zero-shot learning and data-efficient learning experiments, we follow the original benchmarks in and . More details can be found in Appendix B.

2 Annotation-free Semantic Segmentation

Semantic segmentation on standard benchmarks. We first present the results of annotation-free semantic segmentation on standard ScanNet and S3DIS benchmarks, where no supervision is provided during training. This is a rather challenging setting. After training CLIP-FO3D with feature distillation, we can take the text embeddings of each class name’s prompts as classification weights to perform semantic segmentation. We remove the “other furniture” category in ScanNet and the “clutter” category in S3DIS because they do not have specific semantics that can be classified with any text embeddings.

The results are shown in Table 1. The first two rows show the results of using multi-view pixel features to infer each point’s semantics by feature projection (introduced in Section 3.2). MaskCLIP-3D is a baseline method that we use MaskCLIP to compute pixel features for RGB views as in . Target Feature is our proposed method to extract free dense CLIP features from RGB views introduced in Section 3.1. Notice that one should forward the vision encoder for all RGB views and then align 3D points with pixels with camera pose, intrinsics and transformation matrix. This inference process by feature projection is very slow and not applicable for practical use. In contrast, CLIP-FO3D is convenient for processing new 3D data. However, the results can still be compared and give some inspiration.

It is observed that our Target Features extract more reasonable semantics than MaskCLIP-3D, indicating that the proposed method is indispensable in extracting dense CLIP features for 3D contents. Moreover, our CLIP-FO3D achieves remarkable segmentation results on ScanNet and S3DIS in this challenging setting (notice that four categories in S3DIS never appear during the training on ScanNet). CLIP-FO3D even outperforms its target features in ScanNet, indicating that some misleading target features can be corrected during feature distillation.

Semantic segmentation with open vocabularies. The standard ScanNet benchmark only contains a small vocabulary of 20 classes. To examine the open-vocabulary capability of CLIP-FO3D inherited from CLIP, we first extend the original vocabulary size with the NYU-40 label set. We remove the NYU-40 labels that do not have specific semantics (e.g., “other structure”, “other furniture”, “other prop”) and evenly divide all the rest categories into Head, Common and Tail based on the sample numbers of each category. We show the semantic segmentation results on the three category sets in Table 2. Our Target Features outperform the baseline MaskCLIP-3D by a large margin. Since categories in Common and Tail set usually contain objects with smaller sizes (e.g., bag, box, pillow, book), MaskCLIP-3D struggles to extract their semantics due to the limited feature resolution. While our method even extracts higher quality features for Tail categories than Head, thanks to the multi-scale inputs and local super-pixel features. CLIP-FO3D also achieves great results in all categories, while it does not perform as well as our Target Features for Common and Tail categories, due to the limited number of examples in the training set. However, CLIP-FO3D can be easily used for indoor open-vocabulary scene understanding applications.

We further show the semantic segmentation results of more categories beyond the NYU-40 label set: the ten most frequent raw categories provided with the ScanNet dataset that do not belong to NYU-40 categories are used in Table 3. Our Target Features and CLIP-FO3D perform well on these long-tailed categories, showing great open-vocabulary scene understanding properties.

Qualitative Results. We show the visualization of annotation-free semantic segmentation results on standard ScanNet benchmark and long-tailed categories. As shown in Figure 3, decent segmentation masks are obtained compared to the ground truth. Our method differs from ground-truth results for some objects with ambiguous semantics, such as table and desk, cabinet and its door, chair and sofa. In Figure 4, our method successfully segments several long-tailed categories that are not annotated in traditional benchmarks, demonstrating our model’s excellent open vocabulary capability.

3 Zero-shot Semantic Segmentation

Zero-shot semantic segmentation methods train the 3D scene understanding model only with the labels on a subset of classes (seen classes) and evaluate both seen classes and unseen classes. All the existing methods in 3D scene understanding use transductive settings, in which the unlabeled points are accessible during training. CLIP-FO3D can be applied to zero-shot semantic segmentation with minor effort. Specifically, CLIP-FO3D can be used to generate pseudo-labels for the unlabeled points.

We compare our method with state-of-the-art zero-shot semantic segmentation methods on ScanNet and S3DIS following the benchmarks in . We also build a new benchmark following , where six classes are chosen as unseen classes in S3DIS. We use the metric of mean intersection over union (mIoU) for both seen classes (S\mathcal{S}) and unseen classes (U\mathcal{U}), and use the harmonic mean IoU (hIoU) to demonstrate the overall performance of zero-shot learning as in .

As shown in Table 4, our method outperforms the previous state-of-the-art method by 10.1%, 18.3%, 33.8%, 35.0% and 34.7% on hIoU metric when there are 2, 4, 6, 8 and 10 unseen classes during training. Our CLIP-FO3D also outperforms the method of pseudo-labeling with MaskCLIP-3D by large margins. It is observed that our method is more effective than the baselines when there are more unseen classes, demonstrating superior zero-shot capability. The results on S3DIS with two settings in Table 5 reflect a similar phenomenon, although S3DIS’s data is not accessible during the training of CLIP-FO3D.

4 Data-efficient 3D Scene Understanding

As the collection and annotation of 3D point cloud data are very laborious, data-efficient learning methods have been proposed to train a better 3D model when training data or labels are scarce. Existing works have explored self-supervised or cross-modal pre-training methods for data-efficient fine-tuning. CLIP-FO3D can also be considered an unsupervised cross-modal pre-training method to benefit data efficiency. Unlike existing works, our method leverages the richly-structured feature space inherited from CLIP to improve data-efficient fine-tuning results.

We follow the official data-efficient learning benchmarks : Limited Reconstructions (only a few labeled scenes are used for training) and Limited Annotations (only a few points are labeled in each scene). “Scratch” denotes the training from scratch baseline. The results of limited scene reconstructions are shown in Table 6, when only 1%, 5%, 10%, and 20% scenes are used during training. Our method achieves mIoU improvements of 9.3%, 4.8%, 4.7%, and 3.9% over training from scratch, and achieves superior results compared to the previous state-of-the-art supervised pre-training method. The results of limited point annotations are shown in Table 7, when 20, 50, 100, and 200 annotated points are randomly sampled for each scene. Our method achieves mIoU improvements of 11.3%, 6.0%, 5.4%, and 4.1% over training from scratch, and outperforms previous state-of-the-art supervised pre-training methods.

Discussion

Ablation study. Table 8 shows the ablation study on the proposed methods when extracting pixel features from CLIP. “Multi-scale” represents increasing the feature resolution with multi-scale inputs, and “Local Feature” represents extracting super-pixel features using local classification tokens. We show the results of annotation-free and zero-shot semantic segmentation (with four unseen classes) on ScanNet. It is observed that both techniques are crucial for transferring CLIP’s feature space. While each single method outperforms the MaskCLIP-3D baseline, increasing the feature resolution with multi-scale inputs brings more significant improvements than extracting local features.

Open-world 3D scene understanding. Since CLIP is trained with massive image-text pairs mined from the Internet, it inherently encodes rich real-world knowledge which can guide open-world applications such as visual navigation and embodied AI . Since we directly transfer CLIP’s feature space to 3D models, CLIP-FO3D has the potential to accomplish open-world 3D scene understanding.

Inspired by the qualitative results in and , we use text embeddings to query open-world semantics for CLIP-FO3D’s scene representations. Given a text description, we extract its feature with CLIP’s text encoder, calculate its similarity with point features, and then threshold to produce a 3D mask. The visualization results are shown in Figure 5. Our model successfully finds the location most relevant to the text descriptions’ actual meaning. For example, given “write”, CLIP-FO3D finds the whiteboard and desk on which one can write. Our model also understands other affordance, activity, color, and function words. These results demonstrate that our 3D scene representations encode rich and well-structured knowledge about the real world.

Trained with no human annotations, CLIP-FO3D successfully preserves CLIP’s open-world properties, while the model trained with human annotations can only recognize object concepts. Our method broadens the applications for 3D scene understanding. For example, in conjunction with the language foundation models (e.g., BERT , GPT-3 ), CLIP-FO3D may enrich the functionality of robots by connecting 3D vision and language modalities.

Conclusion

This paper proposes directly transferring CLIP’s feature space to 3D scene understanding models without any supervision. We first extract pixel-level features from CLIP for 3D scene’s RGB views. Then we adopt the feature projection scheme to get the target point features and train a 3D model with feature distillation. CLIP-FO3D achieves remarkable annotation-free semantic segmentation results on standard benchmarks as well as open-vocabulary concepts. Our model also outperforms previous state-of-the-art methods in zero-shot and data-efficient learning tasks. Most importantly, our model successfully inherits CLIP’s open-world properties, allowing 3D scene understanding models to encode open-world knowledge beyond object concepts.

References

Appendix A Advantages of Annotation-free Training

Our contribution is directly transferring CLIP feature space to 3D representations without any form of annotation. We demonstrate that this annotation-free training preserves CLIP’s rich information to the maximum. In contrast, fine-tuning CLIP with human annotations restricts the feature space. And the open-world knowledge encoded in CLIP is forgotten during fine-tuning as demonstrated in .

To further demonstrate the advantages of annotation-free training, we compare our unsupervised pixel-level feature extraction with a supervised 2D segmentation method, LSeg . Lseg aligns pixel embeddings to the CLIP text embedding of the corresponding semantic class. It is trained on 7 segmentation datasets, and most importantly, uses the original label sets provided by these datasets without relabeling (details can be found in Section 5.2 of Lseg’s paper ). Large amount of human annotations are required, and the supervision covers hundreds of categories in LSeg.

Although our annotation-free method does not perform well compared with LSeg on some common categories from existing benchmarks, we show that our method has outstanding advantages of recognizing long-tailed categories and open-world knowledge (e.g., color, affordance) directly inherited from CLIP. Although Lseg is trained on large-scale annotated object concepts, its feature space is restricted to these concepts and can not be generalized to other open-world knowledge.

Visualization results of capturing long-tailed concepts and open-world knowledge are shown in Figure 6. We use the ViT-B/16-based Lseg model that is pre-trained on 7 datasets provided by its official code. The similarity of the pixels and text queries are used for visualization (normalized to $$). The second picture in Figure 6 comes from the paper of a concurrent work , which proposes mapping language, vision and audio inputs to the same feature space. We show the results of long-tailed categories (e.g., yogurt, milk, crisps) and some proper names (e.g., Gryffindor, name of a house in the Harry Potter films; Minion, classic cartoon character). These categories are not included in LSeg’s pre-training datasets, so it produces non-discriminative results, while our method successfully highlights the target regions. Besides, we show the results of open-world text queries, such as color and affordance words. It is observed that our model indeed encodes color information into pixel features (e.g., black, yellow) and also understands objects’ affordance (e.g., open, rest and artistic). Since these open-world text queries are irrelevant to the object concepts in pre-training datasets, LSeg does not encode this knowledge into its feature space.

A concurrent work utilizes a pre-trained LSeg model to perform zero-shot 3D semantic segmentation. Given that LSeg is trained on a large amount of human annotations regarding hundreds of raw categories, and pixel-to-point alignments can also be obtained from camera poses and intrinsics in , we tend to consider their method not really zero-shot learning for 2D annotations. In contrast, CLIP-FO3D is entirely annotation-free.

Appendix B Implementation Details

More details of the training process. When extracting free pixel-level CLIP features, we introduce multi-scale region extraction and super-pixel partition in Section 3.1. For multi-scale region extraction, we use the cropping size of 1,121,\frac{1}{2}, and 14\frac{1}{4} of the input view’s width/height. And the cropping window slides with the stride of 12\frac{1}{2} cropping size. For super-pixel partition, we divide each image sample into 5050 super-pixels with SLIC . We use a smaller number of super-pixel compared with SLidR since we already improved the resolution with multi-scale cropping. Through super-pixel partition, we extract locally visually similar regions which represent objects or object parts. By extracting a local feature for each super-pixel, we intend to preserve object-level CLIP semantics.

More details of the ablation study. Here, we introduce the implementation details of the ablation study in Section 5. In the “w/o Multi-scale” setting, we discard the multi-scale region extraction and the corresponding multi-scale feature fusion. We only extract local super-pixel features for the whole input view. In the “w/o Local Features” setting, multi-scale cropping and feature fusion are kept the same as in the main method, and the CLIP local feature extraction process is replaced with MaskCLIP .

Prompt engineering. To classify points’ semantics, we take the text embeddings of each class name’s prompts as classification weights. We use the 8080 hand-craft prompts as in MaskCLIP. In practice, we extract CLIP embeddings for all prompts and average them to obtain a single text embedding for each class.

Appendix C Dataset Partitioning

Semantic segmentation with open vocabularies. To examine the open-vocabulary properties of CLIP-FO3D, we extend the original vocabulary size in ScanNet benchmark with the NYU-40 label set. We remove the NYU-40 labels that do not have specific semantics (e.g., “other structure”, “other furniture”, “other prop”) and evenly divide all the rest categories into Head, Common and Tail. Head classes contain wall, floor, cabinet, bed, chair, bathtub, table, door, toilet, bookshelf, curtain, and ceiling. Common classes contain sofa, counter, desk, dresser, refrigerator, shelves, shower curtain, night stand, window, picture, sink and floor mat. Tail classes contain blinds, mirror, clothes, pillow, book, box, whiteboard, lamp, towel, bag, person, and television.

Zero-shot learning We follow the original benchmarks in for zero-shot semantic segmentation. In ScanNet, we conduct experiments with a different number of unseen classes, including the 2-sofa/desk, 4-bookshelf/toilet, 6-bathtub/bed, 8-curtain/window, 10-door/counter. In S3DIS , beam, column, window and sofa are unseen classes. We also conduct experiments on a new benchmark with six unseen classes in S3DIS, where board, door, floor, sofa, table, window are chosen as unseen classes. This benchmark is first used in PLA , which aligns coarse-to-fine 3D scene representations with paired text embeddings and achieves promising zero-shot learning results. Different with this work, our method does not rely on pre-trained image-captioning model and text data.

Appendix D More Qualitative Results

We show additional qualitative results of annotation-free semantic segmentation in Figure 7 and open-world 3D scene understanding in Figure 8.