Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang, Xianzheng Ma, Jiaming Han, Kexin Chen, Peng Gao, Xianzhi Li, Hongsheng Li, Pheng-Ann Heng

Introduction

In these years, 3D vision has gained significant attention and development, driven by the rising popularity of autonomous driving , navigation , 3D scene understanding , and robotics . To extend its application scenarios, numerous efforts have been made to incorporate 3D point clouds with data from other modalities, allowing for improved 3D understanding , text-to-3D generation , and 3D question answering .

For 3D geometry understanding, previous works either leverage 2D-language embeddings to guide 3D open-world recognition , or harness visual and textual semantics to assist 3D representation learning . However, their perception capabilities are mostly constrained by limited modalities provided in the training phase. Inspired by 2D generative models , a collection of methods has achieved text-to-3D synthesis with high quality and efficiency. Despite this, they lack the ability to generate 3D shapes conditioned on multi-modal input, i.e., any-to-3D generation. Another series of works connects descriptive natural language with 3D data, which is applied to 3D captioning , question answering , and visual grounding . Yet, they fail to utilize the pre-trained linguistic knowledge within large language models (LLMs) to better capture 3D geometrics.

Therefore, how to develop a unified 3D framework aligning with multi-modality for general 3D learning still remains an open question. Very recently, ImageBind was proposed to learn a shared representation space across six different modalities, i.e., images, text, audio, depth, thermal, and IMU data. Motivated by this, we ask the following question: can we construct a joint embedding space between 3D and multi-modality for unified 3D understanding, generation, and insturction following?

To this end, we introduce Point-Bind, a 3D multi-modality framework that aligns point clouds with multiple modalities for general 3D analysis, as shown in Figure 1. Specifically, we collect 3D-image-text-audio pairs as the training data, and construct a joint embedding space guided by ImageBind. We adopt a contrastive loss between the extracted features from a trainable 3D encoder, e.g., I2P-MAE , and the frozen multi-modal encoders of ImageBind. Such a simple strategy can efficiently integrate different modalities into a unified representation space, and allows for various 3D-centric multi-modal tasks in Figure 2.

The main contributions of Point-Bind are as follows:

Aligning 3D with ImageBind. Within a joint embedding space, Point-Bind firstly aligns 3D point clouds with multi-modalities guided by ImageBind, including 2D images, video, language, audio, etc.

Any-to-3D Generation. Based on existing text-to-3D generative models, Point-Bind enables 3D shape synthesis conditioned on any modalities, i.e., text/image/audio/point-to-mesh generation.

3D Embedding-space Arithmetic. We observe that 3D features from Point-Bind can be added with other modalities to incorporate their semantics, achieving composed cross-modal retrieval.

3D Zero-shot Understanding. Point-Bind attains state-of-the-art performance for 3D zero-shot classification. Also, our approach supports audio-referred 3D open-world understanding, besides text reference.

Furthermore, on top of our joint embedding space, we propose to incorporate Point-Bind with LLaMA to develop the first 3D large language models (LLMs), termed as Point-LLM. As shown in Figure 2, our Point-LLM can respond to language instructions with 3D point cloud conditions, and effectively capture spatial geometry characteristics. Referring to ImageBind-LLM , we utilize a bind network along with a visual cache model to bridge Point-Bind with LLaMA, and adopt zero-initialized gating mechanisms for parameter-efficient fine-tuning. With superior data efficiency, the entire training phase of Point-LLM requires no 3D instruction dataset, and only utilizes public vision-language data for vision-language tuning. In this way, we enable LLMs to understand and conduct cross-modal reasoning for 3D and multi-modal data, achieving superior 3D question-answering capacity in both English and Chinese.

The main contributions of Point-LLM are as follows:

Point-LLM for 3D Question Answering. Using Point-Bind, we introduce Point-LLM, the first 3D LLM that responds to instructions with 3D point cloud conditions, supporting both English and Chinese.

Data- and Parameter-efficiency. We only utilize public vision-language data for tuning without any 3D instruction data, and adopt parameter-efficient fine-tuning techniques, saving extensive resources.

3D and Multi-modal Reasoning. Via the joint embedding space, Point-LLM can generate descriptive responses by reasoning a combination of 3D and multi-modal input, e.g., a point cloud with an image/audio.

Related Work

Compared to single-modal approaches, multi-modal learning aims to learn from multiple modalities simultaneously, achieving more robust and diverse representation learning. Numerous studies have proved its efficacy, involving 2D images, videos, texts, and audio , and enhance the cross-modal performance for downstream tasks , and video-text-audio integration for text generation . The representative vision-language pre-training, CLIP , effectively bridges the gap between 2D images and texts, which encourages further exploration of cross-modality learning. Recently, ImageBind successfully aligns six modalities in a joint embedding space, unleashing the power for emergent zero-shot cross-modal capabilities. However, ImageBind fails to investigate its efficacy on 3D point clouds. In the 3D domain, most existing cross-modal works introduce vision-language alignment into 3D point clouds, and mainly focus on open-world recognition tasks, which ignore the potential of multi-modal semantics for wider 3D applications. In this paper, our Point-Bind develops a general 3D multi-modality model that aligns 3D point clouds with six other modalities guided by ImageBind, allowing for more diverse 3D cross-modal understanding.

Large Models in 3D.

Large-scale pre-trained models have achieved remarkable downstream performance in language and 2D image processing. Inspired by this, many efforts have introduced 2D and language large models, to assist in 3D learning. The prior PointCLIP series project 3D point clouds into depth maps, and utilize CLIP for zero-shot recognition. Image2Point instead converts 2D pre-trained models into 3D space as a good network initialization. By contrastive learning, ULIP series and other works pre-train 3D networks guided by the vision-language embedding space of CLIP. Another branch of work employs CLIP to guide the text-conditioned generation of 3D objects or stylized meshes by encoding descriptive textual input. Some works also adopt GPT-3 to enhance the language-based understanding of 3D spatial geometry, such as PointCLIP V2 and ViewRefer . Different from them, we utilize ImageBind to construct a joint embedding space between 3D point clouds and multiple modalities. The derived Point-Bind can well leverage the multi-modal semantics for general 3D cross-modal understanding, generation, and question answering.

Pre-training in 3D.

In recent years, significant progress has been made in supervised learning for 3D vision tasks . However, these approaches lack satisfactory generalization capabilities for out-of-domain data. To address this, self-supervised learning has emerged as a promising solution to enhance 3D transfer learning . Most self-supervised pre-training methods employ an encoder-decoder framework to encode point clouds into latent representations and then reconstruct the original data form . Therein, Point-MAE and Point-M2AE introduce masked autoencoders into 3D point clouds pre-training, achieving competitive results on different 3D tasks. Alternatively, cross-modal pre-training approaches are also leveraged to enhance the 3D generalization ability . For example, ACT and I2P-MAE utilize pre-trained 2D transformers as teachers to guide 3D representation learning. Inspired by previous works, we adopt collected 3D-image-text-audio pairs for self-supervised pre-training, and regard ImageBind’s encoders as guidance for contrastive learning. In this way, the Point-Bind is pre-trained to obtain a joint embedding space between 3D and multi-modality, allowing for superior performance on different 3D downstream tasks.

Point-Bind

The overall pipeline of Point-Bind is shown in Figure 4. In Section 3.1, we first provide a preliminary of ImageBind . Then, in Section 3.2 and 3.3, we elaborate on the training data and multi-modal alignment for Point-Bind, respectively. Finally, in Section 3.4, we introduce several 3D-centric applications derived from Point-Bind.

ImageBind proposes an approach to combine multiple modalities together, which utilizes only image-paired data to learn a joint embedding space of six modalities, i.e., images, text, audio, depth, thermal, and IMU data. It does not need training dataset pairing all six modalities, but leverages the binding property of 2D images, i.e., aligning every single modality to image independently. Specifically, ImageBind feeds multi-modal input into corresponding encoders, and adopts for cross-modal contrastive learning. After training on large-scale image-paired data, ImageBind effectively aligns six modalities into a single shared representation space, enabling emergent cross-modal zero-shot capabilities. Based on existing vision-language models, ImageBind can also be utilized for several multi-modal tasks, such as text-to-audio/video retrieval, audio-to-image generation, and audio-referred object detection. Inspired by this, we propose to develop a 3D multi-modal framework that incorporates 3D point clouds with other modalities for general 3D understanding, generation, and instruction following.

2 Training Data

To align 3D with multi-modalities, we leverage the pre-trained joint embedding space of ImageBind and utilize contrastive loss to simultaneously align 3D point clouds with the other three modalities: image, text, and audio. To obtain the contrastive training data, we collect a cross-modal dataset of 3D-image-audio-text pairs. There are three steps for dataset collection as follows.

We adopt the data pairs of 3D, images, and text from ULIP , which includes 3D-image-text triplets built from ShapeNet , a common-used dataset containing abundant 3D CAD models. Each 3D point cloud is paired with a corresponding text describing the semantic information of its spatial shape, and a 2D counterpart generated by multi-view image rendering. The text description is constructed by a synset of category names and 64 pre-defined templates.

D-audio Pairs.

To provide more contrastive signals, we also collect the data pairs of 3D and audio from ESC-50 and ShapeNet datasets. Specifically, we first select the categories whose objects can make a sound in the real world from the 55 categories of ShapeNet, such as ‘airplane’, ‘clock’, ‘washing machine’, and ‘keyboard’. Then, we preserve only the categories that are also within ESC-50. By this standard, we obtain 9 categories of 3D-audio paired data with extensive audio clips.

D-image-audio-text Pairs Construction.

Finally, we match each 3D-audio pair with its corresponding 3D-image-text data, resulting in a unified 3D-image-audio-text dataset with extensive cross-modal pairs. During training, we simultaneously feed point clouds and their paired data of three modalities for contrastive learning.

3 Aligning 3D with Multi-modality

After collecting the 3D paired data, we conduct contrastive training to learn a joint embedding space aligning 3D and multi-modalities. Each data sample contains a point cloud PP, along with the paired 2D image II, text description TsT^{s}, and audio AA, where TsT^{s} represents a set of 64 pre-defined templates. For the point cloud, we adopt I2P-MAE as the learnable 3D encoder, denoted as Encoder3D(⋅)⁡\operatorname{Encoder_{3D}(\cdot)}, and append a projection network Proj(⋅)⁡\operatorname{Proj(\cdot)} of two linear layers, which transforms the encoded 3D feature into ImageBind’s multi-modal embedding space. We formulate it as

which represents the aggregated text embedding with more robustness. After that, we adopt contrastive loss between 3D and other modalities, which effectively enforces 3D embeddings to align with the joint representation space, formulated as

Note that some categories of the training data do not include the paired audio AA, since they inherently cannot make any sound, e.g., bottle, planter, and couch, for which we ignore their audio features and loss.

4 Multi-modal Applications

Benefiting from the joint embedding space of Point-Bind, we respectively introduce several emergent application scenarios concerning 3D and multi-modalities.

Inherited from 2D generative models, existing 3D generation methods can only achieve text-to-3D synthesis. In contrast, with the joint embedding space of Point-Bind, we can generate 3D shapes conditioned on any modalities, i.e., text/image/audio/point-to-mesh. In detail, we directly connect the multi-modal encoders of Point-Bind with the pre-trained decoders of current CLIP-based text-to-3D models, e.g., CLIP-Forge . Without further training, we can synthesize a 3D car mesh based on an input car horn.

D Embedding-space Arithmetic.

We observe that 3D features encoded by Point-Bind can be directly added with other modalities to incorporate their semantics, further achieving composed cross-modal retrieval. For instance, the combined embeddings of a 3D car and audio of sea waves can retrieve an image showing a car parking by a beach, while the composition of a 3D laptop and audio of keyboard typing can retrieve an image of someone who is working with a laptop.

D Zero-shot Understanding.

For traditional text-inferred 3D zero-shot classification, Point-Bind attains state-of-the-art performance guided by additional multi-modal supervision. Besides, Point-Bind can also achieve audio-referred 3D open-world understanding, i.e., recognizing 3D shapes of novel categories indicated by the corresponding audio data .

Point-LLM

In this section, we illustrate how to leverage Point-Bind to develop 3D large language models (LLMs), termed as Point-LLM, which fine-tunes LLaMA to achieve 3D question answering and multi-modal reasoning. The overall pipeline of Point-LLM is shown in Figure 4.

Our Point-LLM is developed on top of ImageBind-LLM , which conducts multi-modality instruction tuning by injecting the semantics of ImageBind into LLaMA. Our approach exhibits both data and parameter efficiency.

During training, only the public vision-language data is required for fine-tuning LLaMA to learn the 3D-conditioned response capacity. As Point-Bind has built a joint embedding space between 3D and multi-modalities, if any one of the modalities is trained to connect with LLaMA, the others would also be aligned at the same time. Considering this, we select the 2D image modality, since it has the most public data with paired language. By only aligning ImageBind’s image encoder with LLaMA, we avoid the expensive cost of collecting and annotating large-scale 3D instruction data, thus saving extensive resources.

Parameter-efficient Training.

Instead of tuning the entire LLMs , we only unfreeze partial parameters within LLaMA for efficient vision-language instruction tuning. Specifically, a learnable bind network is adopted to bridge the image encoder of ImageBind with the language space of LLaMA. Then, a zero-initialized gating mechanism is proposed to add the image features after the bind network to the words tokens within LLaMA. This mechanism can progressively inject visual instruction cues into LLaMA for stable training at early stages, inspired by LLaMA-Adapter . By such a parameter-efficient fine-tuning strategy, most parameters of LLaMA are kept frozen, and only the zero-initialized gating factors and bias-norm weights are learnable for fine-tuning. Please refer to ImageBind-LLM for further training details. After the vision-language training, the joint embedding space enables LLaMA to naturally align with other modalities, such as audio within ImageBind and 3D point clouds of Point-Bind. Therefore, our Point-LLM effectively provides LLaMA with 3D-instruction following capacity without any 3D instruction data, indicating superior data efficiency.

2 3D Question Answering

For an input language instruction and a 3D point cloud, we feed them into the fine-tuned LLaMA and our Point-Bind, respectively. Then, the encoded 3D feature is enhanced by a visual cache model proposed in ImageBind-LLM, before feeding into the bind network. The cache model is only adopted during inference, and constructed in a training-free manner .

As we adopt the image encoder of ImageBind for training, but switch to Point-Bind’s 3D encoder for inference, the cache model is designed to alleviate such modality discrepancy for better 3D geometry understanding. Referring to ImageBind-LLM, the cache model stores from three ImageBind-encoded image features from the training data, which are regarded as both keys and values for knowledge retrieval. We regard the input 3D feature as the query, and retrieve the top-kk similar visual keys from the cache model. Then, according to the cosine similarity, we aggregate the corresponding cached values (top-kk similar image features), and add the result to the original 3D feature via a residual connection. In this way, the enhanced 3D feature can adaptively incorporate similar visual semantics from the cache model. This boosts the representation quality of 3D shapes, and mitigates the semantic gap of 2D-3D encoders within Point-LLM. After this, the enhanced feature is fed into the bind network for feature transformation and LLaMA for response generation.

D and Multi-modal Reasoning.

In addition to point clouds, our Point-LLM can also conduct cross-modal reasoning and generate responses conditioned on multiple modalities. For an additional input image or audio, we utilize the image or audio encoder of ImageBind to extract the features, and directly add them with the 3D feature encoded by Point-Bind. By injecting such integrated features into LLaMA, Point-LLM can reason cross-modal semantics, and respond with the information of all input modalities. This demonstrates the promising significance of aligning multi-modality with 3D LLMs.

Experiments

In this section, we first present the implementation details of the multi-modal training for Point-Bind. Then, we illustrate the emergent multi-modal applications, i.e., Point-LLM for 3D instruction following, 3D cross-modal retrieval, 3D embedding-space arithmetic, any-to-3D generation, and 3D zero-shot understanding. Finally, we conduct an ablation study to verify the effectiveness of our designs.

We adopt pre-trained I2P-MAE as the 3D encoder of Point-Bind, and utilize the collected 3D-image-text-audio pairs for pre-training. We only update the 3D encoder with the newly added projection network, and freeze the encoders of other modalities in ImageBind . The projection network is simply composed of two linear layers with an intermediate LayerNorm . We train Point-Bind for 300 epochs with a batch size of 64, and adopt AdamW as the optimizer with a learning rate of 0.003.

2 Point-LLM for 3D Q&A

We refer to ImageBind-LLM to conduct parameter- and data-efficient fine-tuning to inject 3D instructions into the pre-trained LLaMA 7B model . In detail, the fine-tuning techniques include zero-initialized gating , LoRA , and bias-norm tuning . We utilize a collection of several datasets for vision-language training, and require no 3D instruction-following dataset due to the learned joint embedding space.

Analysis.

In Figure 3, we provide the question-answering examples of Point-LLM, which shows favorable 3D instruction-following and multi-modal reasoning capacity. As shown, for either English or Chinese instructions, Point-LLM can effectively incorporate the spatial geometry of input point clouds and generate detailed language responses. It obtains a comprehensive 3D understanding for both global and local characteristics, e.g., recognizing the pattern of the piano keyboard and the shape of the airplane’s wing and tail. Then, our Point-LLM can also respond with cross-modal understanding. For an input 3D model with a 2D image or audio, Point-LLM can enable LLaMA to take both two conditions into understanding and reasoning, which thus incorporates multi-modal semantics in the output language response. With superior data- and parameter-efficiency, the examples indicate the 3D multi-modal instruction-following capabilities of Point-LLM.

3 3D Cross-modal Retrieval

To evaluate the multi-modal alignment of Point-Bind, we experiment on several cross-modal retrieval tasks, i.e., 3D-to-3D, 2D-to-3D, 3D-to-2D, and text-to-3D retrieval.

We conduct 3D zero-shot retrieval on multi-modal ModelNet40 dataset, which contains 9,843 CAD models for training and 2,468 for testing of 40 categories. ModelNet40 provides data of three modalities for retrieval, i.e., image, point cloud, and mesh. We obtain the retrieved results by ranking the similarities between embeddings of point clouds and other modalities. Following previous works , we measure the networks via the mean Average Precision (mAP) score, a commonly used evaluation criterion for retrieval tasks.

Analysis.

In Table 1, we report the quantitive results for 3D zero-shot retrieval, where Point-Bind attains state-of-the-art performance on all benchmarks compared with prior works. In particular, for 2D-to-3D and text-to-3D retrieval, Point-Bind surpasses the second-top ULIP significantly by +14.29% and +13.99% improvements, respectively. This indicates the superior cross-modal understanding capacity of our approach.

4 Embedding-space Arithmetic with 3D

With the multi-modal alignment, we further explore the capability of embedding composition, i.e., the embedding-space arithmetic of 3D and other modalities, e.g., audio.

To obtain the multi-modal input for arithmetic, we utilize 3D objects from ShapeNet and TextANIMAR2023 , and audio clips from ESC-50 . We simply add the 3D and audio embeddings from Point-Bind and ImageBind, respectively, and then retrieve 2D images from ImageNet with 1,000 image categories.

Analysis.

In Figure 6, we show the results of 2D image retrieval with the composed embeddings between 3D and audio. As shown in the first row, with the combined embeddings of a 3D dog and sea-wave audio, we effectively retrieve 2D images of dogs by the sea. Similarly, with the combination of a 3D laptop and keyboard-typing audio, the obtained images show someone is working with a laptop, or a cat inadvertently presses on the keyboard. Likewise, the last row retrieves images of bears hunting by the water by using embeddings of a 3D bear and audio of flowing water. The examples demonstrate that the 3D features encoded by Point-Bind can be directly added with other modalities, and well incorporate their semantics, achieving favorable composed cross-modal retrieval capacity.

5 Any-to-3D Generation

Existing text-to-3D generation methods normally adopt CLIP’s text encoder to process the input language prompt. Considering this, we simply replace it with the multi-modalities encoders of Point-Bind and ImageBind without further training, which follows the original generative decoder for 3D shape synthesis. We adopt the decoder of CLIP-Forge by default.

Analysis.

In Figure 7, we show the examples of any-to-3D generation powered by Point-Bind. For text, audio, and point cloud prompts, our approach can all produce satisfactory 3D meshes. This demonstrates the well-aligned embedding space of 3D and multiple modalities.

6 3D Zero-shot Understanding

In this section, we test the open-word understanding ability of Point-Bind, i.e., recognizing novel classes, by 3D zero-shot classification on ModelNet40 dataset.

Following previous works, we utilize the text embeddings from CLIP’s or ImageBind ’s text encoder to construct the zero-shot classification head. Specifically, we apply a simple template of ‘a/an [CLASS]’ for the 40 categories of ModelNet40, and calculate the cosine similarity between 3D and all textual embeddings, selecting the most similar one as the final prediction.

Analysis.

We report the 3D zero-shot classification accuracy in Table 2, where our Point-Bind surpasses existing methods with state-of-the-art performance. This indicates the unified representation space of Point-Bind leads to strong emergent 3D open-world recognition.

7 Ablation Study

To investigate the effectiveness of our designs in Point-Bind, we conduct ablation studies on the projection network and 3D encoders in Table 3. We report the performance of zero-shot classification on ModelNet40 dataset. In the first two columns, we experiment with different projection schemes for embeddings after the 3D encoder. As shown, using two linear layers for embedding projection performs the best. In the last two columns, we utilize different 3D encoders in Point-Bind, i.e., Point-BERT , PointNeXt , and I2P-MAE . As reported, the self-supervised Point-BERT and I2P-MAE achieve much better performance, indicating the importance of 3D pre-training to boost the multi-modal alignment.

Conclusion

In this paper, we propose Point-Bind, a 3D multi-modality model that aligns 3D point clouds with multi-modalities, guided by ImageBind. By aligning 3D objects with their corresponding image-audio-text pairs, Point-Bind obtains a joint embedding space, and exhibits promising 3D multi-modal tasks, such as any-to-3D generation, 3D embedding arithmetic, and 3D open-world understanding. Upon that, we further introduce Point-LLM, the first 3D large language model (LLM) with instruction-following capability in both English and Chinese. Our future work will focus on aligning multi-modality with more diverse 3D data, such as indoor and outdoor scenes, which allows for wider application scenarios.

References