ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
Zekun Qi, Runpei Dong, Shaochen Zhang, Haoran Geng, Chunrui Han, Zheng Ge, Li Yi, Kaisheng Ma
Introduction
3D shape understanding, serving as a fundamental capability for molding intelligent systems in both digital and physical worlds, has witnessed tremendous progress in graphics, vision, augmented reality, and embodied robotics. However, to be effectively deployed by real-world agents, several critical criteria must be fulfilled: (i) Sufficient 3D geometry information needs to be captured for accurate spatial and structure processing . (ii) Models should be endowed with a foundational knowledge of the embodied interaction fashion with objects — often physically — for functional comprehension . (iii) A universal interface is required as a bridge between information encoding and decoding, which could help translate high-order instructions for agent reactions like dialogue response and embodied feedback .
Recent advancements in Large Language Models (LLMs) have demonstrated unprecedented success of foundational knowledge and unified reasoning capabilities across tasks . It makes it possible to utilize language as a universal interface that enables the comprehensive commonsense knowledge embedded in LLMs to enhance understanding of 3D shapes. This is particularly evident in physically-grounded tasks, where the wealth of commonsense knowledge simplifies the interpretation of an object’s functionality, mobility, and dynamics, etc. However, the aforementioned challenges remain when incorporating LLMs for 3D object understanding — especially embodied interaction that relies on precise geometry — currently under-explored.
The question is: What makes better 3D representations that bridge language models and interaction-oriented 3D object understanding? In this work, we introduce ShapeLLM that meets the requirements, which is established based on the following three designing policies:
3D Point Clouds as Inputs Some concurrent works recently propose to use point cloud-rendered images as multimodal LLMs’ inputs and demonstrate effectiveness. However, these works fail to achieve accurate 3D geometry understanding and often suffer from a well-known visual hallucination issue . Compared to 2D images, 3D point clouds provide a more accurate representation of the physical environment, encapsulating sparse yet highly precise geometric data . Moreover, 3D point clouds are crucial in facilitating embodied interactions necessitating accurate 3D structures like 6-DoF object pose estimation .
Selective Multi-View Distillation Interacting with objects typically necessitates an intricate 3D understanding that involves knowledge at various levels and granularities. For instance, a whole-part high-level semantic understanding is needed for interactions like opening a large cabinet, while detailed, high-resolution (i.e., low-level) semantics are crucial for smaller objects like manipulating a drawer handle . However, existing works mainly distill single-view high-resolution object features from 2D foundation models , providing a complementary understanding . The potential of multi-view images, which offer abundant multi-level features due to view variation and geometry consistency , is often neglected. ShapeLLM extends ReCon to ReCon++ as the 3D encoder by integrating multi-view distillation. To enable the model to selectively distill views that enhance optimization and generalization, inspired by Carion et al., ReCon++ is optimized through adaptive selective matching using the Hungarian algorithm .
3D Visual Instruction Tuning Instruction tuning has been proven effective in improving LLMs’ alignment capability . To realize various 3D understanding tasks with a universal language interface, ShapeLLM is trained through instruction-following tuning on constructed language-output data. However, similar to 2D visual instruction tuning , the data-dessert issue is even worse since no object-level VQA data is available, unlike 2D . To validate the efficacy of ShapeLLM, we first construct 45K instruction-following data using the advanced GPT-4V(ision) on the processed Objaverse dataset and 30K embodied part understanding data from GAPartNet for supervised fine-tuning. Following MM-Vet , we further develop a novel evaluation benchmark named 3D MM-Vet. This benchmark is designed to assess the core vision-language capabilities, including embodied interaction in a 3D context, thereby stimulating future research. The 3D MM-Vet benchmark comprises 59 diverse InternetURL & License. 3D objects and 232 human-written question-answer pairs.
Through extensive experimentation, we first demonstrate that our improved 3D encoder ReCon++ sets a new state-of-the-art representation transferring on both downstream fine-tuned and zero-shot 3D object recognition. Specifically, ReCon++ has obtained 95.25% and 95.0% fine-tuned accuracy on ScanObjectNN and ModelNet40, surpassing previous best records by +1.85% on the most challenging ScanObjectNN. Besides, ReCon++ achieved 53.7% and 65.4% zero-shot accuracy on Objaverse-LVIS and ScanObjectNN, which is +0.6% and +1.6% higher than previous best. By utilizing our ReCon++ as ShapeLLM’s 3D encoder, ShapeLLM successfully unifies various downstream tasks, including 3D captioning, 3D VQA, embodied task planning & decomposition, 3D embodied visual grounding, and 3D precise referring dialogue (See Fig. 1). On our newly constructed 3D MM-Vet benchmark, 42.7% and 49.3% Total accuracy have been achieved by ShapeLLM-7B and ShapeLLM-13B, surpassing previous best records that also uses 3D point clouds by +2.1% and +5.1%, respectively. This work initiates a first step towards leveraging LLMs for embodied object interaction, and we hope our ShapeLLM and proposed 3D MM-Vet benchmark could spur more related future research.
ShapeLLM
In this section, we first introduce the overall architecture of ShapeLLM. Then, we delve into two critical challenges faced in interactive 3D understanding: data dessert and representation of 3D point clouds. We present the detailed design of our method to tackle these challenges, respectively.
The main objective of this work is interactive 3D understanding by using the LLM as a universal interface. Drawing inspiration from recent work in visual understanding , the proposed ShapeLLM consists a pre-trained 3D encoder and an LLM for effective 3D representation learning and understanding, respectively. Specifically, we adopt LLaMA as our LLM, building upon the success of previous work . As for the 3D encoder, we propose a novel 3D model named ReCon++ based on the recent work ReCon with multiple improvements as the 3D understanding generally demands more information, such as accurate spatial and multi-view details, etc. To ensure compatibility with the LLM inputs, the representation of a 3D object obtained from ReCon++ undergoes a linear projection before being fed into the LLM. To further improve low-level geometry understanding, which benefits tasks like 6-DoF pose estimation, we append the absolute position encoding (APE) obtained by linear projection of 3D coordinates. Besides, we use prefix-tuning with learnable prompts to adaptively modulate the different semantics of APE and ReCon++ representations.
2 How to alleviate interactive 3D understanding Data Dessert?
Most published 3D data is typically presented as 3D object-caption pairs, lacking an interactive style. Although a few concurrent works have attempted to construct interactive 3D understanding datasets, the questions-and-answers (Q&As) are primarily based on annotated captions, often providing a limited perspective without sufficient details. Additionally, those works have generally been limited to semantic understanding without considering embodied interaction. To address these limitations, our work constructs question-and-answer pairs based on multi-view images of a 3D object using GPT-4V(ision) . For data diversity, we explicitly introduce six aspects as prompts, as illustrated Fig. 3. In the following, we provide the details about data collection and construction regarding general semantic understanding and embodied object understanding, respectively.
Data Objaverse-LVIS and GAPartNet are data sources. Objaverse-LVIS covers 1,156 LVIS categories, and we sample Top-10 “likes”“Likes” statistics can be found at Sketchfab. 3D objects per category and generate Q&A pairs per sample. After filtering out noisy Q&As, we obtain 45K instruction-following samples. We use 12 categories from GAPartNet by removing “Remote” to avoid too many tiny boxes, which leads to filtered 30K Q&A samples constructed from 8K parts of the 4K objects states covering 1.1K different objects.
General Semantic Understanding This aims to enhance the model’s generalization abilities in visual recognition, knowledge integration, spatial understanding, and other aspects. We prompt GPT4-V to generate Q&As in six different aspects based on images captured from four different views of a 3D subject, as illustrated in Fig. 3.
Embodied Object Understanding A comprehensive understanding of the spatial positions and semantics at the part level is crucial to facilitate effective object grasping and interaction in embodied scenarios. Fortunately, the GAPartNet provides rich part annotations, including semantics and poses, which are instrumental in constructing instruction-tuning data for embodied interactive parts of a subject. Specifically, given a 3D object, questions are formulated based on the semantics of its different parts, and answers are constructed in both the semantics and 3D positions. The positions are represented as 6-DoF 3D bounding boxes in a straightened Python multidimensional list format, denoted as [[x1, y1, z1], [x2, y2, z2], …, [x8, y8, z8]], to meet characteristics of the textual dialogues response in LLMs. The canonical space of the object determines the sequence of coordinates. Using bounding box coordinates leverages the inherent spatial relationship, allowing LLMs to readily learn these patterns and generate accurate output coordinates. This approach can offer specific position information for embodied manipulation.
3 ReCon++: Scaling Up 3D Representation Learning
Interaction with objects such as object grasping typically requires accurate perception of 3D shape information at multi-level and multi-granularity. This imposes heightened requirements on 3D representations, calling for a higher standard of a holistic understanding of 3D geometry.
However, existing 3D cross-modal representation learning methods mainly distill high-resolution object features from single-view 2D foundation models, resulting in a unilateral shape understanding. Besides, they generally employ multi-view images as a data augmentation strategy, imposing the learned representation to the average representation of all views. Thus, the accurate 3D shape information is missing. Recently, ReCon utilizes contrast guided by reconstruction to address the pattern disparities between local masked data modeling and global cross-modal alignment. This results in remarkable performance in various tasks, including transfer learning, zero-shot classification, and part segmentation. However, its potential is hindered by the scarcity of pretraining data .
To address the above limitations, this paper proposes ReCon++ with multiple improvements. First, multi-view image query tokens collaboratively comprehend the semantic information of 3D objects across different views, encompassing both RGB images and depth maps. Considering the disorderliness of pretraining data in terms of pose, we propose a cross-modal alignment method based on bipartite matching, which implicitly learns the pose estimation of 3D objects. Second, we scale up the parameters of ReCon and broaden the scale of the pretraining dataset for robust 3D representations.
Denote as the number of multi-view images, is the image feature from -th view, and represents the global query of -th view. Following Carion et al., we search for an optimal permutation of elements with the lowest cost:
where is a pair-wise matching cost between -th view image features and matched query with the permutation . In practice, we employ cosine similarity as the matching cost. In this fashion, the query of each view is learned to gather accurate 3D shape information from the 3D point clouds. Concatenating the features from the local 3D point cloud encoder and global 3D point cloud decoder together provides comprehensive information for 3D understanding of multimodal LLMs.
3D MM-Vet: 3D Multimodal Comprehension Evaluation Benchmark
A wide range of diverse visual-language capabilities is essential to develop a multimodal large language model tailored for embodied scenarios, particularly addressing task and action planning.
The model’s proficiency in processing point clouds enables it to perform general recognition tasks effortlessly, demonstrating a broad understanding of colored point clouds. This capability serves as the groundwork for more intricate tasks. Beyond 3D recognition, the LLM should exhibit competence in addressing tasks in real-world embodied scenarios. This entails unifying the aforementioned abilities to generate decomposed task actions step-by-step in an instruction-following fashion, addressing specific problems.
Hence, to formulate an evaluation system aligned with the aforementioned task description, we establish a multi-level evaluation task system encompassing four-level tasks: General Recognition, Knowledge and Language Generation, Spatial Awareness, and Embodied Interaction. This framework systematically and comprehensively assesses the model’s proficiency in information comprehension and language generation when processing interactive objects. The detailed descriptions of the tasks are listed as follows:
General Recognition: Following MM-Vet , we assess the fundamental comprehension abilities of LLMs involving both coarse- and fine-grained aspects. Coarse-grained recognition focuses on basic object attributes such as color, shape, action, etc. While fine-grained recognition delves into details like subparts and counting, etc.
Knowledge Capability & Language Generation: To examine the models’ capacity to understand and utilize knowledge, drawing inspiration from MMBench , we integrate its reasoning components. This includes knowledge spanning natural and social reasoning, physical properties, sequential prediction, math, etc., evaluating gauges whether multimodal LLMs possess the requisite expertise and capacity to solve intricate tasks. We utilize customized prompts to stimulate models and extract detailed responses to evaluate language generation.
Spatial Awareness: In 3D, spatial awareness holds heightened significance compared to 2D due to the provided geometry information. The point clouds contain location information crucial for discerning spatial relationships between different parts. In 2D, achieving the same information intensity level would necessitate multi-view images. Therefore, our evaluation includes questions probing the ability of LLMs to understand spatial relations.
Embodied Interaction: The utilization scope of multimodal LLMs extends into the field of embodied interaction, facilitated by the utilization of instruction-following data. Our evaluation system tests their capacity by formally requesting LLMs to provide execution steps toward an instruction. This approach aims to establish connections for handling Embodied Interaction tasks .
To prevent any overlap with training data, our collection of 3D models is sourced exclusively from Turbosquid , a platform not included in the acquisition lists of Objaverse and ShapeNet . We meticulously curated a dataset of 59 3D models, generating 232 Q&As for evaluation purposes. In our pursuit of a precise assessment of single-task capabilities, each question is designed to test only one specific capacity outlined earlier. Every question is paired with a corresponding answer tailored to the particular 3D model, serving as the ground truth. More details and analysis can be found in Appendix B.
Experiments
Fine-tuned 3D Object Recognition In Tab. 1, we first evaluate the representation transfer learning capabilities of self-supervised ReCon++ by fine-tuning on ScanObjectNN and ModelNet , which are currently the two most challenging 3D object datasets. ScanObjectNN is a collection of 15K 3D object point clouds from the real-world scene dataset ScanNet , which involves 15 categories. ModelNet is one of the most classical 3D object datasets collected from clean 3D CAD models, which includes 12K meshed 3D CAD models covering 40 categories. Following PointGPT , we adopt the intermediate fine-tuning strategy and use the post-pretraining stage to transfer the general semantics learned through self-supervised pretraining on ShapeNetCore . For a fair comparison, our Base and Large models adopt the same architecture as PointGPT regarding layers, hidden size, and attention heads. Tab. 1 shows that: (i) ReCon++ exhibits representation performance significantly surpassing that of other baselines, achieving state-of-the-art results. (ii) Particularly, ReCon++ achieves a remarkable accuracy of 95.25% on the most challenging ScanObjectNN PB_T50_RS benchmark, boosting the Transformer baseline by +16.14%.
Zero-Shot 3D Open-World Recognition Similar to CLIP , our model aligns the feature space of languages and other modalities, which results in a zero-shot open-world recognition capability. In Tab. 2, we compare the zero-shot 3D open-world object recognition models to evaluate the generalizable recognition capability. Following OpenShape , we evaluate on ModelNet , ScanObjectNN , and Objaverse-LVIS . Objaverse-LVIS is a benchmark involving 47K clean 3D models of 1,156 LVIS categories . We compare ReCon++ with 2D inference methods, ShapeNet pretrained methods, and “Ensembled” datasets-pretrained methods. It can be concluded from Tab. 2: i) Compared to 2D inference and ShapeNet-pretrained methods, ReCon++ demonstrates significantly superior performance, showing the necessity of 3D point clouds as inputs and scaling up. ii) Compared to state-of-the-art methods trained on “Ensembled” datasets, ReCon++ demonstrates superior or on-par performance across all benchmarks. Notably, ReCon++-L achieves a remarkable Top-1 accuracy, which is +0.6% and +7.2% higher than Uni3D-L on the most challenging Objaverse-LVIS and ScanObjectNN benchmarks, respectively.
2 Multimodal Comprehension with ShapeLLM
Quantitative Analysis To assess the comprehensive capabilities of ShapeLLM, we first quantitatively compare various baselines and our model on the proposed 3D MM-Vet benchmark using GPT-4. Following ModelNet-C and ModelNet40-C , we construct 3D MM-Vet-C to benchmark the robustness against 3D corruptions.
3D MM-Vet. Tab. 3 shows the detailed results of ShapeLLM on different tasks of 3D MM-Vet. It can be observed that ShapeLLM significantly outperforms PointLLM across various metrics, particularly in Embodied Tasks. This substantiates our model’s versatile capability in addressing real-world scenario tasks.
3D MM-Vet-C. Tab. 4 shows the comparison of model robustness against “single-view”, “jitter” and “rotate” corruptions, which are the most common corruptions in real-world scenarios. The results demonstrate significantly superior robustness of ShapeLLM against corruption, indicating stronger potential in real-world applicability.
Qualitative Analysis Fig. 5 illustrates qualitative examples of ShapeLLM in multimodal dialogue. ShapeLLM is capable of supporting general VQA, embodied task and action planning, as well as 6-DoF pose estimation. Notably, due to the strict spatial relationship inherent in 6-DoF bounding box coordinates, we observe that LLMs easily grasp such patterns and consistently produce valid coordinates.
Discussions
Fig. 6 illustrates the visualization of the attention map in the last cross-attention layer, documenting the image query to which each local patch in the attention map primarily attends. It provides evidence that multi-view alignment achieves geometrically informed spatial understanding, which may implicitly encompass the estimation of the object pose and a more profound knowledge of 3D spatial relationships.
2 Is ShapeLLM grounded in physical worlds?
Tab. 5 compares ShapeLLM with image-only methods on 3D referring expression grounding (REG) of 6-DoF poses on GAPartNet. The results show that: i) Image-only methods cannot perform zero-shot geometry-necessary 6-DoF pose estimation. ii) Compared to image-only methods with 2D to 6-DoF pose estimation fine-tuning or in-context prompting, ShapeLLM still performs significantly better. It demonstrates the necessity of geometry and the difficulty of the ill-posed 2D to 6-DoF pose estimation problem.
3 Can ShapeLLM generalize to unseen objects?
Fig. 7 shows the part understanding examples of unseen objects. While ShapeLLM’s 6-DoF pose estimation is trained on GAPartNet, which primarily consists of indoor articulated furniture. It has demonstrated promising generalization potential of spatial understanding on the open-world objects, paving ways for scaling up spatial-awareness training.
Related Works
Interaction-oriented 3D Understanding Interaction with 3D objects typically involves concept-only interaction and physical-grounded interaction . The former works focus on 3D perception and semantic parsing, such as 3D object recognition and scene perception . By utilizing language for open-ended interaction in 3D, a number of works demonstrate successful 3D scene QA , grounding , and captioning . Recently, some works propose to utilize foundation models like LLMs or CLIP for open-ended 3D object recognition and scene segmentation . Guo & Zhang et al. utilizes ImageBind and LLaMA-Adapter to realize point cloud-based interactive QA. Following LLaVA, PointLLM conducts supervised fine-tuning by constructing a visual instruction-following dataset. Other works focus on scene-level tasks utilizing comprehensive 2D features or 3D features distilled from 2D images into LLMs . The second kind of interaction typically requires physical understanding in 3D, such as part understanding , 6-DoF pose estimation , particularly useful for human-object interaction (HOI) and robotic manipulation and complex robotic planning . In this work, we focus on both physical and conceptual interactions with 3D shapes for embodied understanding.
Multimodal Large Language Models Multimodal comprehension, which allows human interaction with textual and visual elements, has witnessed significant advancements, particularly in extending LLMs like LLaMA . The early efforts predominantly revolved around integrating LLMs with various downstream systems by employing it as an agent . Significant success has been demonstrated within this plugin-style framework. Due to the remarkable capabilities of LLMs, aligning the visual semantic space with language through parameter-efficient tuning and instruction tuning has emerged as the prevailing approach in current research. To further enhance interactive capabilities, some approaches have been developed towards visual-interactive multimodal comprehension by precisely referring to instruction tuning . Another family advances the developments of LLMs endowed with content creation beyond comprehension, notable efforts include DreamLLM , GILL , Emu , SEED , NeXt-GPT , and Kosmos-G .
Conclusions
This paper presents ShapeLLM, a 3D multimodal LLM for embodied interaction, capable of generalizable recognition and embodied interaction comprehension. We first propose a novel 3D point cloud encoder, ReCon++, by utilizing multi-view distillation and scaling up 3D representation learning, which serves as the foundation 3D representation encoder for ShapeLLM. Then, we perform 3D visual instruction tuning on constructed instruction-following data for general and embodied comprehension. We also established a 3D evaluation benchmark, 3D MM-Vet, severing as assessing the 4-level capacity in embodied interaction scenarios, varying from basic perception to control statements generation.
References
Appendix A Additional Experiments
The local and transformation-invariant 3D geometric embeddings \mathbf{x}_{i}=\mathop{\text{MAX}}\limits_{\mathbf{p}_{i,j}\in{\mathcal{N}_{i}}}\big{(}\Phi_{\gamma}\left(\xi_{i,j}\right)\big{)} for is used as 3D token embeddings of ReCon++, where is a per-point MLP point feature extractor and is the feature of -th neighbour point in the neighbourhood . Let be multi-view image global queries and be the global text query. ReCon++ outputs the local and global 3D point cloud representations by taking 3D embeddings and global queries as inputs:
In addition, inspired by prefix-tuning and dream queries , we append -length learnable embeddings , , as visual prompts representation for adaptively modulating different semantic information encoded in APE, local and global ReCon++ representations, respectively.
Formally, the encoded 3D representations to ShapeLLM can be written as:
Tab. 6 shows the ablation study of each input component by supervised fine-tuning with different input representations, demonstrating that it is necessary to employ all designs for achieving decent performance on both 3D comprehension and real-world grounding.
Visual Prompt Number Fig. 8 shows the performance of ShapeLLM using different numbers of prompts, including 1, 8, 16, 32, and 64. This ablation study has shown that a different number of prompts leads to varied improvements, and the optimal setting is 32. This observation is similar to VPT where the prompts used to modulate Transformer attention should be studied .
A.1.2 Baseline Improvement
Can we improve the baseline to bridge the gap between PointLLM and ShapeLLM? In Tab. 6, we study two technical factors that are contributed by ShapeLLM: 3D point cloud encoder and SFT data.
Improvement from encoder. (Line 1) First, by changing PointLLM’s encoder to ReCon++, a significant improvement of +4.20% is obtained. This demonstrates the significantly better 3D representation extraction of ReCon++ compared to ULIP-2. It is consistent with previous findings in Tab. 1 and Tab. 2 that ReCon++ outperforms ULIP-2 by a large margin regarding 3D representation transferring learning and zero-shot learning.
Improvement from data. (Line 2) As stated in Sec. 2.2, we have constructed instruction-following data for supervised fine-tuning (SFT) using GPT-4V involving comprehensive topics. By further using the SFT data curated by us, PointLLM’s performance gap to ShapeLLM has been fulfilled. This demonstrates the superiority of our SFT data, where the decent quality comes from the more advanced GPT4-V model using multi-view images and the comprehensive topics covered in the data.
A.2 Multimodal Comprehension with ShapeLLM
Generative 3D Object Recognition & Captioning Following PointLLM , we conduct generative 3D recognition and captioning experiments. Tab. 8 shows 3D object classification overall accuracy (%) and captioning performance evaluated by GPT-4 and data-driven metrics: Sentence-BERT (S-BERT) and SimCSE . It can be observed that ShapeLLM consistently outperforms other methods across all metrics, demonstrating robust recognition and instruction-following capabilities.
Note that similar to PointLLM’s findings, we also notice that the 3D captioning performance evaluated by traditional metrics like BLEU-1 , ROUGE-L , and METEIOR are highly unreliable in accurately revealing the response quality. This is further demonstrated by human-oriented evaluation, such as the preference win rate comparison presented next.
Singe-View Point Cloud Inputs As stated in Sec. 4.2, we construct 3D MM-Vet-C which studies three kinds of corruptions commonly met in real-world scenarios: “single-view”, “jitter”, and “rotate”. Among these corruptions, the “singe-view” issue stands out as the most critical challenge since obtaining the objects’ complete point clouds is non-trivial, similar to multi-view images. As a result, everyday real-world robots only get single-view 3D perceptions with sensors such as RGB-D . Fig. 9 shows the qualitative examples of ShapeLLM-13B’s response using single-view point cloud inputs, demonstrating surprisingly outstanding robustness in processing such occluded inputs.
Human Win Rate Comparison GPT-4 is widely used as an evaluator in natural language and vision language processing, as seen in recent modern benchmarks like MM-Bench and MM-Vet. Recent studies have demonstrated that ChatGPT-based evaluation is more closely aligned with human preferences compared to traditional metrics. With GPT4-turbo, the standard deviation of 3D MM-Vet is less than 0.1. To further verify the soundness of the models’ response, we also conduct human evaluation and report the win rate in Fig. 10, where ShapeLLM demonstrates superior preference by humans.
Visual hallucination is a well-known issue in LLMs and MLLMs that generate non-existent objects or identities from the input data, significantly compromising their multimodal comprehension capabilities and may pose safety risks . Recent research suggests that hallucination may stem from biases in training data, particularly within supervised fine-tuning data, or inappropriate generation strategies. In Fig. 11, we qualitatively demonstrate the illusion evaluation of ShapeLLM compared to other methods. We assess the model’s ability to counteract illusions by prompting it with detailed captions and misleading questions. The results in Fig. 11 demonstrate that previous methods Point-Bind&Point-LLM and PointLLM suffer from the problems of mis-recognition and mis-associating non-existing identities.
A.3 Representation Transferring with ReCon++
Linear SVM evaluation can be used to evaluate the discriminative quality of pretrained features . The results on ModelNet40 are shown in Tab. 9. It shows that our ReCon++ outperforms Point-BERT, which also uses plain Transformers with contrastive objectives, by a clear margin of +6.2%. Compared to hierarchical Transformers methods, our ReCon++ outperforms PointM2AE by +0.7%.
Few-shot learning is critical for evaluating the representation transferring capabilities in data and training efficiency. We conduct few-shot 3D object recognition experiments on the ModelNet40 dataset, and the results are shown in Tab. 10. Our ReCon++ achieves state-of-the-art performance in all the benchmarks compared to previous works.
Appendix B Additional Information about 3D MM-vet
Unlike classification or regression tasks, language generation tasks lack a definitive ground truth that can comprehensively cover diverse real-life scenarios. Therefore, evaluating the alignment of model-generated results with the question and assessing their appropriateness becomes a challenging problem, requiring a reasonable quantitative score. Fortunately, we have observed the recent surge in the popularity of GPT, providing us with a dependable tool for conducting open-ended evaluations.
To enhance the performance of GPT, we employ a few-shot style in-context prompt. This involves feeding GPT with prompts from evaluative examples and instructing it to generate scores. Specifically, we present prompts to obtain a score ranging from 0 to 1, indicating the degree of similarity between the model-generated answers and the ground truths we provided. When implementing this approach, we observed that results generated multiple times may vary a lot. To address it, we apply the same evaluation setting to a single answer for iterations, obtaining the average result as the final score for a precise answer. The score of an answer and the total score of answer set are calculated by:
Here we set , and is the score of the test of answer . The average score for a specific capability is the sum of scores in category answer set :
where is the number of answers in each capability set.
To mitigate excessive standard deviation, we opt for GPT-4 in a series of scoring rounds to get rounds of outputs with a standard deviation below 0.1. This choice is motivated by the enhanced stability offered by GPT-4 , in contrast to GPT-3.5 , where scores across different rounds exhibit significant variability.
B.2 Analysis
The 3D MM-Vet evaluation benchmark consists of 5 different categories of questions. In Fig. 12 we report the distribution of problem categories. The knowledge and General Visual Recognition parts contain multiple subparts that comprehensively evaluate these capacities and thus hold higher proportions. Fig. 13 shows an example of how we prompt GPT-4 for 3D MM-Vet evaluation. Fig. 14 and Fig. 15 illustrate additional examples of 3D MM-Vet Q&As.
Appendix C Implementation details
ReCon++ Following the standard ViT architecture, we design four different model structures consistent with prior work . The model parameters are shown in Tab. 11. Following OpenShape , we employ four datasets as pretraining data, namely Objaverse , ShapeNet , ABO , and 3D-FUTURE . Each point cloud sample has a size of 100006, where the first three dimensions represent coordinates, and the latter three dimensions represent values.
Regarding the masked modeling strategy, we experimented with both random masking strategies and the latest causal masking strategy. Using causal masking as initialization significantly improves transfer learning capability, as shown in the ablation experiments in Tab. 12. Specifically, the point encoder of ShapeLLM still employs the original local-guided stop-gradient strategy . Additionally, to enhance global classification and retrieval capabilities, we backpropagate gradients from the global branch to the local branch in open vocabulary zero-shot experiments, as demonstrated in the ablation experiments in Tab. 12.
ShapeLLM We use the LLaMA model as our LLM backbone, with the 7B and 13B Vicuna-1.1 checkpoint as the default settings. We partitioned the point clouds into 512 patches using furthest point sampling and k-nearest neighbors. Similar to other MLLMs , we employ a 2-layer MLP with GELU as the projector, with hidden layer sizes of 1,024 and 2,048, respectively. Note that different projector parameters are utilized for absolute positional encoding, local, and global features. Through training the projector, multi-scale and multi-mode features of the point cloud are mapped into the text space. After adding two special tokens, the vocabulary size becomes 32,003.
Appendix D Training details
ReCon++ Due to the sensitivity of the Chamfer Distance loss to accuracy, all experiments were conducted at FP32 precision using 8 × 80G A800 GPUS. We still use the strategy of contrast with reconstruct . To save parameter tuning time and improve performance, we divide the training process into two stages: the reconstruction stage based on mask modeling and the cross-modal alignment stage based on knowledge distillation. For transfer learning classification tasks, ReCon++ is pretrained on 1,024 points. For zero-shot tasks and ShapeLLM tasks, ReCon++ is pretrained on 10,000 points. Further details regarding the hyperparameter settings are documented in Tab. 13.
ShapeLLM All experiments were conducted using 8 80G A800 GPUs with a BF16 data type. During the multimodal alignment stage, we train our model for one epoch with a batch size 256 and a learning rate 2e-3. During the instruction tuning stage, we train our model for one epoch with a batch size of 128 and a learning rate 2e-5. Throughout both stages, we employ flash-attention , the AdamW optimizer, and a cosine learning rate scheduler . For the entire training process, the 7B and 13B models require approximately 10 and 20 hours, respectively. Further details regarding hyper-parameters are documented in Tab. 13.
Appendix E Additional Related Work
Research on 3D Representation Learning encompasses various methods, including point-based , voxel-based , and multiview-based approaches . Point-based methods have gained prominence in object classification due to their sparsity yet geometry-informative representation. On the other hand, voxel-based methods offer dense representation and translation invariance, leading to a remarkable performance in object detection and segmentation . The evolution of attention mechanisms has also contributed to the development of effective representations for downstream tasks, as exemplified by the emergence of 3D Transformers . Notably, 3D self-supervised representation learning has garnered significant attention in recent studies. PointContrast utilizes contrastive learning across different views to acquire discriminative 3D scene representations. Innovations such as Point-BERT and Point-MAE introduce masked modeling pretraining into the 3D domain. ACT pioneers cross-modal geometry understanding through 2D or language foundation models such as CLIP or BERT . Following ACT, ReCon further proposes a learning paradigm that unifies generative and contrastive learning. Additionally, leveraging foundation vision-language models like CLIP has spurred the exploration of a new direction in open-world 3D representation learning. This line of work seeks to extend the applicability and adaptability of 3D representations in diverse and open-world/vocabulary scenarios .
Appendix F Future Works
ShapeLLM has made significant progress in advancing 3D shape understanding and embodied perception through MLLMs. Future endeavors aim to scale up embodied understanding training using datasets larger than GAPartNet , potentially leading to open-vocabulary part-level comprehension, including 6-DoF pose estimation. Excitingly, there is a vision to establish a unified framework capable of comprehending not only 3D shapes but also entire 3D scenes. To enhance real-world applications on robots, a promising approach involves a robotics co-design that effectively connects 3D representations with downstream language-based tasks . Additionally, addressing efficiency for real-time deployment is crucial, emphasizing techniques like model compression .