GPT4Point: A Unified Framework for Point-Language Understanding and Generation
Zhangyang Qi, Ye Fang, Zeyi Sun, Xiaoyang Wu, Tong Wu, Jiaqi Wang, Dahua Lin, Hengshuang Zhao
Introduction
The recent Large Language Models (LLMs) have demonstrated remarkable advancements in the field of natural language processing. Inspired by their powerful capabilities, researchers have also explored Multimodal LLMs (MLLMs), via adapting LLMs into various modalities like images , audio and videos . The proliferation of extensive image-text pair has crucially enabled 2D MLLMs i.e., Vision Language Models (VLMs) to interpret images through textual representations. Concurrently, there is a growing trend in utilizing these multimodal models for guiding text-to-image generation . This represents a form of compression and reconstruction, exploring how to accurately recover and even edit the input image using controllable image generation models. However, despite the impressive capabilities of MLLMs in handling multiple modalities, they still face significant limitations in understanding and accurately interpreting the 3D world, a critical need for various important downstream applications like intelligent robotics and augmented reality.
Recent efforts to develop 3D MLLMs have notable limitations. Some prioritize the overall scene and focus primarily on the spatial coordinates of objects, often neglecting the geometric details of individual objects. This can lead to a limited understanding of the 3D world. Meanwhile, these methods generally convert 2D image features into 3D representations , which leads to a substantial loss of geometric accuracy. 3D geometry information is important in understanding. As shown at the bottom of Fig. 1, the VLM fails to recognize the four-sided face object while our GPT4Point can figure out the anomalies. Concurrent works focusing on utilizing 3D features directly exhibit notable limitations. PointBind exhibits a deficiency in training and demonstrates restricted text referencing abilities due to the limited dataset. On the other hand, PointLLM necessitates the training of the corresponding Language Model (LLM) component and does not possess the capability to expand into text generation.
We present GPT4PointFirst author is the intern at Shanghai AI Laboratory., a novel unified framework for point-language understanding and generation. GPT4Point introduces the 3D object MLLM, which is a groundbreaking language model that fully utilizes point clouds to perform various point-text tasks as shown in Fig. 1. We utilize a Bert-based Point-QFormer for point-text feature alignment. Aligned features are separately input into the LLMs for text inference tasks and Diffusion for 3D object generation tasks. It is worth noting that, given a low-quality point cloud feature as a condition, GPT4Point can generate higher-quality results while maintaining the geometric shapes and colors by using point-text aligned features called controllable text-to-3D.
To tackle the scarcity of object point-language data , we leverage the Objaverse-XL dataset to develop an automated, effective data annotation engine Pyramid-XL. It employs Vision Language Models (VLMs) for generating text annotations. Pyramid-XL solves the problem that VLMs can not understand multi-view images directly. By synthesizing captions from multi-views obtained by the VLMs, the text annotation is stratified into three hierarchical levels, ranging from low to high, ultimately leading to precise annotations. Apart from the data engine, we establish an object point-text benchmark for assessing point multimodal model capabilities in recognition and text inference tasks, such as 3D object point cloud captioning, and Q&A. This benchmark also provides a critical standard for evaluating 3D object generation, while current assessments often rely on qualitative judgments from rendered images without a direct evaluation in 3D space . Only relying on rendering images may lead to misunderstanding, for instance, in the bottom right of Fig. 1, a failure case produced by 3D generation (a bear has two bodies), makes 2D VLMs and even humans fail to recognize its anomaly but our model can identify with anomalies easily.
Our paper makes three major contributions:
We present the unified framework for point-language understanding and generation GPT4Point, including the 3D MLLM for point-text tasks and controlled 3D generation.
Introducing the automated point-language dataset annotation engine Pyramid-XL based on Objaverse-XL, currently encompassing 1M pairs of varying levels of coarseness and can be extended cost-effectively.
Establishing a novel object-level point cloud benchmark with comprehensive evaluation metrics for 3D point cloud language tasks. This benchmark thoroughly assesses models’ understanding capabilities and facilitates the evaluation of generated 3D objects.
Related Work
Multi-modal large language models (MLLMs). Large Language Models (LLMs) have demonstrated robust capabilities in language comprehension, reasoning, and generalization . Building upon this, Multimodal Large Language Models (MLLMs) extend these reasoning skills to additional modalities such as image , audio , and video . Typically, MLLMs align target features with textual features and then integrate them with LLMs for various text inference tasks. Some train the whole architecture from scratch and others utilize pretrained LLMs. In the realm of 3D MLLMs, existing models either rely on 2D image information or simply align low-quality textual phrases with points . To solve these problems, we introduce a novel 3D MLLM designed for diverse point-text tasks. Our model, featuring a Point Q-Former based on Bert , aligns two domain features and integrates an LLM for text-based reasoning tasks, advancing the field of 3D multimodal understanding.
Language-driven 3D object understanding. 3D point cloud multimodal models encompass a broad spectrum, generally categorized into those focusing on the entire scene containing multiple objects and those focusing on individual objects. The former places more emphasis on the relative positions of objects in the scene rather than their geometric shapes; Here, we primarily focus on the latter. In a self-supervised way, powerful backbones like PointBert for object points have been obtained . Then, point cloud language pretraining attempts to align the point cloud modality and the text modality. Some methods try to convert point clouds to depth images for alignment with text using CLIP . Tri-modal approaches such as ULIP integrate point cloud, text, and image data. However, these methods all exclusively use 2D images, either explicitly or implicitly. Our work differs by directly aligning 3D point-text modalities, completely removing the dependency on image data.
Text-to-3D generation. Text-to-image generation models have experienced significant advancements recently , yet text-to-3D models face challenges due to limited 3D data availability. Current approaches often rely on optimizing Neural Radiance Fields (NeRF) representation with Score-Distillation-Sampling (SDS) loss . While these optimization-based methods still fall short in robustness, speed, and generalization. Alternatively, Point-E and Shap-E employ feed-forward 3D generative models trained on large, undisclosed 3D datasets, offering better generalization and faster processing. However, these models often produce random, uncontrollable outputs with low-quality textures. To solve these limitations, we leverage point-text features to enhance the controllability of feed-forward models. This approach uses a low-quality point-text feature as a condition that allows for maintaining specific shapes and colors, thereby enabling the generation of higher-quality 3D objects.
Methods
This section provides an overview of our data text annotation engine and model architecture. In Sec. 3.1, we introduce Pyramid-XL, our point-language dataset annotation engine, discussing its design, function, and the progression from low-quality descriptions to ultimately precise and detailed ones. Then, in Sec. 3.2, we delve into GPT4Point’s architecture, explaining how to align point and text and demonstrating how LLM and point diffusion models contribute to unified understanding and generation.
The public release of the large-scale Objaverse dataset and its successor Objaverse-XL includes 800K and 10M objects respectively, providing a vast amount of 3D object data. However, these objects lack corresponding text descriptions. We plan to use the rendered images of the objects as input and obtain textual descriptions through a trained Vision Language Model (VLM), however, we find that direct input of multi-view images into the VLM does not enable it to understand their 3D structure and give precise descriptions, as shown in the top right of Fig. 3. Hence, Pyramid-XL employs a hierarchical pipeline, evolving from initial low-quality descriptions to achieve ultimately precise and detailed results.
Pyramid-XL Single-View Caption (Level 1): We use the primary VLM model BLIP-2 to generate concise descriptions, approximately 10 words in length, from a single-view rendered image. Multi-View Caption (Level 2): This level synthesizes multiple Level 1 descriptions by GPT-4 to create comprehensive multi-view captions which has approximately 30 words. VLM Instruction Caption and QA Pair (Level 3): Utilizing the view with the highest CLIP score, selected from textual descriptions, we engage the advanced VLM to produce detailed dense captions and a corresponding QA dataset. In terms of scale, Pyramid-XL is employed to annotate over 1M objects with Level 1 captions, 660K objects with Level 2 captions (same as Cap3D ), and 70K objects with Dense Captions including QA data. To assess the impact of text granularity on training, we designate the 1M Level 1 captions as the pretrain dataset, while a smaller set of detailed Level 3 data is used for instruction tuning. This methodology mirrors practices in the vision field, where models are initially pretrained on large volumes of coarser data and subsequently finetuned on more detailed data from specialized domains. Detailed experimental results of this approach are presented in Sec. 5.3.
2 Model Architecture
GPT4Point consists of two stages as illustrated in Fig. 2. In Stage1, we focus on point-text alignment using the Point-QFormer, a Bert-based structure similar to the Q-Former in BLIP-2 . This stage involves supervision through three tasks related to recognition and text reasoning. In Stage2, only the point cloud is input into the point encoder and Point-QFormer to obtain aligned features, which are then devided into two branches: the LLM Branch and the Diffusion Branch separately. These branches supervise text comprehension and object generation tasks, respectively.
Here, represents the loss for three tasks, and we have set the weight ratios between them all to 1. In the final layer of , a fully connected layer maintains consistency between the dimensions of and .
Stage2: point understing and generation. After the point-text feature alignment, we proceed with understanding and generation tasks. It’s important to note that here we only input the point cloud into the Point Encoder and Point Q-Former to obtain the aligned feature. For the understanding task, a Large Language Model (LLM) is integrated with the Point Q-Former. The semantically integrated point cloud features are represented as . The textual feature tokens are obtained from the LLM’s own tokenizer. The objective function is defined as follows:
indicates Point Q-former including a fully connected layer in its last layer to ensure consistency between the dimensions of and . represents the loss function from the Point Caption task alone.
For 3D object generation, we utilize the features obtained from low-quality point clouds via the Point Q-Former as conditions inputted into the text-to-3D framework. This process results in the generation of refined 3D objects that maintain consistency in shape and color with the original point cloud. A notable distinction from the LLM branch is that we have not only frozen point cloud diffusion but also frozen Point Q-Former. As shown in Fig. 2, we employ a single fully-connected layer to project the aligned features into the CLIP token embedding space, referred to as , and then concatenate these with the original text embeddings using the CLIP tokenizer. The output from the CLIP text encoder, enriched with information from the original point cloud, is instrumental in enabling effective text-to-3D generation. The final output is achieved using Point-E. This framework is inspired by BLIP-Diffusion techniques used in subject-driven 2D generation. However, the key distinction here from BLIP-Diffusion lies in the way we concatenate the Clip text token and Q-Former feature. This difference may also stem from variations in the data volumes between 2D and 3D, which will be thoroughly examined in the appendix.
Benchmarks and Evaluation
Evaluating the performance of multimodal models presents significant challenges due to the lack of mature metrics for assessing the quality of generated texts. For 3D objects, benchmarks primarily rely on human judgment or GPT-based assessments . There are two key issues to consider in this context. Firstly, the evaluation process involves a certain degree of subjectivity. Identical results might receive varying scores, leading to an element of randomness. Secondly, each evaluation incurs time and monetary costs. In this section, we present the evaluation benchmark we have proposed, which is primarily designed to be objective, ensuring repeatability and verifiability. Sec. 4.1 outlines the composition of our test set. Sec. 4.2 addresses the evaluation of recognition capabilities, while Sec. 4.2 provides a detailed assessment of text inference abilities.
We leverage the Objaverse dataset , aligning it with LVIS categories , to create Objaverse-LVIS validation and test sets. In Objaverse-LVIS, we exclude scenes with complex settings, such as indoor houses or outdoor parks, focusing more on scenarios with single objects or combinations of multiple objects. We construct validation and test sets, each containing 1K objects. Compared to the PointLLM , which uses only 200 unfiltered objects as a test set, our larger set of 1K objects better measures the model’s generalization capabilities. For textual descriptions, we initially use Pyramid-XL to get initial annotations, followed by multiple rounds of expert manual revisions, ensuring comprehensive and accurate descriptions.
2 3D Object Recognition
3D object recognition represents the classification capabilities of 3D multimodal models and the ability to match point cloud features with textual features. Objective measures, like accuracy, are typically used for evaluation.
Zero-shot point classification. Zero-shot point classification is considered a classic task in this domain. The widely used ModelNet40 dataset , which includes 2,468 objects across 40 categories, serves as a benchmark to evaluate a model’s classification capabilities. In the multimodal context, the typical approach involves using the text ’a 3D model of [name]’ as input to match with the point cloud modal features. The accuracy metric ACC@1, indicating the precision of top-1 rankings, best reflects the model’s ability to accurately match object categories.
3D point-text retrieval. In 3D Point-Text Retrieval, we initially select 128 candidates based on point-text feature similarity and then re-rank these candidates using matching scores. Unlike classification tasks where the text usually involves simple category names, here the text can be more complex descriptions. The evaluation metrics used are similar to those in image-text retrieval. We employ R1, R5, and R10 metrics to measure the accuracy of the top 1, 5, and 10 results in correctly matching points to text and vice versa.
3 3D Object Text Inference
3D object text inference deeply represents the understanding capabilities regarding objects, including 3D object point cloud captioning and 3D point cloud question answering.
3D point cloud captioning. This task primarily evaluates the model’s ability to provide an overall summary of a 3D object. The captions in the Objaverse-XL-LVIS caption test set are mostly within 30 words and accurately describe the object’s geometry, color, and state. And we predominantly employ common image description metrics, such as BLEU1, BLEU4, METEOR, ROGUE-L, and CIDEr for evaluation.
3D point cloud question answering. In addition to point cloud captioning, 3D point cloud question answering explores object details through multiple rounds of dialogue. For instance, we can further explore the color or shape of specific parts of an object or even infer its simple usage. The curated Objaverse-XL-LVIS short QA 1K test set features concise, straightforward questions and answers, allowing us to conveniently calculate answer accuracy. Besides accuracy, we also use metrics from captioning to evaluate model performance. It is important to note that, for a fair comparison, we solely utilize zero-shot learning, meaning no fine-tuning is conducted on this kind of short QA dataset.
We configure our setup to process 8,192 input point clouds, utilizing Point-BERT as the backbone. This transformer-based network excels in capturing geometric and semantic features of object point clouds. And the backbone is pretrained through retrieval tasks like ULIP-2 . We employ OPT and FlanT5 as Large Language Models (LLMs). For the training process, we adopt an initial learning rate of 1e-4, weight decay of 0.05, batch size of 32, and the AdamW optimizer . All hyperparameters remain unchanged in both stages. The training process takes 10 epochs for each stage on 8 A100 GPUs.
2 Evaluation and Diverse Tasks
We evaluate our model on the benchmark we proposed in Sec. 4, which includes 3D object recognition and 3D object text inference. Additionally, we demonstrate the model’s capability for controllable text-to-3D generation.
3D object recognition. Recognition capabilities are shown in Sec. 4.3, with zero-shot classification results on the right side. Our approach demonstrates superior performance, outperforming the Vision Language Model(VLM) InstructBLIP by 12.42 points and surpassing PointLLM by 2.57 points. Notably, PointLLM employs a generative approach to generate the text results by a prompt, limiting its direct recognition capabilities. The results for 3D point-text retrieval are shown on the left side. Our GPT4Point model outperformed other VLMs . The results quantitatively highlight the challenges of single-viewpoint 3D object occlusions and biases, emphasizing our approach’s advantages over other image-text models.
3D object text inference. Model’s text inference capabilities are displayed in Tab. 2. On the left, the results of 3D object point cloud captioning confirm GPT4Point’s superiority over pretrained VLMs and PointLLM. Notably, the Point Q-Former structure allows freezing the LLM, significantly reducing training parameters. The results for 3D point cloud question answering on the right side show that GPT4Point achieved the best zero-shot accuracy, surpassing InstructBLIP by 11.7 points and outperforming PointLLM by 4.2 points. Alongside quantitative results, Fig. 4 qualitatively demonstrates its detailed answers and multi-turn dialogue capabilities, with more examples in the appendix.
Controllable text-to-3D object generation. Here, we showcase the generative capabilities of our model. Given features of low-quality point clouds along with textual descriptions, we can generate corresponding higher-quality point clouds, making text-to-3D more controllable. Fig. 6 displays experimental results, We compare our point feature condition with text or single image condition in Point-E, demonstrating that aligning features using both point cloud and textual information significantly improves guidance for point cloud generation. It is worth noticing that when compared to a single view image rendered from the original 3D model, our Point Q-former feature serves as a better condition that contains richer information about the geometric shape and detailed color information of 3D objects. We believe this is the first step towards the point cloud editing.
3 Assessing the Effectiveness of Pyramid-XL
In this section, we demonstrate the effectiveness of Pyramid-XL in obtaining high-quality point-text annotations. We focus on two tasks: fine-tuning Point-E for 3D object generation using dense captions and utilizing annotations of varying granularities on the QA benchmark.
Finetune the Point-E with Level 3 Caption. We fine-tuned Point-E base-40M text-vec model using 70K Level 3 VLM instruction captions from Pyramid-XL for 3D object generation. The results in Fig. 5 show significant improvements in geometric details and and color fidelity in point clouds, especially in objects like baskets and Halloween costumes, compared to Cap3D .
Ablation study in model pretraining. Our ablation studies on Pyramid-XL, detailed in Tab. 4, investigated the impact of pretraining data scale and quality on model performance. The comparison between the first two rows indicates that using a large volume of coarse annotations boosts baseline performance. Additionally, incorporating a higher proportion of detailed Level 3 annotations leads to improved QA scores, with 80% yielding near-optimal results.
We introduce the innovative GPT4Point, a Unified Framework for point-language understanding and generation including the 3D MLLM for point-text tasks and controlled text-to-3D generation based on low-quality point feature. We develop Pyramid-XL, a point-language dataset annotation engine. This setup constructs a large-scale database over 1M objects of varied coarseness levels from the Objaverse-XL dataset. Furthermore, we establish an object-level point cloud benchmark with specific metrics for evaluating 3D point cloud-language tasks. This benchmark provides a comprehensive approach to assess both the understanding abilities of 3D multimodal language model and the quality of generated objects.
In this supplementary material, we extend the discussions presented in the main conference paper. Sec. B provides a more in-depth exploration of related work, focusing on defining the scope of large language models family and examining the developments in point-text multimodal approaches. Sec. C supplements more details about the data annotation engine: Pyramid-XL and the diffusion architecture. Moving to Sec. D, we expand on the superiority of our benchmark. Initially, we introduce examples from our ObjaverseXL-LVIS QA 1K dataset, which includes concise QAs for evaluation and long QAs for instructive tuning. Then we show more 3D generation failure cases where GPT4Point can figure it out while 2D VLM can not to underscore the necessity and relevance of our 3D point-text benchmark. Finally in Sec. E, we give more qualitative results of Point-text inference tasks including caption and QA tasks and Controllable point diffusion.
Appendix B Additional Related Work
In this section, we provide detailed insights into related work. Sec. B.1 classifies key concepts of large language models, including LLMs, MLLMs, and VLMs. Sec. B.2 present the evolution of point-text multimodal models through an illustrative flowchart.
Although the concepts related to large language models are already familiar, we still wish to detail these concepts here. We briefly introduce some families of LLMs and MLLMs. First are the LLMs based on the Transformer architecture, such as ChatGPT and GPT-4 . Currently, there are several open-source, deployable models . After extensive pre-training on a vast corpus, they exhibit strong comprehension and reasoning abilities. Multimodal Large Models (MLLMs) aim to enable LLMs to understand information in other modalities. The fundamental approach involves retrieving text features with other modality features. Among them, image-text multimodal large models, also known as 2D MLLMs or Visual Language Models (VLMs), stand out due to the abundant image-text pairs and strong image backbones provided by computer vision . Beyond images, there are other modalities, such as Audio MLLMs that combine with the audio modality and Video MLLMs with the video modality . In the 3D domain some existing work, like 3D-LLM , utilizes 2D image features combined with depth projections to generate 3D features. We propose a unified text understanding and generation model based on point clouds and develop a real 3D MLLM.
B.2 The development of Point-text Multimodal
In this section, we delve into the evolution of point-text multimodal models for single objects.
Backbone Development: The foundational aspect of our methodology lies in the robust development of the backbone for handling point clouds. Similar to the methodologies applied to texts and images, point clouds undergo a self-supervised training strategy to establish a strong foundation. Notably, we leverage the innovative PointBert framework, which divides point clouds into patches and executes a reconstruction process on masked patches. This is achieved through the utilization of a Transformer-based backbone, imparting a powerful and adaptive feature extraction capability to our model.
Text Modality Alignment: Drawing inspiration from the successful model CLIP , our approach incorporates a phase dedicated to aligning point patches with textual features. This strategic alignment augments the backbone’s inherent ability to process textual information seamlessly. By fusing the spatial understanding of point clouds with the semantic richness of textual data, our GPT4Point achieves a more comprehensive and nuanced representation, enhancing its overall performance.
3D MLLMs Integration: Building upon the successful alignment of point patches with textual features, the next crucial step involves the integration of point features into Large Language Models (LLMs). This integration mirrors approaches seen in Vision Language Models (VLMs) and extends their capabilities to comprehend and interpret point cloud data. The seamless fusion of 3D spatial information with the linguistic context empowers Large Language Models (LLMs) with a more holistic understanding of the data, enabling them to discern intricate patterns and relationships within the point clouds.
Appendix C Additional Method
Here, we provide additional information to our method. We first give more details about the data text annotation engine Pyramid-XL in Sec. C.1. And then, in Sec. C.2 about the model architecture, we give the details about the point diffusion branch.
First, we introduce the approach to acquire point cloud from Objaverse-XL . Then we introduce the cost and prompts of the our data annotation engine Pyramid-XL. Finally, we give more qualitative results that finetune the Point-E by our Pyramid-XL level 3 dense captions.
Acquire data from Objaverse-XL. Here we detail our processing approach for the Objaverse-XL dataset . It has 10M objects and is the extension of Objaverse-1.0 which only has 800K 3D objects. Objaverse-XL offers only unprocessed downloads for its 3D objects, most of which originate from sources like GitHub. Downloading these mesh files necessitates obtaining the complete project, as materials and related components are often stored in other separate directories. Downloading the raw dataset in this format is impractical due to excessive memory requirements, with an average project consuming about 1GB of space. Therefore, we render object images and clear the cache upon completion to manage space. We render 20 random views of each object, capturing the RGB, alpha values, and depth, which are then used to generate point clouds. In addition to the 780K objects from Objaverse-1.0, we rendered an additional 220K from Objaverse-XL, totaling 1M objects.
The cost of the Pyramid-XL. We now turn our attention to the cost analysis of our data annotation engine, detailed in Tab. S1. The primary costs, detailed under the ’1K Cost’ column, include GPU resources on the left and GPT API usage on the right. We use the same GPU settings as Cap3D , employing A40s on a identical cloud platform. Given GPUs’ parallel processing, costs are equal for single or multiple units. We calculate usage time assuming a single GPU for simplicity. For Level 1, we use BLIP-2 to generate one short caption for one object. It needs 0.074 hours and costs 0.074h\times\1.28/h=\. For Level 2 the cost is the same as the Cap3D . The GPU resource fees include BLIP-2 and CLIP . BLIP-2 generates 8 views for each object and each view has 5 captions, so the fee is \0.095\times 8\times 5=\. And the CLIP uses 0.3h and costs 0.3h\times\1.28/h=\. All GPU resource fee is \3.76+\0.38=\4.170.03/1k tokens and needs 139.3 tokens for each object and the total cost is \139.3/1000k\times\0.03/1k\times 1000=\4.181.28h\times\1.28/h=\1.64$.
We can observe that Level 2 captions account for most of the costs, primarily due to GPT usage fees. Our findings show that using GPT-4 for text-based multi-view caption synthesis doesn’t substantially outperform ChatGPT. Furthermore, by utilizing open-source Large Language Models (LLMs), we can entirely eliminate API call expenses. The other major cost is the GPU resources, as it uses BLIP-2 to generate five captions for each view, which can lead to redundancy in information. We can reduce the number of captions for each view, and even the number of views.
The prompts of the Pyramid-XL. We present the prompt part of the Pyramid-XL data text annotation engine, as illustrated in Fig. S6 and Fig. S7. We primarily focus on illustrating how to construct GPT-based Level 2 captions, ChatCaptioner-based Level 3 short QA pairs, and MLLM-based Level 3 instruction captions and long QA pairs.
For Level 2 captions, we use Level 1 captions of rendered images from 6 views. Through carefully designed prompts, we integrate captions from the 6 captions to obtain a comprehensive and relatively accurate caption with fewer than 30 words. In our paper, we use GPT-4 to get the comprehensive caption but we find that ChatGPT can be replaced by GPT-4 to generate Level 2 captions to reduce the cost.
For Level 3 short QA, we follow the approach outlined in ChatCaptioner . We use ChatGPT or other LLMs (we choose Vicuna-7B ) as the questioner and BLIP-2 as the answerer. By providing appropriate instructions and context (Level 2 caption) to both the LLM and BLIP-2, we observe that, LLM generate diverse questions that that include aspects such as color, type, material, purpose, and more. Also, without restricting the number of words, BLIP-2 tends to output concise answers. These form the basis for our Objaverse-XL short QA dataset.
For Level 3 dense captions, we use the Level 2 caption as context, feed the rendering image that best matches the context into MLLM, and input suitable instructions. Due to a combination of high-quality conversational performance and cost-effectiveness, we choose the Qwen-VL model to generate. The construction method for Level 3 instruction (long) QA pairs is similar to the above steps, with the key difference lying in the variation of instructions.
The effectiveness of Pyramid-XL Level 3 caption. We use dense captions from Level 3 of Pyramid-XL to fine-tune Point-E and compare the results with those of Cap3D, as shown in Fig. S11. Ours significantly outperform Cap3D’s captions, demonstrating the precision of our captions.
C.2 Point Diffusion Architecture
Currently, there are indeed some explorations into controllable text-to-3D work . However, we are attempting to combine understanding and controllable 3D generation together. Here, we offer an in-depth look at the Diffusion branch’s structure in Stage 2, illustrated in Fig. S4. Initially, the point cloud undergoes processing via the Point Encoder (Backbone) and Point Q-Former, yielding Q-Former Tokens. For text, instead of the Point Q-Former’s text tokenizer, we utilize Point-E’s CLIP tokenizer. The resulting text tokens are then concatenated with the Q-Former Tokens. Subsequently, the CLS token from the Text Token is fed into Point-E. The concatenation method in GPT4Point differs notably from BLIP-Diffusion . In BLIP-Diffusion, Q-Former Tokens are inserted between the CLS token and input tokens. In contrast, GPT4Point appends Q-Former Tokens directly to the text token sequence, allowing the CLS token to integrate both geometric and color information, crucial for guiding the 3D generation.
Appendix D Additional Benchmark
In this section, we mainly introduce some additional contents about the benchmark. In Sec. D.1, we give more examples of the ObjaverseXL QA dataset. Note that the short QA dataset is used for evaluation based on the accuracy metric. Then in Sec. D.2, we show more qualitative results about Generation Failure Cases which can not be recognized by 2D VLMs through a single view but are judged by our GPT4Point.
Short QA Dataset We use the short QA dataset for the evaluation of the 3D point cloud question answering task. We selecte categories that overlap with both Objaverse-XL and LVIS , constructing 1K Point-QA data as the test set. The specific samples are presented in Fig. S8, which includes questions covering various aspects such as color, material, composition, category, etc. The answers are concise, with an average word length of 2.32, making them convenient for testing. We use accuracy top-1 as metric and evaluate the model’s zero-shot short QA capability on this dataset.
Long (Instruction) QA Dataset The long (Instruction) QA dataset is for finetuning the model to significantly enhance the model’s conversational capabilities. We impose length constraints on prompts, requiring approximately 50 words for answers to dense caption questions and not less than 10 words for other questions. As illustrated in Fig. S8, we constructed a Long (Instruction) QA dataset for 70K objects, comprising 344,996 QA pairs. Among these, 69K data are used for fine-tuning, while the remaining 1K are reserved for testing. This aims to encourage LLMs to generate long and more comprehensive results.
D.2 Anomalous Objects: Generation Failure Cases
In this section, we will demonstrate more qualitative results to show the failure case which can not be recognized by 2D VLMs through a single view but can be judged by our GPT4Point. In this section, we mainly show the failure cases produced by the state of the arts text-to-3D generation methods like Dream-Gaussian and Fantasia3d . Due to technical constraints, these models are likely to generate 3D objects with multi-heads or multi-bodies. If provided with render images from only a single perspective, 2D VLMs , and even humans in most cases, may make incorrect judgments, as illustrated in the upper part of Fig. S9. This hinders the assessment of 3D object generation. However, our GPT4Point provides a better solution to this issue. More examples are showcased in Fig. S9.
Appendix E Additional Experiments
In this section, we supplement the details of the experiments. First in Sec. E.1, we list all the hyperparameters through the table. Then We give more qualitative results about our experiments. Sec. E.2 shows the text reference tasks like 3D object point caption and QA and Sec. E.3 shows our point diffusion results.
We detail the hyperparameters of GPT4Point, largely mirroring those used in BLIP-2 during the pretrain stage. These parameters are maintained for Stage1: Point-text alignment and the LLM branch in Stage2. Tab. S2 lists them. The parameters for the LLM branch in Stage2 are almost identical to those of Stage1, except for the warm-up iterations, which changed from 5K to 2K. For BLIP-2, after pretraining on multiple datasets, fine-tuning is performed on a smaller dataset and subtasks. Additionally, different image backbones were used in the pretraining and fine-tuning phases. But in our GPT4Point, we only use the pretrain stage in the BLIP-2 and all tasks are evaluated by zero-shot. For the diffusion branch, we need to make the learning rate very small because here we only train the fully connected layers. The init, min and the warmup learing rate is 1e-7, 0 and 1e-8, and we only train 1 epoch.
E.2 Point-text Captions and QA Demos
In this section, we show more point-text qualitative results of GPT4Point. More specific examples are presented in Fig. S10. We can see that GPT4Point is capable of effectively understanding point clouds and can engage in fluent conversations with humans.
E.3 Point Diffusion Results
Fig. S5 shows more qualitative results of point diffusion results of GPT4Point. We find that GPT4Point can guide text-to-3D processes, generating results with more accurate colors and geometric shapes.