3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment

Ziyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng, Siyuan Huang, Qing Li

Introduction

Aligning the 3D physical world with natural language is a crucial step towards embodied artificial intelligence , where intelligent agents can understand and further execute human instructions in the real world . Recently, 3D vision-language (3D-VL) tasks have attracted growing interest , including 3D visual grounding , dense captioning , grammar learning , question answering , and situated reasoning .

However, most of the models developed for 3D-VL only focus on one or two of these 3D-VL tasks and employ task-specific designs . For instance, 3D-SPS and BUTD-DETR progressively discover the target object by attending VL features and detecting objects in each layer. 3DVG , MVT , and ViL3DRel improve 3D visual grounding by explicitly infusing spatial relation information into the model design. 3DJCG jointly learns 3D dense captioning and visual grounding via a shared 3D object proposal module with two separate task-specific heads . Additionally, training these models often requires manually specified auxiliary losses (e.g., 3D object detection/classification and text classification ) or optimization tricks (e.g., knowledge distillation ). The lack of a simple and unified approach creates a significant gap in developing a general-purpose 3D-VL model.

To fill such gap, we introduce 3D-VisTA, a Transformer-based model for 3D Vision and Text Alignment that can be easily adapted to various downstream tasks. Unlike previous models that design sophisticated task-specific modules, we simply utilize a vanilla self-attention transformer for both single-modal modeling and multi-modal fusion in the 3D-VisTA. As a general approach to further enhance 3D spatial comprehension , we explicitly encode the pairwise spatial relations between objects into the self-attention weights for 3D object modeling.

Inspired by the success of large-scale pre-training in NLP , CV , and 2D-VL , we propose to pre-train 3D-VisTA on 3D scene-text data, aiming for better performances on 3D-VL tasks. To this end, we construct ScanScribe, the first large-scale 3D scene-text pairs dataset for 3D-VL pre-training. We first collect RGB-D scans of indoor scenes from ScanNet and 3R-Scan datasets. We also randomly replace some objects in the scene with objects from the Objaverse 3D object database based on their categories, in order to increase object diversity. To obtain the text, we transform the text from existing datasets based on ScanNet into scene descriptions, including the question-answer pairs from ScanQA and the referring expressions from ScanRefer and ReferIt3D . We further leverage the scene graph annotations of scans from 3R-Scan, and adopt both templates and GPT-3 to generate scene descriptions from their scene graphs. In total, ScanScribe contains 278K 3D scene-text pairs for 2,995 RGB-D scans of 1,185 indoor scenes, with 56.1K unique object instances.

We pre-train 3D-VisTA on the proposed ScanScribe dataset. Our pre-training tasks include masked language modeling, masked object modeling, and scene-text matching. Notably, similar objectives are widely adopted in 2D-VL yet rarely explored in the 3D-VL domain. The proposed pre-training procedure effectively learns the alignment between 3D point clouds and texts, which eliminates the need for auxiliary losses and optimization tricks in downstream task fine-tuning. On six challenging 3D-VL tasks, ranging from visual grounding (i.e., ScanRefer , Nr3D/Sr3D ) and dense captioning (i.e., Scan2Cap ) to question answering (i.e., ScanQA ) and situated reasoning (i.e., SQA3D ), fine-tuned 3D-VisTA raises the SOTA results on ScanRefer by 8.1% (acc@0.5), on Sr3D by 3.6%, on Scan2Cap by 10.1%(C@0.25), on ScanQA by 3.5%/2.1% (EM@1), and on SQA3D by 1.9%. Moreover, 3D-VisTA demonstrates superior data efficiency, obtaining strong results with only 30% of the annotations for these downstream tasks.

Our main contributions can be summarized as follows:

We propose 3D-VisTA, a simple and unified Transformer for aligning 3D vision and text. The proposed Transformer simply utilizes the self-attention mechanism, without any complex task-specific design.

We construct ScanScribe, a large-scale 3D-VL pre-training dataset that contains 278K 3D scene-text pairs for 2,995 RGB-D scans of 1,185 unique indoor scenes.

We introduce a self-supervised pre-training scheme for 3D-VL, with masked language/object modeling and scene-text matching. It effectively learns the 3D point cloud and text alignment and further simplifies and improves downstream task fine-tuning.

We fine-tune 3D-VisTA and achieve state-of-the-art performances on various 3D-VL tasks, ranging from visual grounding and dense captioning to question answering and situated reasoning. 3D-VisTA also demonstrates superior data efficiency, obtaining strong results even with limited annotations.

Related Work

3D Vision-language Learning. Recently, there has been growing interest in 3D vision-language (3D-VL) learning. Unlike traditional scene understanding, 3D-VL tasks connect the physical world to natural language, which is crucial for achieving embodied intelligence . In this emerging area, Chen et al. and Achlioptas et al. concurrently introduce ScanRefer and ReferIt3D datasets for benchmarking natural language grounding to 3D object properties and relations. Besides 3D visual grounding, Azuma et al. develop a 3D question-answering dataset named ScanQA that requires a model to answer a question about objects and their relations given a 3D scene. More recently, Ma et al. propose a situated reasoning task called SQA3D for embodied scene understanding in 3D scenes.

Several models have been proposed for these benchmarks . Notably, 3D-SPS and BUTD-DETR progressively discover the target object by leveraging cross attention mechanism and language guidance. 3DVG , MVT , and ViL3DRel tackle 3D visual grounding by explicitly infusing spatial relation information into their models. Although these works have achieved impressive results in bridging 3D vision and language, they still rely heavily on task-specific knowledge in model design and sophisticated optimization techniques . In contrast, the proposed 3D-VisTA unifies visual grounding, question-answering, and situated reasoning through a simple Transformer-based architecture. Training 3D-VisTA is also straightforward, without requiring any auxiliary losses or sophisticated optimization techniques. Refer to Table 2 for a detailed comparison between 3D-VisTA and other 3D-VL models w.r.t. task, auxiliary Loss, and architecture.

Large-scale Pre-training. In recent years, large-scale pre-training has become a cornerstone of natural language processing (NLP), computer vision (CV), and 2D vision-and-language (2D-VL) domains. The introduction of the transformer-based architecture , especially BERT and GPT , has led to significant improvements in various NLP tasks. The success of these models has led to the development of more advanced pre-training techniques such as XLNet and RoBERTa . These models have achieved state-of-the-art performance on a wide range of NLP tasks, including text classification, question answering, and language generation. The most successful pre-training approach in CV is the ImageNet pre-training, which has been used as a starting point for a wide range of downstream tasks such as object detection and image segmentation. Recently, the introduction of transformer-based models such as ViT and Swin Transformer has led to significant improvements in various CV tasks. The field of 2D-VL has also seen significant advancements due to pre-training techniques. In particular, the introduction of the ViLBERT and LXMERT models has led to state-of-the-art performance on tasks such as visual question answering and image captioning. More recently, the development of CLIP , ALIGN , and Flamingo has shown that large-scale pre-training on image-text pairs leads to better cross-modal understanding and the emerge of in-context learning in a zero-shot or few-shot manner.

Although large-scale pre-training has become a crucial technique in NLP, CV, and 2D-VL, it has rarely been explored in 3D-VL. explore multi-task learning of visual grounding and dense captioning, and then further fine-tune their models on each task. The exploration of 3D-VL pre-training may be hindered by the lack of a large-scale pre-training dataset. Therefore, we construct ScanScribe, the first large-scale 3D scene-text pairs dataset for 3D-VL pre-training. As shown in Table 2, ScanScribe is much larger than existing 3D-VL datasets and also has more diverse text. Pre-training 3D-VisTA on ScanScribe has led to significant improvements on 3D-VL tasks, so we believe ScanScribe can fuel the exploration of 3D-VL pre-training in the future.

3D-VisTA

In this section, we introduce 3D-VisTA, a simple and unified Transformer for aligning 3D scenes and text. As illustrated by Fig. 2, 3D-VisTA takes a pair of scene point cloud and sentence as input. It first encodes the sentence via a text encoding module and processes the point cloud via a scene encoding module. Then the text and 3D object tokens are fused by a multi-modal fusion module to capture the correspondence between 3D objects and text. 3D-VisTA is pre-trained using self-supervised learning and can be easily fine-tuned to various downstream tasks. Next, we describe each module in detail.

We adopt a four-layer Transformer to encode the sentence SS into a sequence of text tokens {wcls,w1,w2,⋅⋅⋅,wM}\{w_{\text{cls}},w_{1},w_{2},\cdot\cdot\cdot,w_{M}\}, where wclsw_{cls} is a special classification token ([CLS]) and MM is the sentence length. This text encoding module is initialized by the first four layers of a pre-trained BERT .

2 Scene Encoding

Given the point cloud of a 3D scene, we first use segmentation masks to break down the scene into a bag of objects. The segmentation masks can be either obtained from ground truth or instance segmentation models . For each object, we sample 1024 points and normalize their coordinates into a unit ball. Then the object point cloud is fed into PointNet++ to obtain its point features and semantic class. We compose the point features fif_{i}, the semantic class embedding cic_{i}, and the location lil_{i} (i.e., 3D position, length, width, height) as the representation of the object token ii:

where WcW_{c} and WlW_{l} are additional projection matrices to map cic_{i} and lil_{i} into the same dimension as fif_{i}.

To further provide a contextual representation of objects, we capture the object-to-object interactions by infusing object tokens into a four-layer Transformer. Motivated by previous works , we explicitly encode the pairwise spatial relations of objects into the Transformer (Spatial transformer in Fig. 2). More specifically, we follow to define the pairwise spatial features for the object pair i,ji,j:

3 Multi-modal Fusion

We simply concatenate the text and the 3D object tokens and send them to a LL-layer Transformer (Unified transformer in Fig. 2) for multi-modal fusion. Learnable type embeddings are added to the tokens to differentiate text and 3D objects. We denote the output of the multi-modal fusion module as {wcls,w1:M,o1:N}\{\textbf{w}_{\text{cls}},\textbf{w}_{1:M},\textbf{o}_{1:N}\} for [CLS], text tokens, and 3D object tokens, respectively.

4 Self-supervised Pre-training

To learn the 3D scene and text alignment in a self-supervised manner, we pre-train 3D-VisTA on 3D scene-text pairs via the following proxy tasks:

Masked Language Modeling (MLM). We follow the BERT pre-training to perform MLM: (1) 15% of the text tokens are randomly chosen; (2) 80% of the time: replace these tokens with [MASK]; (2) 10% of the time: replace these tokens with some random text tokens; (3) 10% of the time: these tokens remain unchanged. The model is trained to predict the masked text tokens given the remaining text and 3D object tokens:

Masked Object Modeling (MOM). Similar to MLM, we mask out 10% of 3D object tokens. However, we mask a 3D object token by only replacing its point features and semantic embedding (i.e., “fi+Wccif_{i}+W_{c}c_{i}” in Eq. 1) with a learnable mask embedding but keep its positional information (i.e., “WlliW_{l}l_{i}” in Eq. 1) unchanged. The model is trained to utilize the position clue of the masked object to predict its semantic class cc given the remaining 3D objects and text:

Scene-Text Matching (STM). While masked language and object modeling enable local text-object alignment in a fine-grained granularity, we also perform scene-text matching to enhance the global fusion of scene and text, which we find very beneficial for downstream question-answering tasks. More specifically, we extract the output corresponds to [CLS] as the global representation of the input scene-text pair, and feed it into a two-layer MLP to predict if the scene and the text are matched:

In practice, 30% of the samples in a training batch are negative pairs, created by replacing the scene point cloud or text with a randomly selected sample.

Final loss. Our final pre-training objective is obtained by simply adding the losses of the proxy tasks above:

Notably, the proposed pre-training scheme is self-supervised and task-agnostic, unlike the supervised multi-task learning used in previous work that requires task supervision.

5 Downstream Task Finetuning

The pre-trained 3D-VisTA can be easily adapted to various 3D-VL tasks by adding lightweight task heads. More specifically, we fine-tune 3D-VisTA on the following tasks:

3D Visual Grounding tasks a model to locate a target object in a 3D scene from a referring expression. To find the referred object, we apply a two-layer MLP to each object token oi\textbf{o}_{i}, and obtain the probability of the object being referred to. The model is fine-tuned using the cross-entropy loss.

3D Dense Captioning is introduced by to test a model’s ability of detecting and describing objects in a 3D scene. Following , we take w1:M\mathbf{w}_{1:M} and predict text tokens autoregressively to generate a sentence. The model is fine-tuned using cross-entropy loss.

3D Question Answering requires a model to answer an object-related question given a 3D scene. Following , we feed the text tokens w1:M\textbf{w}_{1:M} and the object tokens o1:N\textbf{o}_{1:N} into a modular co-attention network (MCAN) to produce answers. The model is fine-tuned using the QA loss and the object localization loss.

3D Situated Reasoning is recently proposed by to benchmark the 3D scene understanding of embodied agents. To adapt 3D-VisTA to this task, we concatenate the situation description and the question into a single input sentence. The answer classification is similar to the 3D question answering task. The model is fine-tuned using the answer loss.

In general, we find adapting 3D-VisTA to these downstream tasks much simpler than previous methods , as 3D-VisTA is simply fine-tuned using the task loss only, without the need for any auxiliary losses (e.g., sentence/object classification loss ) or optimization tricks (e.g., multi-view aggregation and knowledge distillation ). This makes 3D-VisTA a more unified and general-purpose 3D-VL model.

ScanScribe

In recent years, large-scale pre-training has been widely used to improve the performance on downstream tasks in CV , NLP , and 2D-VL . However, large-scale pre-training has barely been touched in the 3D-VL domain, possibly due to the lack of pre-training datasets for 3D-VL. To facilitate the exploration of 3D-VL pre-training, we build a large-scale 3D scene-text pairs dataset, named ScanScribe. As illustrated in Table 3, the construction of 3D scene-text pairs in ScanScribe comprises two parts:

3D scenes. We collect RGB-D scans of indoor scenes from ScanNet and 3R-Scan . To increase the diversity of 3D objects in these scenes, 10% of the object instances in each scene are randomly replaced by objects from the Objaverse 3D object database based on their categories. For each ScanNet and 3R-Scan object category, we download about 40 object instances from Objaverse as candidate object replacements. As a result, we collect 2,995 RGB-D scans of 1,185 indoor scenes, with 56.1K unique object instances.

Text. For the scans from ScanNet, we transform the text from existing datasets based on ScanNet into scene descriptions, including the question-answer pairs from ScanQA and the referring expressions from ScanRefer and ReferIt3D . For the scans from 3R-Scan, we adopt both templates and GPT-3 to generate scene descriptions based on their scene graph annotations . Specifically, for each object, we first extract all the ⟨object, relation, neighbor⟩\langle\text{{object}, {relation}, {neighbor}}\rangle triplets from the scene graph. We then use the template “This is a object, a neighbor is relation to object” to generate the descriptions. Note that we only choose objects with fewer than 7 neighbors in a template-based generation. We further explore using GPT-3 to generate the descriptions with the following prompt “object is relation to neighbor …(repeat until all the neighbors have been used). Where is object? or Summarize the scene.” Ultimately, 278K scene descriptions are generated for the collected 3D scenes.

Experiments

Implementation Details. The pre-training runs for 30 epochs with a batch size of 128. We use the AdamW optimizer with β1=0.9,β2=0.98\beta_{1}=0.9,\beta_{2}=0.98. The learning rate is set to 1e−41e^{-4}, with a warmup of 3,000 steps, and cosine decay. During pre-training, we use ground-truth segmentation masks to generate object-level point clouds.During fine-tuning, we use ground-truth masks or Mask3d , which depends on the task setting. On the ScanRefer dataset, we also incorporate PointGroup for comparison with previous approaches. In ablation studies, we use ground-truth masks in all tasks for simplicity. Both pre-training and fine-tuning are conducted on a single NVIDIA A100 80GB GPU.

3D Visual Grounding. We evaluate our model on three datasets for this task: ScanRefer , Nr3D, and Sr3D . For Nr3D/Sr3D, we follow ReferIt3D to use ground-truth object masks and report the results as the grounding accuracy, i.e., whether the model correctly selects the referred object among ground-truth object proposals. For ScanRefer, we follow to use detector-generated object proposals and report the results as Acc@k(k∈{0.25,0.5})k(k\in\{0.25,0.5\}), i.e., the fraction of referring queries whose predicted box overlaps the ground truth with IoU >k>k.

3D Dense Captioning We evaluate our model on the Scan2cap dataset and report the text similarity metrics under different box overlap ratios.

3D Question Answering. We evaluate our model on the ScanQA dataset and use exact matches (EM@1 and EM@10) as the evaluation metric. We also report several sentence evaluation metrics, including BLEU-4, ROUGE, METEOR, and CIDEr. Both test sets (w/ or w/o objects) of ScanQA are used in our evaluation.

3D Situated Reasoning We evaluate our model on the SQA3D dataset and report the answer accuracy under different types of questions as the evaluation metric.

2 Downstream Task Results

In this section, we discuss the experimental results of the downstream tasks and compare the proposed 3D-VisTA model with the state-of-the-art (SOTA) methods. Results are presented in Tables 4, 5, 6, 7, 8 and 3 and the main observations from these results are as follows:

Even trained from scratch, 3D-VisTA achieves competitive performances with SOTA methods. Specifically, 3D-VisTA (scratch) obtains an overall accuracy of 57.5% and 69.6% on Nr3D and Sr3D, which outperforms most previous models; it gets an EM@1 accuracy of 25.2% on ScanQA, which is 1.7% higher than SOTA. Of note, 3D-VisTA is trained on these datasets simply using the task losses, without any auxiliary losses or optimization tricks, indicating that 3D-VisTA is a very simple yet effective architecture for 3D-VL tasks.

Pre-training on ScanScribe significantly improves the performance of 3D-VisTA. Overall, the pre-training improves the accuracy on Nr3D/Sr3D by 6.7%/6.8%, the acc@0.25/0.5 on ScanRefer by 4.7%/4.3%, the EM@1 on ScanQA by 1.8%/2.6%, the C@0.25 on Scan2Cap by 4.2%, and the average accuracy on SQA3D by 1.8%. These large improvements consolidate the efficacy of ScanScribe for the 3D-VL pre-training.

The pre-trained 3D-VisTA outperforms SOTA by a large margin. 3D-VisTA outperforms ViL3DRel on Sr3D by 3.6% and on ScanRefer by 2.7%/8.1% (acc@0.25/0.5), beats ScanQA by 3.5%/2.1 (EM@1), Scan2Cap SOTA by 10.1%/19.2% (C@0.25/0.5), SQA3D by 1.9% (Avg.). 3D-VisTA sets a new record for these 3D-VL tasks and may inspire future research on 3D-VL pre-training.

Finetuning 3D-VisTA on downstream tasks with limited annotations achieves strong results. As shown in Fig. 3, being fine-tuned using 30% and 40% of the annotations on ScanRefer and ScanQA, the pre-trained 3D-VisTA can achieve better performance than the one trained from scratch with full data. We hypothesize that 3D-VisTA has successfully captured the alignment between 3D objects and text via pre-training and is thus able to readily adapt to downstream tasks of various formats. It also reveals the potential of 3D-VisTA to learn unseen tasks in a zero-shot or few-shot manner, which has emerged in NLP and 2D-VL via large-scale pre-training.

3 Ablation Studies

In this section, we conduct ablation studies to analyze the impact of several important hyperparameters, including Transformer depth, pre-training objectives, and data amount.

Transformer Depth. Since the model size is a key factor in the pre-training of NLP and 2D-VL, we study the effect of the transformer depth by varying the number of layers in the multimodal fusion module. As shown in Table 9(a), using 4 layers achieves the best performance and simply adding more layers does not help. This observation is somewhat contradictory to the ones from NLP and 2D-VL. It points out that although ScanScribe is much larger than existing 3D-VL datasets, it is still far from enough to unleash the full potential of pre-training in the 3D-VL domain.

Pre-training Objectives. Table 9(b) presents the ablation study for the pre-training objectives. The MLM objective alone slightly benefits question answering (QA), but brings a negative impact on visual grounding (VG). Adding MOM and STM boosts the performance of both QA and VG, which highlights the importance of MOM and STM for aligning 3D vision and text. Overall, using all three objectives together leads to the best performance for both tasks, with STM and MOM providing the greatest improvements in accuracy.

Pre-training Data. Table 9(c) presents the results using various configurations of pre-training data. We can see that simply using the ScanNet data for pre-training, which is from the same domain as downstream tasks, leads to a significant improvement in VG and QA. This validates the effectiveness of pre-training, even in the case of no additional 3D data than downstream tasks. Adding 3R-Scan and Objaverse increases the amount and the diversity of 3D data, which further boosts the accuracy of both VG and QA. Overall, the best performance for both tasks is achieved when all three data sources are used. This points out a promising path for improving 3D-VL tasks — collecting more data for pre-training.

4 Qualitative Studies and Additional Results

In this section, we perform additional studies to better understand how pre-training helps. As shown in Fig. 4, pre-training improves the spatial understanding of 3D-VisTA for visual grounding, so it can better align with human prior viewpoint and reason over spatial relations. This is very helpful when the model needs to distinguish the target object from multiple instances of the same class. Pre-training also helps with a better understanding of visual concepts like colors and shapes, and situations for question answering and situated reasoning. Besides, pre-training enhances the capability of aligning long text with 3D scenes, as evidenced by the larger improvement over longer queries in Fig. 5.

Conclusion

This paper proposes 3D-VisTA, a simple yet effective architecture for 3D-VL tasks. The model simply uses self-attention layers and can be easily adapted to various downstream tasks, without requiring any auxiliary loss or optimization trick. We also introduce ScanScribe, the first large-scale 3D scene-text pairs dataset for 3D-VL pre-training. The pre-trained 3D-VisTA achieves state-of-the-art results on a variety of 3D-VL tasks with superior data efficiency, paving the path to future foundation models for 3D-VL tasks.

Future Works. Currently, 3D-VisTA uses an offline 3D object detection module, which may be a bottleneck for further improvement. Jointly optimizing the object detection module in the pre-training phase is an interesting future direction. Besides, the data amount in ScanScribe is still insufficient for large-scale 3D-VL pre-training, so scaling up the pre-training dataset as well as the model size is a promising direction to further improve the 3D-VL learning.

Acknowledgements. The authors would like to thank Hongming Xu at BIGAI for the help on Mask3D. This work is supported in part by the National Key R&D Program of China (2022ZD0114900) and the National Science Foundation of China (NSFC) under Grant No. 62176134.

References

Appendix A Implementation Details

ScanRefer : The ScanRefer dataset contains 51,583 sentences written by humans to describe 800 scenes in ScanNet. We used the official split and allocated 36,665 and 9,508 samples for training and validation, respectively. The dataset is categorized into unique and multiple subsets based on whether the target object is a unique class in the scene. In this task, we need to find the target object described by a sentence. The evaluation metric for this task is accuracy under intersection over union (IoU) 0.25 and 0.5.

Nr3D/Sr3D : The Sr3D dataset comprises of 83,572 utterances that are automatically generated using a template that focuses on the target-anchor spatial relationship. The Nr3D contains 45,503 human utterances. Both Sr3D and Nr3D are split by “Easy”/“Hard” and “ViewDep”/“ViewIndep”. Hard samples are the ones with two or more distractors in a scene. The view-dependent samples contain language descriptions that rely on viewing directions. These two datasets are also used for visual grounding like ScanRefer. But grounding accuracy with ground truth object proposal is evaluated in this setting.

ScanQA : ScanQA is a dataset for 3D question answering with 41,363 questions and 58,191 answers. Different from 2D QA, ScanQA focuses more on spatial relations. We follow to use exact matches EM@1 and EM@10 as the evaluation metric. EM@K means the percentage of top K answers from the model matches one of the ground-truth answers. Also, we include text similarity metrics to evaluate answers, including BLEU-4, ROUGE, METEOR, and CIDEr.

SQA3D : SQA3D is a benchmark for scene understanding of embodied agents with 6.8k unique situations, 20.4k descriptions, and 33.4k diverse reasoning questions. Given a situation, an embodied agent must understand embodied activities, navigation instructions, and common sense, and perform multi-hop reasoning. The evaluation metric is answer accuracy under different types of questions.

Scan2Cap : Scan2Cap is a dataset for 3D dense captioning. Object descriptions are produced from ScanRefer dataset. For each sentence, two special tokens including [SOS] and [EOS] are added.

A.2 Model Architecture

For the scene encoder, we use a three-layer Pointnet++ with radius 0.2, 0.4, and sample all points to aggregate a 768-dimension feature. For all text and object tokens, the dimension is 768 in the following multi-modal fusion layers. In the unified encoder, the number of attention heads is set to 12 and the dimension of feedforward layers is set to 2048. For the visual grounding head, we use a two-layer MLP with a hidden dimension of 384. For the question-answering head and the situated reasoning head, we use a two-layer MLP with input dimensions 512 (from the attention flat layer) and 768.

A.3 Training settings

The settings of pre-training including mask ratio, and optimization hyperparameters are introduced in the main paper. We exclude the ScanNet validation and test scenes from pre-training to ensure a fair comparison with other methods. All scenes from 3R-Scan are used for pre-training. In this part, we elaborate on the fine-tuning details.

3D Visual Grounding: We only use a cross-entropy loss for fine-tuning 3D-VisTA on ScanRefer, Nr3D, and Sr3D. For all these grounding tasks, we set the batch size to 64, and the learning rate to 1e-4, We multiply the learning rate of the text encoder by 0.1 to stabilize the training process. We fine-tune the pre-trained 3D-VisTA for 100, 100, and 50 epochs for ScanRefer, Nr3D, and Sr3D, respectively. AdamW with β1=0.9,β2=0.98\beta_{1}=0.9,\beta_{2}=0.98 is chosen as the optimizer. We use a warmup of 5,000 steps and a cosine annealing learning rate schedule.

3D Question Answering: We use a cross-entropy answer classification loss and a visual grounding loss for ScanQA. The batch size is 64 and the learning rate is 1e-4. 3D-VisTA is fine-tuned for 30 epochs with 2000 warmup steps for this task. Other optimization parameters are the same as the visual grounding task.

3D Situated Reasoning: Answer classification loss is used for SQA3D. We fine-tune 3D-VisTA for 50 epochs. Other optimization parameters are the same as the 3D question-answering task.

3D Dense Captioning: Cross entropy loss is used for fine-tuning Scan2Cap. We use the BERT tokenizer to process input sentences and use the casual mask for language transformer. During both fine-tuning and inference, object tokens are not allowed to attend text tokens because of information leaks. 3D-VisTA is fine-tuned for 100 epochs with batch size 64 and learning rate 1e-4. During inference, text tokens are generated by the greedy selection policy.

A.4 ScanScribe

In the main paper, we introduce our method of generating new scene-text pairs from scene graphs and large language models. More examples and cases are provided in this section. We support 40 relations and the mapping of relations to descriptions for the template-based generation is shown in Table A1.

With these relations, we can use templates like “This is a object, a neighbor is relation to object” and utilize GPT-3 to increase text diversity. During pre-training, to balance the proportion of template and GPT-3 generated texts in the 3R-Scan dataset, we duplicate texts from GPT-3 to 15 times for pre-training. Examples from both template-based generation and GPT-3 are presented in Fig. A1. We can observe that given entities and relations in a scene, GPT-3 can summarize them into a fluent and natural sentence.

Appendix B Additional Results

We provide ablation studies on the use the template-generated text and GPT-3-generated text. As shown in Table A2, GPT-3-generated text improves Sr3D and Nr3D by 1.0% and 1.5%, while having little impact on ScanRefer and ScanQA. More qualitative results including failure cases are provided in Fig. A2.