An Embodied Generalist Agent in 3D World

Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, Siyuan Huang

Introduction

Building one generalist model that can achieve comprehensive tasks like humans has been a long-existing pursuit in artificial intelligence and neuroscience (Lake et al., 2015; 2017; Zhu et al., 2020; Mountcastle, 1979; Schmidhuber, 2018; Huang et al., 2022a). Recent advances in large language models (LLMs) (Brown et al., 2020) and “foundation model” (Bommasani et al., 2021) emerge as a promising paradigm in building such generalist models in natural language processing (OpenAI, 2022; 2023), computer vision (Kirillov et al., 2023), and robotics (Brohan et al., 2022; 2023). The keys to the success of this paradigm lie in large-scale internet-level datasets from numerous tasks and domains and scalable Transformer architectures (Vaswani et al., 2017) that can absorb generalizable and task-agnostic knowledge from the data. Such efforts are further extended to multi-modal (Alayrac et al., 2022; Lu et al., 2023; Li et al., 2023c) and generalist models (Reed et al., 2022; Driess et al., 2023) where the agents can solve versatile tasks based on the language-specified task descriptions and show certain generalizability to novel situations. Nonetheless, their abilities are primarily demonstrated within 2D domains, thereby limiting the comprehension of the 3D physical environment that envelops humans and other intelligent species. This limitation acts as an obstacle, preventing current models from successfully executing real-world tasks and the attainment of general intelligence. Therefore, we ask a fundamental question: how to equip the generalist agent with a comprehensive understanding of and the ability to interact with the real 3D world?

The development of such generalist agents encounters three primary challenges: the creation of suitable datasets, the design of unified models, and the design of effective learning strategies. Despite substantial progresses in scaling up image-text models (Tsimpoukelli et al., 2021; Alayrac et al., 2022) and the curation of corresponding datasets (Radford et al., 2021; Schuhmann et al., 2022), advancements in 3D scene-level understanding have significantly lagged behind. This is largely attributed to the limited scale and manual labeling of 3D datasets (Dai et al., 2017; Wald et al., 2019; Chen et al., 2020), given the higher cost associated with collecting 3D data compared to 2D data. Furthermore, previous models have often been designed with strong priors (Zhao et al., 2021; Chen et al., 2022), with limited exploration of large-scale unified pretraining and efficient fine-tuning based on LLMs. Notably, recent works (Zhu et al., 2023c; Hong et al., 2023) utilize the unified Transformers or LLMs to enhance the model’s capability in grounded 3D scene understanding. However, they still lack the agency to act within 3D environments and efforts in unleashing LLMs for 3D vision-language-action (VLA) learning. How to equip the 3D agent with a simple unified architecture and effective learning strategy to establish VLA capability remains rarely explored.

In this work, we introduce the generalist agent LEO, which is generically embodied, multi-modal, and multi-task. It can take egocentric 2D images, 3D point clouds, and texts as task input and achieve comprehensive tasks within the 3D environment. As shown in Fig. 1, LEO exhibits the capability of perceiving, grounding, reasoning, planning, and acting with shared model architectures and weights. LEO perceives through an egocentric 2D image encoder for the embodied view and an object-centric 3D point cloud encoder for the third-person global view. The 3D encoder generates an object-centric token for each observed entity. Such encoder design can be flexibly adapted to tasks with various embodiments. These output tokens are then interleaved with text tokens to form a scene-grounded instructional task sequence, which further serves as the input to a decoder-only LLM. All the tasks are re-formulated as sequence prediction problems. Therefore, LEO can be trained with task-agnostic inputs and outputs using autoregressive training objectives. To accommodate embodied tasks that predict action tokens, we employ a pool of special tokens to represent actions and replace the least used tokens (Brohan et al., 2023) in LLM. We perform LoRA (Hu et al., 2022) to fine-tune and adapt the innate knowledge of LLM to the multi-modal generalist. LEO demonstrates two essential capabilities: 3D vision-language grounding and 3D VLA. They are injected with two training stages: 3D vision-language alignment and VLA instruction-tuning. To facilitate the training, we curate a large-scale dataset with object-level and scene-level tasks by fusing existing datasets with high-quality data prompted from the LLMs. We propose scene-graph-based prompting and refinement methods, along with Object-centric Chain-of-Thought (O-CoT) for improving the quality of generated data, largely enriching the data scale and diversity, and further eliminating the hallucination of LLMs.

We quantitatively evaluate and ablate LEO on diverse 3D tasks, including object-level and scene-level captioning (Luo et al., 2023; Chen et al., 2021), 3D question answering (Azuma et al., 2022), situated question answering (Ma et al., 2023), embodied navigation (Ramrakhya et al., 2022), and robotic manipulation (Shridhar et al., 2021). The results indicate (i) LEO achieves state-of-the-art results on most tasks; (ii) through task-agnostic instruction tuning with a unified model, LEO outperforms most previous task-specific models on various domains; (iii) the pretraining of 3D vision-language alignment greatly elevates the performance of VLA instruction-tuning; (iv) scaling up the training data boosts the performance of generalist agent, similar to scaling laws (Kaplan et al., 2020; OpenAI, 2023) in LLMs and generalist agent (Reed et al., 2022). We also show qualitative results of chatting with LEO over diverse tasks, demonstrating its capability in scene-aware planning and dialogue.

In summary, our main contributions are: (i) we propose LEO, the first generalist agent with the capability to perceive, ground, reason, plan, and act in the 3D world; (ii) we demonstrate that a generalist agent can be built via fine-tuning the LLM with object-centric multi-modal representations and mixing the training data with embodied action sequences, enabling it to excel in embodied tasks; (iii) we meticulously curate a large-scale dataset to train such agent and propose several techniques aimed at enhancing the quality of prompted data from LLMs; (iv) we extensively evaluate LEO and demonstrate its proficiency on diverse tasks including embodied navigation and robotic manipulation, we also observe consistent performance gains while simply scaling up the training data; (v) we will release the data, code, and model weights to endow the future research of generalist agents.

Model

The leading design principles of LEO are two-fold: 1) It should handle the input of egocentric 2D, global 3D, and textual instruction, and the output of textual response as well as embodied action commands in a unified architecture; 2) It should leverage pretrained large language models (LLMs) as a powerful prior for the downstream tasks. We therefore convert all data of different modalities into a sequence of tokens, illustrated below:

With this representation, we formulate the learning of LEO as GPT-style autoregressive language modeling (Brown et al., 2020) given the prefix (from system message to instruction), i.e. prefix language modeling (Raffel et al., 2020). Therefore, a pretrained LLM can be used to process such sequences. In the following, we will detail the tokenization of multi-modal data, model architecture, training loss, and inference settings. An overview of our model can be found in Fig. 1.

We follow prior practices in 2D VLM (Liu et al., 2023b; Alayrac et al., 2022) and 3D VLM (Zhu et al., 2023c) to tokenize the multi-modal data in LEO. We use SentencePiece tokenizer (Kudo & Richardson, 2018) to encode text with 32k subwords; 2D image tokens for egocentric 2D images; and object-centric 3D tokens extracted over Mask3D-based (Schult et al., 2022) object proposals for 3D point cloud inputs. For embodied action commands, continuous actions (e.g. in manipulation) are discretized (details in Appendix B) to join the discrete actions (e.g. navigation) and form a unified discrete action space. We follow Brohan et al. (2023) to map these discrete actions to the least used tokens in SentencePiece. After tokenization, all tokens are ordered into the format in (1).

2 Token embedding & LLM

We apply several token embedding functions to process the tokens in the sequence before sending them to the LLM. The LLM will then align these tokens of different modalities, and produce the response. Most of the responses are text and can be decoded directly. For responses that include embodied actions, we will map the reserved SentencePiece text tokens back to action commands.

For text tokens (including embodied actions that have been mapped to the reserved text tokens), an embedding look-up table is used to map them into vectors. While the egocentric 2D image is encoded by a pretrained OpenCLIP ConvNext (Liu et al., 2022) for obtaining image token embeddings. We apply MLP adapters to match the dimensions of all token embeddings.

Each 3D object token (i.e., the point cloud of a 3D object) is first encoded by a pretrained point cloud encoder (e.g., PointNet++ (Qi et al., 2017)). We then adopt the Spatial Transformer introduced in Chen et al. (2022) to further process the point cloud embedding of all objects into object-centric 3D token embeddings. In a nutshell, Spatial Transformer biases the standard attention score with relative position and size for capturing 3D relations between objects. Due to space limit, the readers are referred to Chen et al. (2022) and Sec. D.2 for more details.

We choose Vicuna-7B (Chiang et al., 2023) to process the token sequence. In order to tackle the challenging alignment and grounding problem of multi-modal tokens (2D, 3D, text, embodied action) while preserving the LLM pretrained knowledge, we employ LoRA (Hu et al., 2022) to introduce additional tunable parameters to the frozen pretrained LLM.

3 Training & Inference

We formulate the learning objective of LEO following (Brown et al., 2020; Raffel et al., 2020) in a prefix language modeling fashion. For a batch B\mathcal{B} of token sequence ss, we optimize LEO via:

where sprefixs_{\text{prefix}} denotes the prefix token (from system message to instruction) in (1). During training, we freeze the pretrained 3D point cloud encoder and the LLM and finetune the 2D image encoder, the Spatial Transformer and the LoRA parameters. In total, LEO has ~7B parameters and ~142M of them will be tuned. During inference, we use beam search to generate textual responses. For tasks that require action commands, we map the textual outputs to action commands as discussed in Sec. 2.1. More details on the model and training can be found in Appendix D.

Datasets

Since LEO is a generalist agent that receives multi-modal inputs and follows instructions, we adopt the two-stage training proposed in Liu et al. (2023b) and categorize the data into two sets: (i) LEO-align that focuses on 3D vision-language alignment at object-level and scene-level, to bridge the gap between 3D scene representations and natural language; and (ii) LEO-instruct that targets at 3D VLA instruction tuning for endowing LEO with the generalist capability of accomplishing myriad tasks in the 3D world including perceiving, reasoning, and acting. We provide the statistics of the two separate splits of data in Tab. 1. Examples of both datasets can be found in Appendix C.

In LEO-align, we focus on 3D vision-language alignment. We follow the alignment method proposed by BLIP-2 (Li et al., 2023d) and train the model to follow the instructions for captioning given 3D input. As a result, we consider collecting the following types of 3D caption data:

To facilitate object-level grounding of detailed object attributes, we leverage Cap3D (Luo et al., 2023), which contains language descriptions for objects in Objaverse (Deitke et al., 2023). Given a single 3D object as input, LEO will be asked to predict its caption.

For a better understanding of how an object can be related to others (spatial relations, etc.) when situated in a 3D scene, we collect referring expressions of objects in scenes from existing datasets, including ScanScribe (Zhu et al., 2023c) and ReferIt3D (Achlioptas et al., 2020). Further, we generate additional object-referring expressions on 3RScan (Wald et al., 2019) scenes by prompting LLMs (details in Sec. A.1). During alignment, LEO needs to predict these referring expressions given the object-centric 3D input of the scene and the referred object.

Finally, we encourage LEO to capture scene-level descriptions of a 3D scene. These scene-level captions focus on global information depicting key objects in the scene as well as their attributes and functionalities, relations among multiple objects, and room types and styles. We leverage scene graph annotations (Wald et al., 2019) and prompt LLMs to produce a total of ~20K captions. To further increase caption diversity, we propose a subgraph sampling strategy to prevent LLMs from always attending to certain notable facets of the scene (details in Sec. A.5). Similar to previous settings, LEO needs to predict these captions given the corresponding 3D input.

2 LEO-instruct: Instruction Tuning for Tasks in the 3D world

After alignment, LEO will be tuned to follow instructions and accomplish various 3D VLA tasks. Below, we provide a comprehensive illustration of the data preparation process for these tasks and an overview of generated data in Fig. 2. We list the corresponding instructions in Appendix C.

The task is to produce a generic caption given 3D input. We adopt the Scan2Cap dataset (Chen et al., 2021), which is based on the ScanNet (Dai et al., 2017) 3D scenes and covers various levels (object-level and scene-level) and aspects (attributes, relations, etc.) of scene details.

The 3D-QA task is an extension of VQA (Antol et al., 2015) to 3D scenes with a focus on 3D knowledge, ranging from spatial relations to functionalities of objects. For this task, we first aggregate two existing 3D-QA datasets: ScanQA (Azuma et al., 2022) and SQA3D (Ma et al., 2023). To further generate questions concerning rich 3D knowledge, we prompt LLMs to generate ~35K QA pairs on 3RScanQA with our quality refinement techniques discussed in Sec. 3.3.

The goal of this task is to support natural conversations between LEO and users about a given 3D scene. This task necessitates coherence and continuity across multiple rounds of conversational interactions. We build such dialogues on 3RScan scenes by prompting LLMs with a variant of the Chain-of-Thought prompting method discussed in Sec. 3.3 to facilitate diverse dialogues about relevant and accurate details about the 3D scene. In total, ~11K dialogues are collected.

In this task, LEO is required to decompose high-level tasks into step-by-step low-level plans given 3D scenes. We expect LEO to generate feasible plans based on the current 3D scene and ground its inherent common sense knowledge about procedures to the scene configurations, including, objects, their attributes, relations, and functional characteristics, etc. By prompting LLMs, we end up collecting ~14K task-plan pairs on 3RScan scenes.

We follow imitation learning setting in Habitat-web (Ramrakhya et al., 2022) for the embodied navigation task. We choose ObjNav, where LEO needs to map navigation instructions (e.g. “find bed”), object-centric 3D input, and an egocentric 2D input into discrete habitat motor commands. For simplicity, we use shortest path navigation trials rather than human demonstrations for learning as they are less noisy and therefore easier to learn when provided with the 3D scene. In total, we generate ~60K navigation episodes out of the MP3D ObjNav training scenes (Savva et al., 2019) for this task.

We employ a subset of the manipulation tasks introduced in CLIPort (Shridhar et al., 2021). The input of this task includes instructions, egocentric 2D observations, and object-centric 3D information. As discussed in Sec. 2.1, we discretize the continuous action space of CLIPort into bins to unify the action decoding of navigation and manipulation (more details in Appendix B). We generate 100K demonstrations for each selected manipulation task.

3 LLM-assisted 3D-language Data Generation

As mentioned above, at the core of producing a large proportion of LEO-align and LEO-instruct is the assistance of LLMs. We now detail the key techniques of prompting LLMs (more specifically, ChatGPT) to generate 3D-text data. An overview can be found in Fig. 2.

We use 3D scene graph from 3DSSG (Wu et al., 2021) to provide scene contexts in prompts. Compared to recent efforts that utilize object boxes (Yin et al., 2023; Hong et al., 2023; Wang et al., 2023d), we observed that our method provides high-quality object attributes and spatial relation information among objects, allowing LLMs to capture and generate more accurate and relevant 3D details (see comparisons in Sec. A.6). To further improve data quality in open-ended generation and reduce the hallucination of LLMs (Bang et al., 2023), we propose the Object-centric chain of thought (O-CoT) prompting that requires the LLM to explicitly provide the label and ID of object candidates as thoughts during question and dialogue generation. We provide examples of O-CoT in Fig. 2 and comparative experiments to verify the effectiveness of O-CoT on improving answer reliability in Sec. A.2. We also utilize subgraph sampling to further enhance the diversity of 3D scene graph (see details in Sec. A.5).

We pass raw LLM-generated responses into several human-defined filtering procedures based on the 3D scene graph. Notably, negative responses (e.g., lacking necessary information to answer) will be removed; unnatural narratives will be rewritten. For generated text that involves logical reasoning (e.g., counting) or hallucination, we manually fix the wrong responses based on the information provided by the 3D scene graph. We provide details about these procedures in Sec. A.3 and statistics in Sec. A.4.

Capabilities and Analyses

We present a comprehensive demonstration of LEO’s capabilities by evaluating it on the full spectrum of embodied 3D tasks encompassing perceiving, grounding, reasoning, planning, and acting. We provide both quantitative comparisons between LEO and competitive task-specific baselines and qualitative visualizations (see in Fig. 3) to showcase the power of LEO as an embodied generalist agent. We provide additional experimental details about the model and implementation in Appendix D. We further ablate LEO with various data configurations and analyze the scaling law.

Understanding and reasoning about object attributes, object relations, and other facets of 3D scenes from an agent’s egocentric perspective is a fundamental capability of an embodied generalist agent in the 3D world. We investigate how well can LEO perform 3D VL understanding and embodied reasoning tasks, especially when being compared against task-specific models and existing generalist agents. Specifically, we consider three renowned 3D tasks: 3D captioning on Scan2Cap (Chen et al., 2021), 3D QA on ScanQA (Azuma et al., 2022), and 3D embodied reasoning on SQA3D (Ma et al., 2023). By prompting LEO to follow instructions for these tasks, we follow the standard evaluation metric to report conventional captioning scores (CIDEr, BLEU, METEOR, and ROUGE) and SentenceSim (Reimers & Gurevych, 2019) for open-ended VL generation, as well as exact-match accuracies for QA tasks. Following 3D-VisTA (Zhu et al., 2023c), we use object proposals from Mask3D (Schult et al., 2022) in our object-centric 3D encoder.

For quantitative comparisons, we include both task-specific approaches and generalist models: 1) state-of-the-art specialists in 3D dense captioning (Chen et al., 2021; Cai et al., 2022; Chen et al., 2023); 2) state-of-the-art specialists in 3D QA (Azuma et al., 2022; Ma et al., 2023); 3) task-specific fine-tuned generalist models like 3D-VisTA (Zhu et al., 2023c) and 3D-LLM (Hong et al., 2023). To the best of our knowledge, LEO is the first model that, in stark contrast to prior models, can handle the aforementioned VL tasks in a unified architecture without additional fine-tuning. This lends greater credence to LEO’s comparative superiority.

As shown in Tab. 2, LEO surpasses state-of-the-art task-specific and task-specific fine-tuned models significantly on both 3D dense captioning and 3D QA tasks. In contrast to the specialist models that utilize task-specific heads, we demonstrate our LLM-based approach not only affords the flexibility of generating open-ended responses but also can achieve excellent scores in terms of standard metrics. Notably, considering the difference between close-set classification (e.g., 3D-VisTA) and open-ended text generation, we refine the protocol of exact match (see details in Sec. H.1). On the other hand, compared with the complicated 3D feature aggregation in 3D-LLM, we suggest that object-centric 3D representation is a simple yet effective option to connect 3D scenes with LLM while harnessing the inherent knowledge of LLM.

2 Chatting and planning about a 3D scene

Upon the 3D VL understanding and reasoning, we anticipate LEO to support more sophisticated and grounded interaction with human users, i.e., responding to complex multi-round user instructions in the 3D world. To verify these capabilities, we choose two tasks: 3D dialogue and scene-aware task planning. We provide qualitative examples of unseen scenarios from the held-out test sets of LEO-instruct, highlighting LEO’s merits of instruction following and scene-grounded responses. We defer the quantitative results of dialogue and planning to our ablation study in Sec. 4.4. Quantitative comparison with other approaches is infeasible due to the lack of a common benchmark.

As shown in Fig. 3, LEO is capable of generating high-quality responses to complete the tasks of dialogue and planning. We highlight two features: 1) The responses of LEO are precisely grounded to the 3D scenes. In particular, the proposed plan by LEO contains concrete objects that are present in the scenes, as well as concrete actions regarding these objects. 2) The responses of LEO incorporates rich informative spatial relations. Such information is necessary to refer to specific objects in complex 3D scenes and affords considerable assistance for humans.

3 Embodied Action in 3D World

Finally, we hope to directly probe the embodied acting and interacting capacity of LEO in the 3D World. We select two canonical embodied AI tasks: embodied navigation with ObjNav on AI Habitat (Ramrakhya et al., 2022) and robotic manipulation on CLIPort (Shridhar et al., 2021). Specifically, for ObjNav, although LEO is trained on a customized dataset (see Sec. 3.2), the scenes are all included in the original MP3D ObjNav training split (Savva et al., 2019). Therefore, we still evaluate LEO on the original MP3D ObjNav validation split against baselines. Additionally, we test LEO on the validation split of the newly introduced HM3D ObjNav task (Ramakrishnan et al., 2021). We report the success rate and SPL metrics following Ramrakhya et al. (2022). For CLIPort robotic manipulation, we evaluate LEO on the three training tasks listed in Tab. 3 and their corresponding unseen tasks and report the average reward across the evaluation episodes.

We present the results of CLIPort manipulation and object navigation in Tabs. 3 and 4. Our findings are as follows: 1) In robotic manipulation, LEO exhibits comparable performances to many strong baselines and even achieves significantly better results on some challenging unseen tasks. Note that compared to baselines that rely on heatmap output, LEO produces motor commands directly. 2) On ObjNav, LEO attains a reasonable success rate and better SPL on MP3D-val compared with baselines, suggesting that LEO can learn to leverage the object-centric 3D scene input (potentially offering a coarse global map) and take a shorter path to the target. Further, results on HM3D-val confirm LEO’s zero-shot generalization to novel scenes. Please note that all baselines use an RNN policy while LEO can be viewed as a transformer-based feed-forward policy similar to RT-2 (Brohan et al., 2023) considering the training efficiency, which could lead to a lower success rate. More discussions on this can be found in Sec. H.2. 3) Overall, the align-then-instruct tuning scheme endows LEO with semantic-level generalization (novel objects, etc.) in both manipulation and navigation tasks (we provide results on ObjNav with unseen objects in Sec. I.1).

4 Ablative Study

We ablate LEO on different data configurations. Specifically, we compare LEO’s performance under the following settings: (1) NoAlign: tuning LEO from scratch on LEO-instruct, skipping pre-training on LEO-align; (2) PartialData: uniformly sampling 10% of data in LEO-instruct during instruction-tuning; (3) ScanNetOnly: excluding data generated by LLM and embodied tasks (i.e., navigation and manipulation) during instruction tuning; (4) NoAct (VL): excluding embodied task data during instruction tuning. We provide additional results and findings in Appendix G.

We provide a more comprehensive quantitative evaluation of LEO on all 3D VL tasks, including 3D object/scene captioning, 3D QA, scene-grounded dialogue and task planning. Following prior works (Achlioptas et al., 2020), we use ground-truth object proposals in our ablation to pinpoint the reasoning and planning capabilities of LEO. We report exact match scores for QA tasks and SentenceSim for all other tasks. Fig. 4 shows a holistic view of the results.

1) The two-stage align-then-instruct pipeline is critical for LEO learning. The lack of alignment harms detailed understanding of scenes, while the decrease in instruction-tuning data affects reasoning and planning. 2) Compositional generalization poses considerable challenges. ScanNetOnly, having been exposed to 3RScan scenes or QA skills during the two stages respectively, still struggles to handle the QA task in 3RScan scenes (3RQA). 3) General vs. specific. We observe model performance drops on in-domain tasks (e.g., Scan2Cap) when adding data from other domains or new tasks (ScanNetOnly vs. NoAct (VL)). Scaling up the instruction-tuning data brings significant improvements, though the embodied acting data counteracts such effects due to the domain gap (PartialData vs. Full (VLA) vs. NoAct (VL)).

5 Scaling Law Analysis

Following the analysis in Sec. 4.4, we study the scaling effect (Kaplan et al., 2020; Reed et al., 2022) of data and model in LEO. We use the instruction-tuning loss (on the test set) of LEO with the growth of data and model scale as an indicator. Based on NoAct (VL) with Vicuna-7B (referred to as Vicuna-7B Aligned), we add two variants: (1) Vicuna-7B Scratch, trained without the alignment stage; and (2) Vicuna-13B Scratch, trained without the alignment stage and scaling up the LLM to 13B. The curves of test loss are visualized in Fig. 4(b).

1) The instruction tuning of LEO conforms to the scaling law (Kaplan et al., 2020; Reed et al., 2022). For all three settings, we find the test loss of LEO decreases log-linearly as it is fed with more data. 2) The lack of alignment causes significantly higher loss. Vicuna-7B Scratch shows consistently higher loss than Vicuna-7B Aligned. This corresponds to the inferior performances of NoAlign in Sec. 4.4 and emphasizes the importance of alignment. 3) Scaling up LLM leads to degradation. Vicuna-13B Scratch shows consistently higher loss than Vicuna-7B Scratch, which echos previous findings (Dai et al., 2023; Xu et al., 2023). We conjecture there are two possible reasons: multi-modal instruction tuning data is insufficient to reveal the benefit of scaling up LLMs, or a small-scale LLM (e.g., Vicuna-7B) already suffices for connecting the visual modality.

Related Work

The AI community has witnessed the rising generalist models in both vision (Lu et al., 2023; Wang et al., 2023b; Kirillov et al., 2023) and language (OpenAI, 2022; 2023) domains. A generalist agent requires additional embodiment knowledge to interact with the environment and complete embodied acting tasks. Existing efforts towards generalist agents include: grounded reasoning and task planning in the real world (Ahn et al., 2022; Huang et al., 2022b), skill generalization in open-world environment (Fan et al., 2022; Cai et al., 2023a; Wang et al., 2023e; a; Cai et al., 2023b; Gong et al., 2023b), general robotic manipulation (Brohan et al., 2022; Jiang et al., 2023; Gong et al., 2023a), and unified vision-language-action (VLA) models such as Gato (Reed et al., 2022), PaLM-E (Driess et al., 2023), EmbodiedGPT (Mu et al., 2023), and RT-2 (Brohan et al., 2023). LEO belongs to the VLA model, however, its goal is to build a generalist agent that can understand the real 3D world beyond 2D images, which is absent in existing works.

Pre-trained LLMs demonstrated practical for solving vision-language tasks (Tsimpoukelli et al., 2021; Alayrac et al., 2022; Guo et al., 2023; Li et al., 2023d; Zhao et al., 2023). Meanwhile, the instruction-tuning paradigm exhibited strong zero-shot generalization in NLP tasks (Wei et al., 2022; Sanh et al., 2022; Ouyang et al., 2022; Chung et al., 2022). The two streams merged into instruction-tuned LVLMs (Liu et al., 2023b; Zhu et al., 2023b; Ye et al., 2023; Gao et al., 2023; Li et al., 2023b; Gong et al., 2023c; Dai et al., 2023). Despite the burst, these models are confined to 2D visual modalities, e.g., image or video. Concurrent works (Yin et al., 2023; Hong et al., 2023; Wang et al., 2023d; Xu et al., 2023) extend to 3D vision tasks, but these models either lack the acting capability or unified efficient architecture.

One key obstacle to building LEO is grounding the 3D world with natural languages. There exist diverse methods of grounded scene understanding, e.g., spatial relation modeling (Zhao et al., 2021; Chen et al., 2022; Zhu et al., 2023c) and fine-grained open-scene understanding (Peng et al., 2023b; Kerr et al., 2023). However, due to data scarcity, how to utilize LLMs to ground the 3D scene is rarely explored. Recently, 3D-LLM (Hong et al., 2023) leverages multi-view images and Chat-3D (Wang et al., 2023d) uses object-centric point clouds to enable the LLMs with 3D grounding. In this work, we devise both 2D and 3D encoders for grounding various visual representations and employ LoRA (Hu et al., 2022) to efficiently fine-tune the LLMs.

LLMs exhibit extraordinary capabilities of text generation and serve as a source for collecting diverse instruction-following data (Wang et al., 2023c; Taori et al., 2023; Peng et al., 2023a). However, the lack of access to visual modalities makes it troublesome to collect visual instruction-tuning data. To address this issue, existing methods provide bounding boxes (Liu et al., 2023b) and add dense captions (Li et al., 2023a; Liu et al., 2023a) as image descriptions or directly use off-the-shelf large vision-language models (LVLM) (Zhu et al., 2023a; Luo et al., 2023) to help collect such data. Unlike concurrent attempts (Yin et al., 2023; Hong et al., 2023; Wang et al., 2023d) in collecting 3D instruction-tuning data, our approach features a scene-graph-based prompting and refinement method to prompt and correct the data.

Conclusions

The proposed agent LEO extends the current generalist ability of LLMs from text towards the 3D world and embodied tasks. It is a crucial initial step toward building embodied generalist agents. In light of this work, we identify several promising directions that hold the potential for substantial advancement: (1) enhancing the 3D vision-language grounding capability by leveraging larger-scale paired data from richer real-world 3D domains; (2) continually bridging the gap between 3D vision-language and embodied action, as our experiments reveal the feasibility of their joint learning; (3) investigating the issues of safety and alignment in the context of embodied generalist agents, particularly given that our scaling law analysis suggests that such agents can experience significant enhancements through data scaling in the near future.

References

Appendix A Data

In Fig. 5–9, we show the prompts for five types of LLM-assisted 3D-language data generation. We provide few-shot examples as the context. In each example, the “content” contains a scene graph, and the “response” refers to a human-labeled response. The query is a new scene graph, based on which ChatGPT (OpenAI, 2022) generates responses.

Fig. 5 shows the prompt for generating 3D dialogue data. Red fonts outline our requirements of the dialogue content, including object attributes, spatial relations, and commonsense topics. Purple fonts formulate the template of the response. We require the response generated by the ChatGPT should include the dialogue context as well; the “thought” contains the involved objects in the question, which is used to enhance the reliability of the answer. These two components will be removed after the refinement procedures.

A.2 Analysis of the Object-Centric Chain-of-Thought

To further investigate the impact of Object-centric Chain-of-Thought (O-CoT) on data quality, we analyze the answer accuracy for Object Counting questions. Specifically, we collect several demonstrations, and for each run, we select two of them as the prompt seed. With these seeds, we generate dialogues across all scenes in 3DSSG (Wu et al., 2021) and then assess the answer accuracy for Object Counting questions. The results are presented in Tab. 5.

The results in Tab. 5 indicate that O-CoT consistently improves the answer accuracy for Object Counting questions. Though there remain errors after applying O-CoT, we will conduct refinement to fix them. Examples of Object Counting questions are provided in Sec. A.3.

A.3 Refinement Details

We conduct refinement by passing raw LLM-generated responses into several human-defined filtering procedures based on the 3D scene graph. The refinement considers five raw response categories:

Object Counting. The question concerns counting the target object.

Object Existence. The response claims the existence of objects, which can be actually either existent or non-existent.

Object Non-existence. The response claims the non-existence of objects, which can be actually either existent or non-existent.

Negative Response. The scene graph cannot provide a solid response to the question, which means the question cannot be answered and will be discarded.

Response with ID. The response contains unexpected object IDs.

Specifically, we employ regular expression matching to detect errors in these five categories. And we also employ this method to correct the responses except for Response with ID, which will be rewritten by ChatGPT instead. The QA pair will be eliminated if multiple rounds of rewriting fail to remove the IDs. Tab. 6 and Tab. 7 show some examples of the responses subject to the above five categories as well as the effect of our refinement.

A.4 Statistics of Raw Responses

Based on the aforementioned five raw response categories, we assess their quality by statistics and clarify the refinement effect accordingly. In Tab. 8, we quantify the answer accuracy for Object Counting, Object Existence, and Object Non-existence in dialogue and QA tasks. Results of the two tasks are averaged over all 3DSSG scenes across 6 prompt seeds and 3 prompt seeds, respectively. For these three categories of responses, we can fix almost all the detected errors by referring to the scene graph. In Tab. 9, we present the proportion of Negative Response and Response with ID in the total set of responses. Negative responses will be removed, as well as the responses with remaining IDs after multiple rounds of rewriting. All results are based on the O-CoT method.

A.5 Subgraph Sampling

To enhance the diversity of the 3D scene graphs used for prompting, we perform subgraph sampling on the 3DSSG according to a sampling rate, which denotes the ratio of preserved nodes. The sampled subgraphs are used for generating scene captions and planning data. We analyze the distribution of node numbers across the 3DSSG dataset in Fig. 10 and set different sampling rates for scenes with different numbers of nodes in Tab. 10. For each sampling rate, we set 4 random prompt seeds to further enhance the diversity of prompted data.

To verify whether the subgraph sampling strategy can maintain the consistency and diversity of scene captions, we generate scene captions for the same scene using both the full graph and subgraph. We then employ GPT-4 (OpenAI, 2023) to evaluate the similarities and differences between the two captions. The results in Tab. 11 indicate that our subgraph sampling strategy can maintain both consistency and diversity.

A.6 Scene-graph-based Prompting vs. Box-based Prompting

In this section, we provide a comparative analysis of scene-graph-based prompting and box-based prompting (Hong et al., 2023). We refer the readers to Figure 6 in 3D-LLM (Hong et al., 2023) for details of the box-based prompting method. Fig. 11 shows the contents of two methods. To present a fair comparison between the two methods, we prompt with 1) demonstrations that have similar content under the same scene (see Fig. 12) and 2) identical new scene queries. Since 3D-LLM does not elaborate on attribute-related prompts, we mainly compare the spatial relations in the responses. As shown in Fig. 13, we highlight some spatial relations in red. The comparison shows that our method provides more diverse and reliable spatial relations, which are important for 3D scene understanding.

A.7 Dataset Statistics

We provide statistics on the instruction-tuning datasets. We visualize the distribution of the question types in 3RQA (Fig. 15) and 3RDialog (Fig. 15). The pie chart’s inner circle represents the first word of the questions, while the outer circle accounts for the second or third word in the corresponding questions. The results show that the questions cover the attributes and spatial relations of the objects, as well as high-level topics such as room types and functionalities.

We also provide statistics of the root noun-verb pairs for instructions and responses in 3RDialog and 3RPlan, as shown in Fig. 17–19.

Appendix B Action Tokenization

To empower LEO to exert control over an embodiment or a robot, we encode all actions within the context of Object Navigation (Ramrakhya et al., 2022) and CLIPort (Shridhar et al., 2021) tasks using the least frequently employed language tokens. Specifically, for the Object Navigation task, we allocate 4 tokens to represent actions of move forward, turn right, turn left, and stop. For the CLIPort task, we use a total of 516 tokens to discretize action poses, with 320 tokens dedicated to the x-axis pose bins, 160 tokens for the y-axis pose bins, and 36 tokens for the z-rotation bins.

Appendix C Data Examples

Please refer to Tabs. 24–26 for examples of our dataset.

Appendix D Model Details

The first portion of prompts sent into the LLM is a system message. It consists of two parts: a role prompt and a situation prompt. The role prompt is the same for all tasks:

You are an AI visual assistant situated in a 3D scene. You can perceive (1) an ego-view image (accessible when necessary) and (2) the objects (including yourself) in the scene (always accessible). You should properly respond to the USER’s instructions according to the given visual information. The situation prompt begins with a common sentence:

You are at a selected location in the 3D scene. For SQA3D (Ma et al., 2023), the situation prompt is further extended with the situation description in the dataset. The situation prompt is only used jointly with the embodiment token to support tasks that require information about the embodiment. Details can be found in Sec. D.2.1.

Next are the visual tokens, including 2D image tokens and object-centric 3D tokens. Each token sequence is interleaved within text tokens and starts with a text prefix.

Ego-view image: {IMAGE_TOKENS} Objects (including you) in the scene: {OBJECT_TOKENS} The last portion of prompts is a task-specific instruction. For object-level caption and object-in-the-scene caption, we randomly chose one sentence from 151 sentences to be the instruction. Some examples can be found in Tab. 12. For scene-level caption, we randomly choose one from 183 instructions. Examples can be found in Tab. 13. For 3D question answering task, we simply use the question as the instruction. The dialog history is used as the instruction for 3D dialogue to provide continuity across multiple rounds of interactions. A planning instruction pool consisting of 202 instructions is introduced for scene-aware task planning and we randomly choose one from it as done in the caption tasks. Examples from the pool can be found in Tab. 14. The chosen instruction is further followed by an instruction that specifies the task, e.g., set up a home office.

With past action tokens {PAST_ACTIONS} appended at the end, the instruction for embodied navigation is as follows, where {GOAL} stands for the goal specified by the target object name:

The task is navigation. Your goal is to find {GOAL} by moving around in the scene. Past actions: {PAST_ACTIONS}. The instruction for robotic manipulation is similar to the one in embodied navigation. Here {GOAL} is the task description in CLIPort:

D.2 Feature Encoding

We have several modules to encode the multi-modal features.

Object-centric 3D token embedding. The encoder for 3D object-centric point clouds is a PointNet++ (Qi et al., 2017) pre-trained on ScanNet (Dai et al., 2017) with object-classfication task. We sample 1024 points for every object as in Chen et al. (2022). The architecture parameters all remain the same with Chen et al. (2022). We freeze the PointNet++ for empirically better results.

Spatial Transformer (Chen et al., 2022). Spatial Transformer is a modified transformer architecture that explicitly encodes spatial relations between object pairs. Specifically, consider the vanilla self-attention (Vaswani et al., 2017) mechanism which takes as input a feature matrix X∈RN×dX\in\mathbf{R}^{N\times d}, where NN stands for the number of tokens and dd is the feature dimension. Vanilla self-attention first compute Q=XWQ,K=XWK,V=XWVQ=XW_{Q},K=XW_{K},V=XW_{V} from XX using learnable projection matrices WQ,WK,WV∈Rd×dhW_{Q},W_{K},W_{V}\in\mathbf{R}^{d\times d_{h}} where dhd_{h} stands for the output feature dimension. Then the attention weight matrix is computed by (ωijo)N×N=Ωo=softmax(QKTdh)(\omega^{o}_{ij})_{N\times N}=\Omega^{o}=softmax(\frac{QK^{T}}{\sqrt{d_{h}}}) and finally used for re-weighting ΩoV\Omega^{o}V. The intuition of Spatial Transformer is that we can re-scale the elements ωijo\omega_{ij}^{o} in the weight matrix Ωo\Omega^{o}.

In the object-centric reasoning setting, the input feature matrix is O∈RN×dO\in\mathbf{R}^{N\times d}. Consider an object pair (Oi,Oj)(O_{i},O_{j}) with their geometric centers ci,cjc_{i},c_{j}. Spatial Transformer (Chen et al., 2022) computes the Euclidean distance dij=∣∣ci−cj∣∣2d_{ij}=||c_{i}-c_{j}||_{2} and the horizontal and vertical angles θh,θv\theta_{h},\theta_{v} of the line connecting cic_{i} and cjc_{j}. The spatial feature between the two objects (Oi,Oj)(O_{i},O_{j}) is a 5-dimensional vector fij=[dij,sin⁡(θh),cos⁡(θh),sin⁡(θv),cos⁡(θv)]f_{ij}=[d_{ij},\sin{(\theta_{h})},\cos{(\theta_{h})},\sin{(\theta_{v})},\cos{(\theta_{v})}]. To combine this feature with objects, the spatial attention computes ωijs=gifij\omega^{s}_{ij}=g_{i}f_{ij} where gi=WSToig_{i}=W_{S}^{T}o_{i} is a 5-dimensional vector. The spatial attention further reweights the original self-attention weight matrix as

Readers are referred to Chen et al. (2022) for more details. In summary, Spatial Transformer explicitly computes pairwise spatial relations and fuses them with vanilla self-attention to provide better spatial reasoning ability. We use a three-layer Spatial Transformer with 8 heads to process the object-centric features produced by PointNet++ and output object tokens for LLM. For other settings, We follow all the default hyperparameters in Chen et al. (2022).

2D token embedding. We use OpenCLIP ConvNext-base model (Liu et al., 2022) pre-trained on LAION2B (Schuhmann et al., 2022) to process the egocentric 2D image.

CLIP fusion. To enhance the alignment between visual tokens and instruction tokens, we use the text encoder from CLIP (Radford et al., 2021) to process the instruction tokens to obtain a global feature of the instruction. Next, we update the visual tokens with the element-wise product between the CLIP instruction feature and each image & object token embedding.

In addition to the egocentric 2D input, we introduce an embodiment token to help LEO reason in an embodiment-aware fashion. We find it useful to use it together with the situation prompt and 2D egocentric input. Specifically, an embodiment token ee is introduced in embodied navigation, embodied reasoning, and object-in-the-scene caption tasks. Specifically, ee is a learnable embedding that will be inserted into the 3D object list.

So what does embodiment information mean in these tasks? In embodied navigation, it means the agent’s position and orientation in the scene, which can be derived from a GPS and a compass sensor. The orientation of the agent is further represented by a rotation which is Fourier-embedded and mapped to a feature vector rr by a linear layer. It is the same in embodied reasoning task. In the object-in-the-scene caption task, we assume the agent is situated at the location of the object that is being referred to. Therefore, embodiment information also means the location of the referred object. We obtain this location by randomly choosing a spot inside the referred object bounding box. To sum up, we could simply treat the embodiment token as a special self object, where its object embedding is learnable, and its location/orientation corresponds to the actual or assumed “agent”.

After inserting the embodiment token, we obtain a new 3D object token list: e,s3D(1),s3D(2),…,s3D(N)e,s_{\text{3D}}^{(1)},s_{\text{3D}}^{(2)},\dots,s_{\text{3D}}^{(N)}, where s3D(i),i∈{1,2,…,N}s_{\text{3D}}^{(i)},i\in\{1,2,\dots,N\} are 3D object token embeddings produced by PointNet++, along with location specified for each object (including the self-object). We can concatenate them together to get a feature matrix O∈R(N+1)×dO\in\mathbf{R}^{(N+1)\times d} and send them to the Spatial Transformer to explicitly fuse the spatial information of all the 3D objects and the self-object.

D.3 LLM Hyperparameters

We set the maximum output length of our Vicuna-7B to be 256. The maximum context length is also set to 256 and if the length of the input is greater than 256, we truncate it to 256 by deleting tokens from the left (i.e., only the rightmost 256 tokens are preserved). We set rank and α\alpha in LoRA (Hu et al., 2022) to be 16 and the dropout rate to be 0. LoRA is implemented for all the projection matrices in the LLM, i.e., (Wq,Wk,Wv,Wo)(W_{q},W_{k},W_{v},W_{o}) in attention modules and (Wgate,Wup,Wdown)(W_{gate},W_{up},W_{down}) in MLPs.

The hyperparameters for beam search during inference are as follows:

Appendix E Alignment Setup

The hyperparameters for the first-stage 3D vision-language alignment are presented in Tab. 16

Appendix F Instruction-tuning Setup

The hyperparameters for 3D VLA instruction tuning are presented in Tab. 17

Appendix G Ablation Details

We present the numerical results in Tab. 18 as complements to Fig. 4(a).

We also make some explorations in model ablation and simply present qualitative findings here. For the point cloud encoder, we choose Point-BERT (Yu et al., 2022) as an alternative to the default PointNet++ (Qi et al., 2017). We utilize the checkpoint from PointLLM (Xu et al., 2023), which has adapted Point-BERT to 6-channel (XYZRGB) input and learned a language-aligned representation for 3D objects. Despite larger capacity, Point-BERT shows significantly worse performances than PointNet++ as the point cloud encoder. Similarly, for the Spatial Transformer and LLM, we ablate with different model scales but find no improvement. The influence of different modules remains an interesting question that deserves further exploration.

Appendix H Evaluation Details

We argue that exact match (EM), as a conventional metric for 3D QA, is unsuitable for evaluating the open-ended answer generated by LLMs. For example, given the question “On what side of the towel is a bathroom curtain?” with ground-truth answer “left side of towel”, it is never wrong to answer “left”. However, this will be deemed incorrect if we adopt the strict exact match protocol. Such a misjudgment is quite likely to occur when evaluating the answers from LLMs. By contrast, the classifier heads for QA (e.g., MCAN) are less affected because they collect all possible answers in advance to formulate the QA as a close-set classification problem. Hence, we refine the strict exact match protocol as follows.

In a nutshell, we squeeze the pred and gt, and then check whether one is a subset of the other. To justify our refined exact match protocol, in Tab. 19 we provide some representative examples in the ScanQA validation set. Despite the improvements, we speculate such a simple refinement is still insufficient for a sound evaluation metric considering the flexibility of human language.

H.2 Embodied Navigation

To construct our training set, we adopt all 57 scenes in the MP3D ObjNav training split (Savva et al., 2019; Ramrakhya et al., 2022) and generate ~60K shortest-path navigation episodes. The evaluation is conducted on the original validation split of the MP3D ObjNav task and the newly introduced HM3D ObjNav task (Ramakrishnan et al., 2021).

In contrast to most ObjNav agents that utilize recurrence through either RNN (Ramrakhya et al., 2022) or DT-style Transformer (Suglia et al., 2021), LEO only employs a simplistic feed-forward policy, i.e., the Transformer in LEO only takes in the instruction, current state (2D and 3D observation), and past 4 actions, and predicts the next action, similar to RT-2 (Brohan et al., 2023). Therefore, the only information relayed from the past is about past actions. The absence of recurrence in LEO’s acting policy is indeed the result of a trade-off between better performances and training efficiency. We will commit to exploring the possibility of looping in more sophisticated policy architectures (e.g., recurrence) in future work.

Appendix I Additional Results

Quantitative results of ObjNav. We provide additional results of LEO 1) generalizing to unseen objects on MP3D, and 2) learning with 70K human demonstrations provided by Habitat-web (Ramrakhya et al., 2022) instead of shortest path. Below is a list of the objects used during training (seen) and for OOD evaluation (unseen). Evaluation results are shown in Tab. 20. Note that the baseline Habitat-web is unable to generalize to novel objects as it uses categorical embedding rather than natural language to represent object goals.

# Objects (seen) ‘‘gym_equipment’’, ‘‘tv_monitor’’, ‘‘picture’’, ‘‘counter’’, ‘‘chair’’, ‘‘cabinet’’, ‘‘table’’, ‘‘stool’’, ‘‘plant’’, ‘‘towel’’, ‘‘sofa’’, ‘‘cushion’’, ‘‘sink’’, ‘‘fireplace’’, ‘‘toilet’’, ‘‘seating’’, ‘‘chest_of_drawers’’, ‘‘bed’’, ‘‘shower’’, ‘‘bathtub’’, ‘‘clothes’’ # Objects (unseen) ‘‘shelf’’, ‘‘pillow’’, ‘‘lamp’’, ‘‘box’’, ‘‘desk’’, ‘‘refrigerator’’, ‘‘vase’’, ‘‘armchair’’ The results show that LEO can generalize to novel objects. On the other hand, human demonstrations include more explorations, compared with shortest-path data. Therefore, it will be much harder for agents without a recurrent module (e.g., LEO) to learn from human demonstrations (see Sec. H.2), leading to significantly weaker performances.

Qualitative results. We provide more qualitative results of robotic manipulation and embodied navigation in the supplementary video.

I.2 Scan2Cap

We provide additional qualitative results on Scan2Cap validation set in Tab. 21. The results show that LEO can correctly refer to the queried object and provide accurate descriptions, including spatial relationships with other objects. However, LEO’s responses are confined to simple formats that lack diversity. How to unlock more flexible responses while maintaining accuracy can be a direction for future research.

I.3 ScanQA

We provide additional qualitative results on ScanQA validation set in Tab. 22 and categorize the responses into several types:

Wrong. The response is inaccurate and deemed wrong.

Wrong but reasonable. The response is deemed wrong but is reasonable to some extent, probably due to ambiguities in the scene. Consider the second case in Tab. 22. There are many objects such as a coat rack, a coat, and a mini fridge-shaped cabinet on the right side of the organizer. Though LEO’s response “mini fridge” does not match the ground truth “coat rack”, it is consistent with the 3D scene layout.

Wrong but accurate. The response is accurate according to the scene but is deemed wrong due to imperfect ground truth annotations.

Correct. The response is accurate and deemed correct.

Correct and more accurate. The response is more accurate than the ground truth annotations.

I.4 SQA3D

We provide additional qualitative results on SQA3D test set in Tab. 23 and follow the aforementioned response types. The embodied reasoning in SQA3D requires the understanding of not only the scene but also the situation of embodiment. In Tab. 23, answering “What am I sitting at?” necessitates that LEO accurately identifies the objects at its current location. And the response to “How many beds are in front of me?” indicates that LEO can reason based on the understanding of its orientation.