A Language Agent for Autonomous Driving
Jiageng Mao, Junjie Ye, Yuxi Qian, Marco Pavone, Yue Wang
Introduction
Imagine a car navigating a quiet suburban neighborhood. Suddenly, a ball bounces onto the road. A human driver, leveraging extensive experiential knowledge, would not only perceive the immediate presence of the ball, but also instinctively anticipate the possibility of a chasing child and consequently decide to decelerate. In contrast, an autonomous vehicle, devoid of such reasoning and experiential anticipation, might continue driving until sensors detect the child, only allowing for a narrower margin of safety. The importance of human prior knowledge in driving systems becomes clear: driving is not merely about reacting to the visible, but also to the conceivable scenarios where the system needs to reason and respond even in their absence.
To integrate human prior knowledge into autonomous driving systems, previous approaches deconstruct the human driving process into three systematic steps following Figure 1 (a). Perception: they interpret the human perceptual process as object detection or occupancy estimation . Prediction: they abstract human drivers’ foresight of upcoming scenarios as the prediction of future object motions . Planning: they emulate the human decision-making process by planning a collision-free trajectory, either using hand-crafted rules or by learning from data . Despite its efficacy, this perception-prediction-planning framework overly simplifies the human decision-making process and cannot fully model the complexity of driving. For instance, perception modules in these methods are notably redundant, necessitating the detection of all objects in a vast perception range, whereas human drivers can maintain safety by only attending to a few key objects. Moreover, prediction and planning are designed for collision avoidance with detected objects. Nevertheless, they lack deeper reasoning ability inherent to humans, e.g. deducing the connection between a visible ball and a potentially unseen child. Furthermore, it remains challenging to incorporate long-term driving experiences and common sense into existing autonomous driving systems.
In addressing these challenges, we found the major obstacle of integrating human priors into autonomous driving lies in the incompatibility of human knowledge and neural-network-based driving systems. Human knowledge is inherently encoded and utilized as language representations, and their reasoning process can also be interpreted by language. However, conventional driving systems rely on deep neural networks that are designed to process numerical data inputs, such as sensory signals, bounding boxes, and trajectories. The discrepancy between language and numerical representations poses a significant challenge to incorporating human experiential knowledge and reasoning capability into existing driving systems, thereby widening the chasm from genuine human driving performance.
Taking a step towards more human-like autonomous driving, we propose Agent-Driver, a cognitive agent empowered by Large Language Models (LLMs). The fundamental insight of our approach lies in the utilization of natural language as a unified interface, seamlessly bridging language-based human knowledge and reasoning ability with neural-network-based systems. Our approach fundamentally transforms the conventional perception-prediction-planning framework by leveraging LLMs as an interactive scheduler among system components. As depicted in Figure 1 (b), on top of the LLMs, we introduce: 1) a versatile tool library that interfaces with neural modules via dynamic function calls, streamlining perception with less redundancy, 2) a configurable cognitive memory that explicitly stores common sense and driving experiences, infusing the system with human experiential knowledge, and 3) a reasoning engine that processes perception results and memory data to emulate human-like decision-making. Specifically, the reasoning engine performs chain-of-thought reasoning to recognize key objects and events, task planning to derive a high-level driving plan, motion planning to generate a driving trajectory, and self-reflection to ensure the safety of the planned trajectory. These components, coordinated by LLMs, culminate in an anthropomorphic driving process. To conclude, we summarize our contributions as follows:
We present Agent-Driver, an LLM-powered agent that revolutionizes the traditional perception-prediction-planning framework, establishing a powerful yet flexible paradigm for human-like autonomous driving.
Agent-Driver integrates a tool library for dynamic perception and prediction, a cognitive memory for human knowledge, and a reasoning engine that emulates human decision-making, all orchestrated by LLMs to enable a more anthropomorphic autonomous driving process.
Agent-Driver significantly outperforms the state-of-the-art autonomous driving systems by a large margin, with over collision improvements in motion planning. Our approach also demonstrates strong few-shot learning ability and interpretability on the nuScenes benchmark.
We provide a variety range of ablation study to dissect the proposed architecture and understand the efficacy of each module, to facilitate future research in this direction.
Due to the page limit, we leave further discussions of related works in the appendix for reference.
Agent-Driver
In this section, we present Agent-Driver, an LLM-based intelligent agent for autonomous driving. We first introduce the overall architecture of our Agent-Driver in Section 2.1. Then, we introduce the three key components of our method: tool library (Section 2.2), cognitive memory (Section 2.3), and reasoning engine (Section 2.4).
Conventional perception-prediction-planning pipelines leverage a series of neural networks as basic modules for different tasks. However, these neural-network-based systems lack direct compatibility with human prior knowledge, constraining their potential for leveraging such priors to enhance driving performance. To handle this challenge, we propose a novel framework that leverages text representations as a unified interface to connect neural networks and human knowledge. The overall architecture of Agent-Driver is shown in Figure 2. Our approach takes sensory data as input and introduces neural modules for processing these sensory data and extracting environmental information about detection, prediction, occupancy, and map. On top of the neural modules, we propose a tool library where a set of functions are designed to further abstract the neural outputs and return text-based messages. For each driving scenario, an LLM selectively activates the required neural modules by invoking specific functions from the tool library, ensuring the collection of necessary environmental information with less redundancy. Upon gathering the necessary environmental information, the LLM leverages this data as a query to search in a cognitive memory for pertinent traffic regulations and the most similar past driving experience. Finally, the retrieved traffic rules and driving experience, together with the formerly collected environmental information, are utilized as inputs to an LLM-based reasoning engine. The reasoning engine performs multi-round reasoning based on the inputs and eventually devises a safe and comfortable trajectory for driving. Our Agent-Driver architecture harnesses dynamic perception and prediction capability brought by the tool library, human knowledge from the cognitive memory, and the strong decision-making ability of the reasoning engine. This synergistic integration results in a more human-like driving system with enhanced decision-making capability.
2 Tool Library
The profound challenge of incorporating human knowledge into neural-network-based driving systems is reconciling the incompatibility between text-based human priors and the numerical representations from neural networks. While prior works have attempted to translate text-based priors into semantic features or regularization terms for integration with neural modules, their performances are still constrained by the inherent cross-modal discrepancy. By contrast, we leverage text as a unified interface to connect neural modules and propose a tool library built upon the neural modules to dynamically collect text-based environmental information.
The cornerstones of the tool library are four neural modules, i.e., detection, prediction, occupancy, and map modules, which process sensory data from observations and generate detected bounding boxes, future trajectories, occupancy grids, and maps respectively. The neural modules cover various tasks in perception and prediction and extract environmental information from observations. However, this information is largely redundant, with a significant portion insignificant to decision-making. To dynamically extract necessary information from the neural module outputs, we propose a tool library—where a set of functions are designed—to summarize the neural outputs into text-based messages, and the information collection process can be established by dynamic function calls. An illustration of this process is shown in Figure 3.
Functions. We devised various functions for detection, prediction, occupancy, and mapping, in order to extract useful information from the neural module’s outputs respectively. Our tool library contains more than 20 functions covering diverse usages. Here are some examples. For detection, get_leading_object returns a text description of the object in front of the ego-vehicle on the same lane. For prediction, get_pred_trajs_for_object returns a text-based predicted future trajectory for a specified object. For occupancy, get_occ_at_loc_time returns the probability that a specific location is occupied by other objects at a given timestep. For map, get_lanes returns the information of the left and right lanes to the ego-vehicle, and get_shoulders returns the information of the left and right road shoulders to the ego-vehicle. Detailed descriptions of all functions are in the appendix.
Tool Use. With the functions in the tool library, an LLM is instructed to collect necessary environmental information through dynamic function calls. Specifically, the LLM is first provided with initial information such as the current state for its subsequent decision-making. Then, the LLM will be asked whether it is necessary to activate a specific neural module, i.e. detection, prediction, occupancy, and map. If the LLM decides to activate a neural module, the functions related to this module will be provided to the LLM, and the LLM chooses to call one or some of these functions to collect the desired information. Through multiple rounds of conversations, the LLM eventually collects all necessary information about the current environment. Compared to directly utilizing the outputs of the neural modules, our approach reduces the redundancy in current systems by leveraging the reasoning power of the LLM to determine what environmental information is of real importance to the decision-making process. Furthermore, the neural modules are only activated when the LLM decides to call the relevant functions, which brings flexibility to the system.
3 Cognitive Memory
Human drivers navigate using their common sense, such as adherence to local traffic regulations, and draw upon driving experiences in similar situations. However, it is non-trivial to adapt this ability to the conventional perception-prediction-planning framework. By contrast, our approach tackles this problem through interactions with a cognitive memory. Specifically, the cognitive memory stores text-based common sense and driving experiences. For every driving scenario, we utilize the collected environmental information as a query to search in the cognitive memory for similar past experiences to assist decision-making. The cognitive memory contains two sub-memories: commonsense memory and experience memory.
Commonsense Memory. The commonsense memory encapsulates the essential knowledge a driver typically needs for driving safely on the road, such as traffic regulations and knowledge about risky behaviors. It is worth noting that the commonsense memory is purely text-based and fully configurable, that is, users can customize their own commonsense memory for different driving conditions by simply writing different types of knowledge into the memory.
Experience Memory. The experience memory contains a series of past driving scenarios, where each scenario is composed of the environmental information and the subsequent driving decision at that time. By retrieving the most similar experiences and referencing their driving decisions, our system enhances its capacity for making more informed and resilient driving decisions.
Memory Search. As exhibited in Figure 4, we present an innovative two-stage search algorithm to effectively search for the most similar past driving scenario in the experience memory. The first stage of our algorithm is inspired by vector databases , where we encode the input query and each record in the memory into embeddings and then retrieve the top-K similar records via K-nearest neighbors (K-NN) search in the embedding space. Since the driving scenarios are quite diverse, the embedding-based search is inherently limited by the encoding methods employed, resulting in insufficient generalization capabilities. To overcome this challenge, the second stage incorporates an LLM-based fuzzy search, where the LLM is tasked to rank these records according to their relevance to the query. This ranking is based on the implicit similarity assessment by the LLM, leveraging its capabilities in generalization and reasoning.
The cognitive memory equips our system with human knowledge and past driving experiences. The retrieved most similar experiences, together with commonsense and environmental information, collectively form the inputs to the reasoning engine. Text as a unified interface aligns the environmental information with human knowledge, thereby enhancing our system’s compatibility.
4 Reasoning Engine
Reasoning, a fundamental ability of humans, is critical to the decision-making process. Conventional methods directly plan a driving trajectory based on perception and prediction results, while they lack the reasoning ability inherent to human drivers, resulting in insufficient capability to handle complicated driving scenarios. Conversely, as shown in Figure 5, we propose a reasoning engine that effectively incorporates reasoning ability into the driving decision-making process. Given the environmental information and retrieved memory, our reasoning engine performs multi-round reasoning to plan a safe and comfortable driving trajectory. The proposed reasoning engine consists of four core components: chain-of-thought reasoning, task planning, motion planning, and self-reflection.
Chain-of-Thought Reasoning. Human drivers are able to identify the key objects and their potential effects on driving decisions, while this important capability is typically absent in conventional autonomous driving approaches. To embrace this reasoning ability in our system, we propose a novel chain-of-thought reasoning module, where we instruct an LLM to reason on the input environmental information and output a list of key objects and their potential effects in text. To guide this reasoning process, we instruct the LLM via in-context learning of a few human-annotated examples. We found this strategy successfully aligns the reasoning power of the LLM with the context of autonomous driving, leading to improved reasoning accuracy.
Task Planning. High-level driving plans provide essential guidance to low-level motion planning. Nevertheless, traditional methods directly perform motion planning without relying on this high-level guidance, leading to sub-optimal planning results. In our approach, we define high-level driving plans as a combination of discrete driving behaviors and velocity estimations. For instance, the combination of a driving behavior change_lane_to_left and a velocity estimation deceleration results in a high-level driving plan change_lane_to_left_with_deceleration. We instruct an LLM via in-context learning to devise a high-level driving plan based on environmental information, memory data, and chain-of-thought reasoning results. The devised high-level driving plan characterizes the ego-vehicle’s coarse locomotion and serves as a strong prior to guide the subsequent motion planning process.
Motion Planning. Motion planning aims to devise a safe and comfortable trajectory for driving, and each trajectory is represented as a sequence of waypoints. Following , we re-formulate motion planning as a language modeling problem. Specifically, we leverage environmental information, memory data, reasoning results, and high-level driving plans collectively as inputs to an LLM, and we instruct the LLM to generate text-based driving trajectories by reasoning on the inputs. By fine-tuning with human driving trajectories, the LLM can generate trajectories that closely emulate human driving patterns. Finally, we transform the text-based trajectories back into real trajectories for system execution.
Self-Reflection. Self-reflection is a crucial ability in humans’ decision-making process, aiming to re-assess the former decisions and adjust them accordingly. To model this capability in our system, we propose a collision check and optimization approach. Specifically, for a planned trajectory from the motion planning module, the collision_check function in the tool library is first invoked to check its collision. If collision detected, we refine the trajectory into a new trajectory by optimizing the cost function :
Details about the formulation of collision probability can be found in the appendix. Self-reflection greatly improves the safety of our system.
Our reasoning engine models the human decision-making process in driving as a step-by-step procedure involving reasoning, hierarchical planning, and self-reflection. Compared to prior works, our approach effectively emulates the human decision-making process, leading to enhanced decision-making capability and superior planning performance.
Experiments
In this section, we demonstrate the effectiveness, few-shot learning ability, and other characteristics of Agent-Driver through extensive experiments on the large-scale nuScenes dataset . First, we introduce the experimental settings and evaluation metrics in Section 3.1. Next, we compare our approach against state-of-the-art driving methods on the nuScenes dataset (Section 3.2). Subsequently, we investigate the few-show learning ability (Section 3.3), interpretability (Section 3.4), compatibility (Section 3.5), and stability (Section 3.6) of Agent-Driver. Finally, we conduct empirical studies to evaluate the effectiveness of different components and design choices of our approach in Section 3.7.
Dataset. The nuScenes dataset is a large-scale and real-world autonomous driving dataset that contains driving scenarios and approximately key frames encompassing a diverse range of locations and weather conditions. We follow the general practice and split the whole dataset into training and validation sets. We utilize the training set to train the neural modules and instruct the LLMs, and we utilize the validation set to evaluate the performance of our approach, ensuring a fair comparison with prior works.
Implementation details. We utilize gpt-3.5-turbo-0613 as the foundation LLM for different components in our system. For motion planning, we follow and fine-tune the LLM with human driving trajectories in the nuScenes training set for one epoch. For neural modules, we adopted the modules in . More details can be found in the appendix.
Evaluation metrics. As argued in , autonomous driving systems should be optimized in pursuit of the ultimate goal, i.e., planning of the self-driving car. Hence, we focus on the motion planning performances to evaluate the effectiveness of our system. There are two commonly adopted metrics for motion planning on the nuScenes dataset: L2 error (in meters) and collision rate (in percentage). The average L2 error is computed by measuring each waypoint’s distance in the planned and ground-truth trajectories, reflecting the proximity of a planned trajectory to a human driving trajectory. The collision rate is calculated by placing an ego-vehicle box on each waypoint of the planned trajectory and then checking for collisions with the ground truth bounding boxes of other objects, reflecting the safety of a planned trajectory. We follow the common practice and evaluate the motion planning result in a -second time horizon.
We further note that in different papers there are subtle discrepancies in computing these two metrics. For instance, in UniAD both metrics at -th second are measured as the error or collision rate at this certain timestep, while in ST-P3 and following works , these metrics at -th second is an average over seconds. There are also differences in ground truth objects for collision calculation in different papers. We explain the details of metric implementations in appendix for reference. For a fair comparison, we group different methods based on their metric implementations (UniAD metric or ST-P3 metric) and evaluate Agent-Driver on both two metric implementations.
2 Comparison with State-of-the-art Methods
As shown in Table 1, Agent-Driver surpasses state-of-the-art methods in both metrics and decreases the collision rate of the second-best performance by a large margin. Specifically, under ST-P3 metrics, Agent-Driver realizes the lowest average L2 error and greatly reduces the average collision rates by 35.7% compared to the second-best performance. Under UniAD metrics, Agent-Driver achieves an L2 error of 0.74 and a collision rate of 0.21%, which are 11.9% and 32.3% better than the second-best methods GPT-Driver and UniAD , respectively. The promising performance on the collision rate verifies the effectiveness of the reasoning ability of Agent-Driver, which considerably increases the safety of the proposed autonomous driving system.
3 Few-shot Learning
To assess the generalization ability of motion planning in our approach, we conduct a few-shot learning experiment, where we keep other components the same and fine-tuned the core motion planning LLM with 0.1%, 1%, 10%, 50%, and 100% of the training data for one epoch. For comparison, we adopted the motion planner in UniAD trained with 100% data as the baseline. The results are shown in Figure 6. Notably, with only 0.1% of the full training data, i.e., 23 training samples, Agent-Driver realizes a promising performance. When exposed to 1% of training scenarios, the proposed method surpasses the baseline by a large margin, especially under the average collision rate. Furthermore, with increased training data, Agent-Driver stably achieves better motion planning performance.
4 Interpretability
Unlike conventional driving systems that rely on black-box neural networks to perform different tasks, the proposed Agent-Driver inherits favorable interpretability from LLMs. As shown in Figure 7, the output messages of LLMs from the tool library, cognitive memory, and reasoning engine are recorded during system execution. Hence the whole driving decision-making process is transparent and interpretable.
5 Compatibility
Compatibility with different neural modules. As shown in Table 2, Agent-Driver constantly maintains a favorable performance with combinations of variable neural modules. We argue that discrepancy in perception and prediction performance can be compensated by strong reasoning systems and is no longer the bottleneck of our system. Notably, unlike conventional frameworks which need retraining upon any module change, attributed to the flexibility of Agent-Driver, all neural modules in our system can be displaced in a plug-and-play manner, indicating our system’s compatibility.
Compatibility with different LLMs. We tried leveraging the Llama-2-7B , gpt-3.5-turbo-1106, and gpt-3.5-turbo-0613 models as the foundation LLMs in our system. Table 4 demonstrates that Agent-Driver powered by different LLMs can yield satisfactory performances, verifying the compatibility of our system with diverse LLM architectures.
6 Stability
LLMs typically suffer from arbitrary predictions—they might produce invalid outputs (e.g., hallucination or invalid formats)— which is detrimental to driving systems. To investigate this effect, we conducted a stability test of our Agent-Driver. Specifically, we used different amounts of training data to instruct the LLMs in our system, and we tested the number of invalid outputs during inference on the validation set. As shown in Table 5, Agent-Driver exposed to only 1% of the training data sees zero invalid output during inference of 6,019 validation scenarios, suggesting that our system attains high output stability with proper instructions.
7 Empirical Study
Effectiveness of system components. Table 3 shows the results of ablating different components in Agent-Driver. All variants utilize 10% training data for instructing the LLMs. From ID 1 to ID 5, we ablate the main components in Agent-Driver, respectively. We deactivate the self-reflection module and directly evaluate the trajectories output from LLMs to better assess the contribution of each other module. When the tool library is disabled, all perception results form the input to Agent-Driver without selection, which yields 2 times more input tokens and harms the system’s efficiency. The removal of the tool library also increases the collision rate, indicating the effectiveness of this component. In addition, the collision rate gets worse by removing the commonsense and experience memory, reasoning, and task planning modules, demonstrating the necessity of these components. Besides, we further note that self-reflection also greatly reduces the collision rate.
In-context learning vs. fine-tuning. Two prevalent strategies to instruct an LLM for novel tasks are in-context learning and fine-tuning. To determine which is the most effective strategy, we apply these two strategies to the LLMs of the chain-of-thought reasoning, task planning, and motion planning modules respectively, benchmarking them on the downstream motion planning performance. As indicated in Table 6, in-context learning performs slightly better than fine-tuning in collision rates for reasoning and task planning, suggesting that in-context learning is a favorable choice in these modules. In motion planning, the fine-tuning strategy significantly outperforms in-context learning, demonstrating the necessity of fine-tuning LLMs in motion planning.
Conclusion
This work introduces Agent-Driver, a novel human-like paradigm that fundamentally transforms autonomous driving pipelines. Our key insight is to leverage LLMs as an agent to schedule different modules in autonomous driving. On top of the LLMs, we propose a tool library, a cognitive memory, and a reasoning engine to bring human-like intelligence into driving systems. Extensive experiments on the real-world driving dataset substantiate the effectiveness, few-shot learning ability, and interpretability of Agent-Driver. These findings shed light on the potential of LLMs as an agent in human-level intelligent driving systems. For future works, we plan to optimize the LLMs for real-time inference.
Acknowledgments
We’d like to acknowledge Xinshuo Weng for fruitful discussions. We also acknowledge a gift from Google Research.
References
Related Works
Perception-Prediction-Planning in Driving Systems. Modern autonomous driving systems rely on a perception-prediction-planning paradigm to make driving decisions based on sensory inputs. Perception modules aim to recognize and localize objects in a driving scene, typically in a format of object detection or object occupancy prediction . Prediction modules aim to estimate the future motions of objects, normally represented as predicted trajectories or occupancy flows . Planning modules aim to derive a safe and comfortable trajectory, using rules or learning from human driving trajectories . These three modules are generally performed sequentially, either trained separately or in an end-to-end manner . This perception-prediction-planning framework overly simplifies the human driving process and cannot effectively incorporate human priors such as common sense and past driving experiences. By contrast, our Agent-Driver transforms the conventional perception-prediction-planning framework by introducing LLMs as an agent to bring human-like intelligence into the autonomous driving system.
LLMs in Autonomous Driving. Trained on Internet-scale data, LLMs have demonstrated remarkable capabilities in commonsense reasoning and natural language understanding. How to leverage the power of LLMs to tackle the problem of autonomous driving remains an open challenge. GPT-Driver handled the planning problem in autonomous driving by reformulating motion planning as a language modeling problem and introducing fine-tuned LLMs as a motion planner. DriveGPT4 proposed an end-to-end driving approach that leverages Vision-Language Models to directly map sensory inputs to actions. DiLu introduced a knowledge-driven approach with Large Language Models. These methods mainly focus on an individual component in conventional driving systems, e.g. question-answering , planning , or control . Some approaches are implemented and evaluated in simple simulated driving environments. By contrast, Agent-Driver presents a systematic approach that leverages LLMs as an agent to schedule the whole driving system, leading to a strong performance on the real-world driving benchmark.
Tool Library
In this section, we will first introduce the detailed descriptions of all functions in the tool library (Section 2.1). Next, we will provide a detailed example of how the agent interacts with the tool library (Section 2.2).
We include all function definitions in the tool library in Tables 2 and 3. The proposed functions cover detection, prediction, occupancy, and mapping, and enable flexible and versatile environmental information collection.
2 Tool Use
A detailed example of how the LLM leverages tools to collect environmental information is shown in Figures 1 and 2. System prompts are shown in blue, the response of LLMs is shown in green, and the collected data is shown in orange. System prompts provide sufficient context and guidance for instructing the LLM to dynamically invoke the functions in the tool library to collect necessary environmental information.
Cognitive Memory
In this section, we detail the data format and retrieving process of the commonsense and experience memory.
As shown in Figure 5, the commonsense memory consists of essential knowledge for safe driving, which is cached in a text-based format and is fully configurable.
We build the experience memory by caching the environmental information of driving scenarios and corresponding driving trajectories in the training set. Please note that this experience memory can be editable online, meaning that expert demonstrations conducted by human drivers can be easily inserted into the memory and benefit the subsequent decision-making of Agent-Driver.
2 Memory Search
We propose a two-stage searching strategy for retrieving the most similar past driving scenario to the query scenario.
Finally, top-K samples with the K highest similarity scores are selected as candidates for the second-stage search.
In the second stage, we propose an LLM-based fuzzy search. A detailed example of this stage is illustrated in Figures 3 and 4. The top-K past driving scenarios selected in the first stage are provided to the LLM. Then, the LLM is tasked to understand the text descriptions of these scenarios and determine the most similar past driving scenario to the query scenario. The driving trajectory corresponding to the selected scenario is also retrieved for reference.
With the proposed vector search and LLM-based fuzzy search, Agent-Driver can effectively retrieve the most similar past driving experience. The past experience and driving decision could help the current decision-making process.
Reasoning Engine
In this section, we provide detailed information on the workflow of the reasoning engine. The reasoning engine takes environmental information and memory data as inputs, performs chain-of-thought reasoning, task planning, motion planning, and self-reflection, and eventually generates a driving trajectory for execution. Figures 6, 7, and 8 show an example of how the reasoning engine works. We denote the input environmental information as and the retrieved memory data as . Please note that and are in the text format.
Chain-of-thought reasoning aims to emulate the human reasoning process and generate text-based reasoning results , which can be formulated as:
where is a LLM. To avoid arbitrary reasoning outputs of LLMs which might lead to hallucination and results not relevant to planning, we constrain to contain two essential parts: notable objects and potential effects. Specifically, we first instruct the LLM to identify those notable objects that have critical impacts on decision-making from the input environmental information. Then, we instruct the LLM to assess how these notable objects will influence the subsequent decision-making process. The instruction can be established by two strategies: in-context learning and fine-tuning.
For in-context learning, each time we leverage two human-annotated examples of notable objects and potential effects in addition to and collectively as inputs to the LLM:
For fine-tuning, we auto-generate the reasoning targets leveraging the technique proposed in . Then we fine-tune the LLM to make its reasoning outputs approaching the targets .
Both in-context learning and fine-tuning effectively reduce invalid outputs. As shown in Table 6 of the main paper, compared to the fine-tuning strategy, in-context learning enables the LLMs to generate more diverse reasoning outputs, and results in better motion planning performance.
2 Task Planning
Task planning aims to generate high-level driving plans for autonomous vehicles, taking the reasoning results as well as the environmental information and memory data as inputs. The process can be formulated as
We define the driving plan as a combination of discrete driving behaviors and speed estimations. In this paper, we proposed 6 discrete driving behaviors: move_forward, change_lane_to_left, change_lane_to_right, turn_left, turn_right, and stop. We also propose 6 speed estimations: constant_speed, deceleration, quick_deceleration, deceleration_to_zero, acceleration, quick_acceleration. The combinations of driving behaviors and speed estimations result in 31 different driving plans (stop has no speed estimation). These driving plans can cover most driving scenarios and they are fully configurable, which means that we can add more behavior and speed types to cover those long-tailed scenarios. See Table 1 for details.
Similar to the reasoning module, in-context learning and fine-tuning can also be applied to instruct the LLM to generate driving plans. As shown in Table 6 of the main paper, in-context learning is more appropriate for instructing the LLM for task planning.
3 Motion Planning
Motion planning aims to plan a safe and comfortable driving trajectory , with the driving plan , reasoning results , environmental information , and memory data as inputs. The process can be formulated as
The planned trajectory can be represented as 6 waypoint coordinates in 3 seconds: . A recent finding suggests a fine-tuned LLMs can generate text-based coordinates quite accurately. That is, the LLM can generate a text string “(1.23, 0.32)” representing a coordinate, and this can be easily transformed back into its numerical format for subsequent execution. In particular, given a trajectory , we first transform it into a sequence of language tokens using a tokenizer :
With these language tokens, we then reformulate motion planning as a language modeling problem:
where and are the language tokens of the planned trajectory from the LLM and the human driving trajectory respectively. By learning to maximize the occurrence probability of the tokens derived from the human driving trajectory , the LLM can generate human-like driving trajectories. We suggest readers refer to for more details.
With proper fine-tuning, the LLM is able to generate a text-based trajectory that can be further transformed into its numerical format. This step maps natural-language-based perception, memory, and reasoning into executable driving trajectories, enabling our agent to perform low-level actions.
4 Self-Reflection
Self-reflection is designed to reassess and rectify the driving trajectory planned by the LLM. Specifically, the collision_check function is first invoked to check the collision of utilizing the estimated occupancy map. Specifically, we place the ego-vehicle at each waypoint in a trajectory, and then we will mark the trajectory as collision if there is an obstacle within a safe margin to the ego-vehicle in the occupancy map. If a trajectory is not marked with collision, we directly use this trajectory as output without further rectification. Otherwise, for those trajectories that have collisions, we leverage an optimization approach to rectify the trajectory into a collision-free one . Following , we sample the obstacle points near each waypoint at timestep in the occupancy map. Then, an optimization problem is formulated and solved through the Newton iteration method:
where and are hyperparameters. The first term regulates the optimized trajectory to be similar to the original , and the second term pushes the waypoint in the trajectory away from the obstacle points for each timestep .
Attributing to self-reflection, those unreliable decisions made by the LLM can further be corrected, and collisions in the planned trajectories can be effectively mitigated.
Experiments
We employ gpt-3.5-turbo-0613 as the foundation LLM in our Agent-Driver. We leverage the same LLM for every task except motion planning, and we fine-tuned another LLM specially for motion planning.
For tool use and memory search, the LLM is guided by system prompts without fine-tuning or exemplar-based in-context learning. In chain-of-thought reasoning and task planning, the LLM is instructed by two randomly selected exemplars derived from the training set. This approach encourages the models to develop various insightful chain-of-thought processes and detailed plans for tasks independently. Please note that we utilize the original pre-trained LLM with different system prompts and user input for the above tasks. In motion planning, we fine-tune an LLM with human driving trajectories in the training set for only one epoch.
2 Evaluation Metrics
In this section, we provide more details about the evaluation metrics used in this work. The output trajectory is formatted as 6 waypoints in a 3-second horizon, i.e., . For in each driving scenario, the L2 error is computed as:
In the UniAD metric , the L2 error at the -th second () is reported as the error at this timestep:
The average L2 error is then computed by averaging of the timesteps.
In the ST-P3 metric and following works , the L2 error at the -th second is reported as the average error from to second:
The average L2 error is computed by averaging the of the timesteps again (average over average in other words).
Similarly, UniAD reports the collision at the -th second () as , while ST-P3 reports as the average from to second:
In addition to the differences in calculation methods, there is also a difference in how the ground truth occupancy maps are generated in the two metrics. Specifically, UniAD only considers the vehicle category when generating ground truth occupancy, while ST-P3 considers both the vehicle and pedestrian categories. This difference leads to varying collision rates for identical planned trajectories when measured by these two metrics, yet it does not impact the L2 error.
In this paper, we faithfully evaluated our approach and the baseline methods using the officially implemented evaluation metrics in the two papers , ensuring a completely fair comparison with other methods.
3 Limitations
Due to the limitations of the OpenAI APIs, we are unable to obtain the accurate inference time of our Agent-Driver. Thus it remains uncertain whether our approach can meet the real-time demands of commercial driving applications. However, we argue that recent advances in accelerating LLM inference shed light on a promising direction to enable LLMs for real-time applications. We believe the inference problem of LLMs will be resolved in the future.
Qualitative Analysis
Figure 9 provides more examples to show the interpretablility of Agent-Driver. Figures 10, 11 and 12 visualize critical objects identified by Agent-Driver and the planned driving trajectories, from which we can see our system progressively identifies critical objects via tool use and reasoning, and eventually plans a safe trajectory for driving. Driving videos are shown in Figures 13, 14, and 15. The qualitative results verify the effectiveness and interpretability of our Agent-Driver.
2 Failure Cases
We provide failure cases in Agent-Driver in Figure 16. We observed that heading errors of large objects, e.g., buses, have a critical impact on motion planning, which indicates the importance of accurate heading prediction in detection networks.