Human-Centric Autonomous Systems With LLMs for User Command Reasoning
Yi Yang, Qingwen Zhang, Ci Li, Daniel Simões Marta, Nazre Batool, John Folkesson
Introduction
Autonomous driving (AD) has experienced significant advances in recent years , enhanced by advancements in both hardware and machine learning techniques. The field has seen diverse developments: multi-task learning approaches over traditional modular systems , end-to-end methodologies with planning-oriented techniques , and widespread adoption of Bird-Eye-View representations . Many explorations have been conducted towards safe and robust autonomous systems.
Nonetheless, several aspects are often overlooked, such as taking into consideration human intent when designing AD systems . The future of AD should integrate human-centered design with advanced AI capabilities to reason and interpret the user’s intent . This integration offers numerous advantages e.g. language interaction , driving scene understanding , contextual reasoning , as well as explainability and trust. In a recent milestone survey, Chen et al. emphasize the necessity of considering human behaviors and AD systems to ensure communication transparency and efficiency. They also hint towards a human-machine hybrid intelligence approach, asserting that the reliability of intelligent systems depends on incorporating learnability and the influence of human intelligence and mentorship.
A current relevant challenge stems from allowing humans to interact using natural language with AD systems , aiming to adhere to both human preferences and intent. Our contribution here is a series of insights and experiments on a variety of LLM models and prompt designs exploring the interaction between users and AD systems via verbal commands.
While the use of natural language inherently poses a challenge in itself, we leverage recent advancements in large pre-trained foundational models alongside their zero-shot capabilities to adapt to new tasks . In our approach, we adopt a few-shot learning strategy, keeping an LLM frozen and enhancing reasoning through sequential prompting, drawing inspiration from previous works . Ultimately, our goal is to condition an LLM for the task of multivariate binary classification of AD systems based on in-cabin user commands.
Related Work
Few-shot Reasoning Capabilities of LLMs. Recent research indicates that LLMs with sufficient expansiveness possess the ability to perform sophisticated reasoning tasks . In this work, we adopt an approach somewhat akin to Socratic models , where models–encompassing vision, language, and sound–can operate in a zero-shot or few-shot fashion. A prevalent strategy in this type of approach involves keeping certain model subcomponents static, particularly those relevant to a single modality, preserving them for use in subsequent tasks . This methodology interchangeably aligns well with few-shot transfer learning , wherein an LLM, originally trained on a vast array of internet-scale text prompts for various tasks such as text completion and sentiment analysis, is repurposed for a new, specific task. In our case, this involves leveraging general domain knowledge of AD for a classification challenge . However, the success of adapting language models to new tasks largely depends on the prompting strategy and the quality of the prompts . Wei et al. demonstrated that standard prompting is often insufficient. They introduced a chain-of-thought prompting strategy, where a sequence of thoughtful demonstrations significantly improves performance in tasks such as symbolic reasoning, commonsense, and arithmetic. Inspired by the above works, we aim to keep the weights of an LLM frozen and design a chain-of-thought strategy for a downstream task of classification from in-cabin user commands.
Human-centric Autonomous Driving. Incorporating human intent in the form of natural language within AD is a relatively new and active field of research. The emergence of recent datasets and works has significantly boosted its development. Datasets like NuPrompt , NuScenes-QA , DRAMA , enhance the autonomous driving datasets with provided texts, for different tasks including object tracking, visual question answering, image caption and so on. As many works have shown promising results, a possible goal would be to mediate planning-oriented AD , which includes perception, prediction and planning, with human intent. In recent exploratory work, there is verified evidence for the effectiveness of integrating LLMs into human-centric AD . In a survey by Zhou et al. , large language and vision models show potential to contribute to the AD system in different submodules including perception and understanding, navigation and planning, decision-making and control, as well as end-to-end pipelines. For instance, as explored by Fu et al. , understanding the common sense driving intentions embedded in human commands is achieved by utilizing the reasoning capacities of LLMs in addressing long-tailed cases. This type of approach could be instrumental in the evolution of AD technologies that mirror human-like driving nuances. Ding et al. use multimodal LLM to inform AD system the localization of risk objects and provide suggestions for safety and robustness. More closely related to our work is that reported by Cui et al. and Jain et al. , demonstrating that utilizing linguistics and understanding from verbal commands can improve driving decisions and enhance personalized driving experiences through ongoing verbal feedback. In this exploratory work, we aim to holistically address some of the weaknesses posed when handling long-tail data or out-of-distribution driving scenarios by tapping into the overall implicitly acquired domain knowledge of LLMs .
Method
Given user commands as input text sentences, the model is expected, through reasoning and sentiment analysis, to provide binary results (‘Yes’ / ‘No’) for 8 specific questions. These questions inquire about the command’s diverse requirements- ranging from perception, in-cabin monitoring, localization, control, network access, and entertainment to human privacy and traffic laws.
Since LLMs undergo training on extensive datasets, integrating them within autonomous systems significantly enhances AD systems’ capability to grasp both scene dynamics and user intent. In this paper, we utilize both online (from the GPT series ) and offline (from the Llama series ) LLMs. As demonstrated in Tab. 1, we begin by informing the LLM that its task is to engage in a question and answer (Q&A) format by responding with ‘Yes’ or ‘No’. By instructing the model with an in-depth explanation for each of the eight questions, it is prompted to process information step by step. To further optimize accuracy, the LLMs are provided with few examples of in-context few-shot learning . Finally, conditioned on the real user command, LLMs are instructed to produce both step-by-step explanations and results in a predetermined format.
Experiment
Datasets and Metrics: We assess the performance of LLMs on 1,099 in-cabin user commands from the UCU Dataset . The details are available on the challenge websitehttps://llvm-ad.github.io/challenges/. The evaluation involves determining if a command requires any of the eight specified modules’ help to achieve autonomy. Official evaluation metrics include accuracy at the question level (accuracy for each individual question) and at the command level (accuracy is only acknowledged if all questions for a particular command are correctly identified).
LLM Models and Baseline: We benchmark our prompt approaches using a range of LLMs, including GPT-3.5-turbo / GPT-4 https://platform.openai.com. Note that during the time we used the models, GPT-3.5-turbo pointed to GPT-3.5-turbo-0613 version, and GPT-4 pointed to GPT-4-0613 version., and CodeLlama-34b-Instruct / Llama-2-70b-Chat . The GPT models are accessed online, whereas Llama models offer an on-board solution, with both their code and pretrained weights open-sourced. CodeLlama-34b-Instruct is an enhanced version of the original Llama additionally trained on code generation and instruction problems with 34 billion parameters in the model. Llama-2-70b-Chat has roughly double the number of parameters and is trained mainly on conversational interactions enhanced with human reinforcement learning.
We compare the LLMs’ performance against two baselines. The first employs a simple random guessing strategy. The second is a rule-based method that identifies specific keywords in the user command to determine system requirements. It is noteworthy that these rules are automatically generated by ChatGPT-4, augmented with the Advanced Data Analysis plug-in.
Evaluation: The results are presented in Tab. 2. Compared to the random guess and rule-based approaches, LLMs exhibit notably higher accuracy, especially at the command level. Among the LLMs tested, the GPT series surpasses the Llama models, where GPT-4 achieves a peak accuracy of 89.02% at the question level and 38.03% at the command level. Diving into question-specific accuracy, GPT consistently behaves excellently in determining if a command requires perception, localization, entertainment submodules, or if it might violate traffic laws. However, for the in-cabin monitoring, GPT’s predictions display some inconsistency and lower accuracy, especially for GPT-4 (74.89%). It is observed that in the evaluation of 1,099 commands, GPT-4 incorrectly responded to 276 queries related to in-cabin monitoring. A majority proportion of these errors, amounting to 259 instances, involved GPT-4 saying “Yes, it requires in-cabin monitoring” when the ground truth doesn’t, with only a few errors being the reverse. To further investigate why GPT-4 says yes often, 2 keywords are found in its explanations: multimedia and alert system. Both terms are typically associated with in-cabin monitoring in our given instructed prompts. Specifically, 155 commands (56%) are related to multimedia - GPT-4 identifies these as requiring in-cabin multimedia activities, such as making calls, playing radio or videos, or screen displays. Additionally, in 26% of the instances (71 commands), GPT-4 associates the command with the utilization of in-cabin monitoring systems for alerting or notifying the vehicle’s occupants. Given this context, it’s difficult to conclusively label GPT-4’s responses as incorrect. This is largely due to the vague definition of what exactly constitutes in-cabin monitoring in the context of our study.
Ablation Study: We perform two ablation studies to explore the impacts of providing a detailed explanation of instructions and the given number of shot examples. To reduce costs, we run experiments on GPT-3.5 with a small test subset that matches the distribution of the original test data created by ChatGPT-4. More specifically, ChatGPT-4 is asked to analyze the data distribution for each question and sample 100 commands according to it.
The performance of combining three prompting methods is shown in Tab. 3. Note that for providing a detailed instruction explanation, we experimented with two formats: a step-by-step approach as illustrated in Tab. 1, and a consolidated paragraph format (labeled as ‘w/o step’ in Tab. 3). For the latter case, we omit the word Step:# which causes less structured prompts. Each prompting method exhibits enhanced performance, with the few-shot method outperforming two detailed explanation methods. Notably, combining one of the prompting methods from detailed explanation and the few-shot method significantly augments the performance.
Tab. 4 shows the impact of the number of given few shot examples with the step-by-step detailed explanation method. We observe a notable increase when two examples are provided, which indicates the importance of providing shot examples for performance gain. However, as the number of examples continues to increase, the impact on the final results becomes subtle, suggesting the existence of a saturation threshold or limitations.
Conclusion
In this paper, we offered a human-centric perspective by providing several key insights with different prompting designs to condition LLMs towards achieving AD system requirements from verbal user commands. Our work included an analysis of several online and offline popular LLMs with a variety of ablation studies. We hope that this work inspires our community to consider human intent in the design of an AD system. Future work points towards the direction of incorporating human feedback in a more intelligent and effective way. This would pave the way for the development of AD systems that are not only more reliable and capable in their reasoning and comprehension skills but also more closely aligned with human-like preferences and behaviors. The goal is to establish an AD system with human-centric values, thereby fostering greater trust and acceptance among users.
Acknowledgement
Thanks to RPL’s members: Xuejiao Zhao and Boyue Jiang, who gave constructive comments on this work. We also thank the anonymous reviewers for their constructive comments.
This workWe have used ChatGPT for editing and polishing author-written text. was funded by Vinnova and Wallenberg AI, Autonomous Systems and Software Program (WASP), Sweden (research grant). The computations were enabled by the supercomputing resource Berzelius provided by National Supercomputer Centre at Linköping University and the Knut and Alice Wallenberg foundation, Sweden.
References
Appendix A Qualitative Results of GPT4
There are three examples of the GPT4 real results with explanations for different user commands in Tab. 5, Tab. 6, Tab. 7.