A Survey on Multimodal Large Language Models for Autonomous Driving

Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, Tianren Gao, Erlong Li, Kun Tang, Zhipeng Cao, Tong Zhou, Ao Liu, Xinrui Yan, Shuqi Mei, Jianguo Cao, Ziran Wang, Chao Zheng

Introduction

GPT-4V Answer:

Input LiDAR Point Cloud: [163]

Question / Prompt:

GPT-4V Answer:

Input Driving Front View:

Question / Prompt:

GPT-4 Code Genration:

Simulation [92]:

Development of Autonomous Driving

The quest for autonomous driving has been a progressive journey, marked by a continuous interplay between visionary aspirations and technological capabilities. The first wave of comprehensive research on autonomous driving started in the late 20th century. For example, the Autonomous Land Vehicle (ALV) project launched by Carnegie Mellon University utilized sensor readings from stereo cameras, sonars, and the ERIM laser scanner to perform tasks like lane keeping and obstacle avoidance . However, these researches were constrained by limited sensor accuracy and computation capabilities.

The last two decades have seen rapid improvements in autonomous driving systems. A classification system published by the Society of Automotive Engineers (SAE) in 2014 defined six levels of autonomous driving systems . The classification method has now been widely acknowledged and illustrated important milestones for the research and development progress. The introduction of Deep Neural Networks (DNNs) has also played a significant role . Backed by deep learning, computer vision has been crucial for interpreting complex driving environments, offering state-of-the-art solutions for problems such as object detection, scene understanding, and vehicle localization . Deep Reinforcement Learning (DRL) has additionally played a pivotal role in enhancing the control strategies of autonomous vehicles, refining motion planning, and decision-making processes to adapt to dynamic and uncertain driving conditions ,. Moreover, sensor accuracy and computation power improvements allow larger models with more accurate results to be run on the vehicle. With such improvements, More L1 to L2 level Advanced Driver Assistance Systems (ADAS) like lane centering and adaptive cruise control are now available on everyday vehicles . Companies like Waymo, Zoox, Cruise, and Baidu are also rolling out Robotaxis with Level 3 or higher autonomy. Nevertheless, such autonomous systems still fail in many driving edge cases such as extreme weather, bad lighting conditions, or rare situations .

Inspired by current limitations, part of the research on autonomous driving is now focusing on addressing the safety of autonomous systems and enhancing the safety of autonomous systems . As Deep Neural Networks are often considered black boxes, trustworthy AI aims at making the system more reliable, explainable, and verifiable. For example, generating adversarial safety-critical scenarios for training autonomous driving systems such that the system is more capable of handling cases with low probability . Another way to improve the overall safety is through vehicle-to-infrastructure and vehicle-to-vehicle communication. With information from nearby instances, the system will have improved robustness and can receive early warnings . Meanwhile, as Large Language Models show their powerful reasoning and scene-understanding capability, research is being conducted to utilize them to improve the safety and overall performance of the autonomous driving system.

Development of Multimodal Language Models

The development of language models has been a journey marked by significant breakthroughs. Since the early 1960s, many linguists, most renowned Noam Chomsky, attempted to model natural languages . Early efforts focused mainly on rule-based approaches . However, in the late 1980s and early 1990s, the spotlight shifted onto statistic models, such as N-gram , hidden Markov models , which relied on counting the frequency of words and sequences in text data. The 2000s witnessed the introduction of neural networks into natural language modeling. Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks were used for various NLP tasks.

Despite their potential, early neural models had limitations in capturing long-range dependencies and struggled with complex language tasks. In 2013, Tomas Mikolov and his team at Google introduced Word2Vec , a groundbreaking technique for representing words as dense vectors, providing a better understanding of semantic relationships between words. This laid down the foundation for the rise of deep learning , which eventually led to the pivotal work, Attention is all you need , which kick-started the new era of large language models..

2 Advancements in Large Language Models

LLMs are a category of Transformer-based language models known for their extensive number of parameters, often numbering in the hundreds of billions. These models are trained on vast amounts of internet data, which enables them to perform a wide range of language tasks, primarily through text generation. Some well-known examples of LLMs include GPT-3 , PaLM , LLaMA , and GPT-4 . One of the most notable characteristics of LLMs is their emergent abilities, such as in-context learning (ICL) , instruction following , and reasoning with chain-of-thought (CoT) .

There is a growing area of research that utilizes LLMs to develop autonomous agents with human-like capabilities. These agents leverage the extensive knowledge stored in pre-trained LLMs to create coherent action plans and executable policies . Embodied language models directly integrate real-world sensor data with language models, establishing a direct connection between words and perceptual information. Voyager introduces lifelong learning by incorporating three main components: an automatic curriculum that promotes exploration, a skill library to store and retrieve complex behaviors, and an iterative prompting mechanism to generate executable code for embodied control. Voxposer utilizes LLMs to generate robot trajectories for a wide range of manipulation tasks, guided by open-ended instructions and objects.

In parallel with these advancements, the use of LLMs in the field of autonomous driving is gaining momentum. Recent research has investigated the application of LLMs to comprehend driving environments. These studies have demonstrated the impressive ability of LLMs to handle complex scenarios by converting visual information into text representation, enabling LLMs to interpret the surrounding world. Similarly, in RRR , authors propose a human-centric autonomous driving framework that breaks down user commands into a series of intermediate reasoning steps, accompanied by a detailed list of action descriptions to accomplish the objective.

3 Early Efforts in Modality Fusion

Over the past few decades, the fusion of various modalities such as vision, language, video, and audio has been a key objective in artificial intelligence (AI). Initial efforts in this domain focused on simple tasks, such as image or video captioning and text-based image retrieval, which were mostly rule-based and relied on hand-crafted features. A classic example of early AI problems in the 1970s and 1980s was the ”Blocks World” , where the goal was to rearrange colored blocks on a table based on textual instructions. This early attempt bridged vision (understanding block configurations) with language (interpreting and executing instructions), even though it was not based on deep learning.

4 Advancements in Vision-Language Models

In the following years, the field of multimodal models saw significant advancements. Over the last decade, the advent of deep learning has revolutionized approaches to visual-language tasks. Convolutional Neural Networks (CNNs) became the de facto standard for image and video processing, while Recurrent Neural Networks (RNNs) emerged as the go-to models for processing sequential data, such as natural languages. During this period, popular tasks included image and video captioning, which involves generating descriptive sentences for images and videos, and visual question answering (VQA), where models answer questions related to visual data. Typical vision-language models employed joint embeddings, with image features (processed by CNNs) and text features (processed by RNNs or Transformers ) mapped to a shared semantic space to facilitate multimodal learning . Beyond vision and language, researchers also proposed models for other modalities, such as audio, speech, and 3D data. For instance, Mroueh et al. (2015) developed a deep multimodal learning model for audio-visual speech recognition that utilizes CNNs for visual data and RNNs for audio data . Arandjelović and Zisserman (2017) explored the relationship between visual and auditory data by developing a model that learns shared representations from unlabeled videos, using CNNs for both image and audio processing . Furthermore, Qi et al. (2016) introduced models that process 3D data, including point clouds, for object classification tasks, employing CNNs to learn representations from volumetric data and multiple 2D views of 3D objects . These works highlight the potential of multimodal learning in capturing complex relationships between different types of data, leading to richer and more accurate representations.

5 Pre-Training and Multimodal Transformers

Building on this momentum, the field of multimodal models has continued to evolve, with researchers exploring the potential of pre-training multimodal models on extensive datasets before fine-tuning them on specific tasks. This approach has resulted in significant performance improvements across a range of applications. Inspired by the success of pre-trained NLP models like BERT , T5 , and GPTs , researchers developed multimodal Transformers that can process cross-modality inputs such as text, image, audio, pointcloud . Notable examples of visual-language models include CLIP , ViLBERT , VisualBERT , SimVLM , BLIP-2 and Flamingo , which were pre-trained on large-scale cross-modal datasets comprising images and languages. Other works have explored the use of multimodal models for tasks such as video understanding , audio-visual scene understanding , and even 3D data processing . Pre-training allows the models to align different modalities and enhance the representation learning ability of the model encoder. By doing so, these models aim to create systems that can generalize across tasks without the need for task-specific training data. Furthermore, the evolution of multimodal models has also given rise to new and exciting possibilities. For instance, DALL-E extends the GPT-3 architecture to generate images from textual descriptions, Stable Diffusion and ControlNet utilized CLIP and UNet-based diffusion model to generate images controlled by text prompt. They showcase the potential for using multimodal models in many application scenarios such as healthcare , civil engineering , robotics and, art .

6 Emergence of Multimodal Large Language Models

Recently, MLLMs have emerged as a significant area of research. These models leverage the power of LLMs, such as ChatGPT , InstructGPT , FLAN , and OPT-IML to perform tasks across multiple modalities such as text and images. They exhibit surprising emergent capabilities, such as writing stories based on images and performing OCR-free math reasoning, which are rare in traditional methods. This suggests a potential path to artificial general intelligence. Key techniques and applications in MLLMs include Multimodal Instruction Tuning, which tunes the model to follow instructions across different modalities ; Multimodal In-Context Learning, which allows the model to learn from the context of multimodal data ; Multimodal Chain of Thought, which enables the model to maintain a chain of thought across different modalities ; and LLM-Aided Visual Reasoning (LAVR), which uses LLMs to aid in visual reasoning tasks . MLLMs are more in line with the way humans perceive the world, offering a more user-friendly interface and supporting a larger spectrum of tasks compared to LLMs. The recent progress of MLLMs has been ignited by the development of GPT-4V , which, despite not having an open multimodal interface, has shown amazing capabilities. The research community has made significant efforts to develop capable and open-sourced MLLMs, exhibiting surprising practical capabilities.

Multimodal Language Models for Autonomous Driving

In the autonomous driving industry, MLLMs have the potential to understand traffic scenes, improve the decision-making process for driving, and revolutionize the interaction between humans and vehicles. These models are trained on vast amounts of traffic scene data, allowing them to extract valuable information from different sources like maps, videos, and traffic regulations. As a result, they can enhance a vehicle’s navigation and planning, ensuring both safety and efficiency. Additionally, they can adapt to changing road conditions with a level of understanding that closely resembles human intuition.

Traditional perception systems are often limited in their ability to recognize only a specific set of predefined object categories. This restricts their adaptability and requires the cumbersome process of collecting and annotating new data to recognize different visual concepts. As a result, their generality and usefulness are undermined. In contrast, a new paradigm is emerging that involves learning from raw textual descriptions and various modalities, providing a richer source of supervision.

Multimodal Large Language Models (MLLMs) have gained significant interest due to their proficiency in analyzing non-textual data like images and point clouds through text analysis . These advancements have greatly improved zero-shot and few-shot image classification , segmentation , and object detection .

Pioneering models like CLIP have shown that training to match images with captions can effectively create image representations from scratch. Building on this, Liu et al. introduced LLaMa , which combines a vision encoder with an LLM to enhance the understanding of both visual and linguistic concepts. Zhang et al. further extended this work with Video-LLaMa , enabling MLLMs to process visual and auditory information from videos. This represents a significant advancement in machine perception by integrating linguistic and visual modalities.

Furthermore, researchers have explored the use of vectorized visual embeddings to equip MLLMs with environmental perception capabilities, particularly in autonomous driving scenarios. DriveGPT4 interprets video inputs to generate driving-related textual responses. HiLM-D focuses on incorporating high-resolution details into MLLMs, improving hazard identification and intention prediction. Similarly, Talk2BEV leverages pre-trained image-language models to combine Bird’s Eye View (BEV) maps with linguistic context, enabling visuo-linguistic reasoning in autonomous vehicles.

At the same time, progress in autonomous driving is not limited to discriminative perception models; generative models are also gaining popularity. One example is the Generative AI for Autonomy model (GAIA-1), which generates realistic driving scenarios by integrating video, text, and action inputs. This generative world model can anticipate various potential outcomes based on the vehicle’s maneuvers, showcasing the sophistication of generative models in adapting to the changing dynamics of the real world . Similarly, UniSim aims to replicate real-world interactions by combining diverse datasets, including objects, scenes, actions, motions, language, and motor controls, into a unified video generation framework. Moreover, the Waymo Open Sim Agents Challenge (WOSAC) is the first public challenge to develop simulations with realistic and interactive agents.

2 Multimodal Language Models for Planning and Control

The use of language in planning and control tasks has a longstanding history in robotics, dating back to the use of lexical parsing in natural language for early demonstrations of human-robot interaction , and it has been widely studied being used in the robotics area. There exists comprehensive review works on this topic . It has been well-established that language acts as a valuable interface for non-experts to communicate with robots . Moreover, the ability of robotic systems to generalize to new tasks through language-based control has been demonstrated in various works . Achieving specific planning or control tasks or policies, including model-based , imitation learning , and reinforcement learning , has been extensively explored.

Due to the significant ability in zero-shot learning , in-context learning and reasoning , many works showed that LLMs could enable reasoning of planning and perceiving the environment with textual description to develop user in the loop robotics . broke down natural language commands into sequences of executable actions through a combination of text completion and semantic translation to control the robot. SayCan utilized weighted LLMs to produce reasonable actions and control robots while uses environmental feedback, LLMs can develop an inner monologue, enhancing their capacity to engage in more comprehensive processing within robotic control scenarios. Socratic Models employs visual language models to replace perceptual information within the language prompts used for robot action generation. introduces an approach that uses LLMs to directly generate policy code for robots to do control tasks, specify feedback loops, and write low-level control primitives.

In autonomous driving, LLMs could serve as the bridge to support human-machine interactions. For general purposes, LLMs can be task-agnostic planners. In , the authors discovered that pre-trained LLMs contain actionable knowledge for coherent and executable action plans without additional training. Huang et al. proposed the use of LLMs for converting arbitrary natural language commands or task descriptions into specific and detail-listed objectives and constraints. proposed integrating LLMs as decision decoders to generate action sequences following chain-of-thoughts prompting in autonomous vehicles. In , authors showcased that LLMs can decompose arbitrary commands from drivers to a set of intermediate phases with a detailed list of descriptions of actions to achieve the objective.

Meanwhile, it is essential to enhance the safety and explainable of autonomous driving. The multimodal language model provides the potential to comprehend its surroundings and the transparency of the decision process. showed that video-to-text models can help generate textual explanations of the environment aligned with downstream controllers. Deruyttere et al. compared baseline models and showed that LLMs can identify specific objects in the surroundings that are related to the commands or descriptions in natural language. For the explainability of the model, Xu et al. proposed to integrate LLMs to generate explanations along with the planned actions. In , the authors proposed a framework where LLMs can provide descriptions of how they perceive and react to environmental factors, such as weather and traffic conditions.

Furthermore, the LLMs in autonomous driving can also facilitate the fine-tuning of controller parameters, aligning them with the driver’s preferences and thus resulting in a better driving experience. integrates LLMs into low-level controllers through guided parameter matrix adaptation.

Besides the development of LLMs, great progress has also been witnessed in MLLMs. The MLLMs have the potential to serve as a general and safe planner model for autonomous driving. The ability to process and fuse visual signals such as images enhanced navigation tasks by combining visual cues and linguistic instructions . Interoperability challenges have historically been an issue for autonomous planning processes . However, recent advancements in addressing interoperability challenges in autonomous planning have leveraged the impressive reasoning capabilities of MLLMs during the planning phases of autonomous driving . In one notable approach, Chen et al. integrated vectorized object-level 2D scene representations into a pre-trained LLM with adapters, enabling direct interpretation and comprehensive reasoning about various driving scenarios. Additionally, Fu et al. employed LLMs for reasoning and translated this reasoning into actionable driving behaviors, showing the versatility of LLMs in enhancing autonomous driving planning. Additionally, GPT-Driver reformulated motion planning as a language modeling problem and utilized LLM to describe highly precise trajectory coordinates and its internal decision-making process in natural language in motion planning. SurrealDriver simulated MLLM-based generative driver agents that can perceive complex traffic scenarios and generate corresponding driving maneuvers. investigated the utilization of textual descriptions along with pre-trained language encoders for motion prediction in autonomous driving.

3 Industrial Applications

The integration of MLLMs in the autonomous driving industry has been developed by several significant initiatives. Wayve introduces LINGO-1, which enhances the learning and explainability of foundational driving models by integrating vision, language, and action . They also developed GAIA-1, a generative world model for realistic driving scenario generation, offering fine-grained control over vehicle behavior and scene features .

Tencent T Lab generated traffic, map, and driving-related context from their HD map AI system , creating MAPLM, a large map and traffic scene dataset for scene understanding.

Waymo’s contribution, MotionLM, improved motion prediction in multi-agent environments. By conceptualizing continuous trajectories as discrete motion tokens, it transfers multi-agent motion prediction to a language modeling task . This approach transforms the dynamic interaction of road agents into a manageable sequence-to-sequence prediction problem.

Research from the Bosch Center focuses on using natural language for enhanced scene understanding and predicting future behaviors of surrounding traffic . Meanwhile, researchers from the Hong Kong University of Science and Technology and Huawei Noah’s Ark Lab have leveraged MLLMs to integrate various autonomous driving tasks, including risk object localization and intention and suggestion prediction from videos .

These developments in industry illustrate the expanding role of MLLMs in enhancing the capabilities and functionalities of autonomous driving systems, marking a significant improvement in vehicle intelligence and situational awareness.

Datasets and Benchmarks

Publicly available datasets have played a crucial role in advancing autonomous driving technologies. Tab. 3 provides a comprehensive overview of the latest representative datasets for autonomous driving. In the past, datasets mainly focused on 2D annotations, like bounding boxes and masks, primarily for RGB camera images . However, achieving autonomous driving capabilities that can match human performance requires precise perception and localization in the 3D environment. Unfortunately, extracting depth information from purely 2D images poses significant challenges.

To enable robust 3D perception or mapping, researchers have created many multimodal datasets. These datasets include not only camera images but also data from 3D sensors like radar and LiDAR. An influential example in this field is the KITTI dataset , which provides multimodal sensor data, including front-facing stereo cameras and LiDAR. KITTI also includes annotations of 3D boxes and covers tasks such as 3D object detection, tracking, stereo, and optical flow. Subsequently, NuScenes and the Waymo Open dataset have emerged as representative multimodal datasets. These datasets set new standards by offering a large number of scenes. These datasets represent a significant advancement in the availability of large data for advancing research in autonomous driving.

2 Multimodal-Language Datasets for Traffic Scene

Several pioneering studies have explored language-guided visual understanding in driving scenarios. These studies either enhance existing datasets with additional textual information or create new datasets independently. The former category includes works such as Talk2Car , nuScenes-QA , DriveLM , and NuPrompt . Among these, Talk2Car stands out as the first object referral dataset, which contains natural language commands for autonomous vehicles. On the other hand, datasets like BDD-X and DRAMA were independently created. DRAMA specifically focuses on video and object-level inquiries regarding driving hazards and associated objects. This dataset aims to enable visual captioning through free-form language descriptions and uses both closed and open-ended responses to multi-tiered questions. It allows for the evaluation of various visual captioning abilities in driving contexts.

Despite the advancements in language comprehension in traffic scenes with MLLMs, their performance is still far below the human level. This is because traffic data-text pairs contain diverse modalities, such as 3D point clouds, panoramic 2D imagery, high-definition map data, and traffic regulations. These elements significantly differ from conventional domain contexts and question-answer pairs, highlighting the unique challenges of deploying MLLMs in that autonomous driving context. The datasets mentioned above are limited in terms of scale and quality, which hinders efforts to fully address these emerging challenges.

LLVM-AD Workshop Summary

The 1st LLVM-AD is held together with WACV 2024 on Jan 8th, 2024 in Waikoloa, Hawaii. we seek to bring together academia and industry professionals in a collaborative exploration of applying MLLMs to autonomous driving. Through a half-day in-person event, the workshop will showcase regular and demo paper presentations and invited talks from famous researchers in academia and industry. Additionally, LLVM-AD will launch two open-source real-world traffic language understanding datasets, catalyzing practical advancements. The workshop will host two challenges based on this dataset to assess the capabilities of language and computer vision models in addressing autonomous driving challenges.

Tencent’s THMA HD Map AI labeling system is utilized to create descriptive paragraphs from HD map labels, offering nuanced portrayals of traffic scenes . Participants worked with various data modalities, including 2D camera images, 3D point clouds, and Bird’s Eye View (BEV) images, enhancing our understanding of the environment. This innovative initiative explores the intersection of computer vision, AI-driven mapping, and natural language processing, highlighting the transformative potential of Tencent’s THMA technology in reshaping our understanding and navigation of our surroundings.

UCU Dataset.

The primary objective of this challenge is the development of algorithms that are proficient in understanding drivers’ commands and instructions represented as natural language input. These commands and instructions could encompass a diverse array of command types, ranging from safety-oriented instructions such as “engage the emergency brakes” or “adjust headlight brightness”, to driving operational instructions such as “shift to park mode” or “set the cruise control to 70 mph”, and comfort-related requests such as “turn up the AC” or “turn off seat heating”. The scope of commands can even be extended to vehicle-specific instructions like “open sunroof” or “enable ego mode”.

2 Workshop Summary

Nine papers were accepted in the inaugural Workshop on Large Language and Vision Models for Autonomous Driving (LLVM-AD) at the 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). They cover topics on MLLMs for autonomous driving focusing on integrating LLMs into user-vehicle interaction, motion planning, and vehicle control. Several papers explored the novel use of LLMs to enhance human-like interaction and decision-making in autonomous vehicles. For example, “Drive as You Speak” and “Drive Like a Human” presented frameworks where LLMs interpret and reason in complex driving scenarios, mimicking human behavior. “Human-Centric Autonomous Systems With LLMs” emphasized the importance of user-centric design, utilizing LLMs to interpret user commands. This approach represents a significant shift towards more intuitive and human-centric autonomous systems.

In addition to LLM integration, the workshop featured methodologies in vision-based systems and data processing. “A Safer Vision-based Autonomous Planning System for Quadrotor UAVs” and “VLAAD” demonstrated advanced approaches to object detection and trajectory planning, enhancing the safety and efficiency of UAVs and autonomous vehicles.

Optimizing technical processes was also a significant focus. For instance, “A Game of Bundle Adjustment” introduced a novel approach to improving 3D reconstruction efficiency, while “Latency Driven Spatially Sparse Optimization” and “LIP-Loc” explored advancements in CNN optimization and cross-modal localization, respectively. These contributions represent notable progress towards more efficient and accurate computational models in autonomous systems.

Furthermore, the workshop presented innovative approaches to data handling and evaluation. For example, NuScenes-MQA introduced a dataset annotation technique for autonomous driving. Collectively, these papers illustrate a significant stride in integrating language models and advanced technologies into autonomous systems, paving the way for more intuitive, efficient, and human-centric autonomous vehicles.

Discussion

Despite the success of LLMs in language understanding, applying them to autonomous driving presents a unique challenge. This is due to the necessity for these models to integrate and interpret inputs from diverse modalities, such as panoramic images, 3D point clouds, and HD map annotations. The current limitations in data scale and quality mean that existing datasets struggle to address all these challenges comprehensively. Furthermore, almost all multimodal LLMs like GPT-4V have been pre-trained on a wealth of open-source datasets including traffic and driving scenes, the visual-language datasets annotated from nuScenes may not provide a robust benchmark for visual-language understanding in driving scene. Consequently, there is an urgent need for new, large-scale datasets that encompass a wide range of traffic and driving scenarios, including numerous corner cases, to effectively test and enhance these models in autonomous driving applications.

Hardware Support for Large Language Models in Autonomous Driving.

In the use case of LLMs as the planner for autonomous driving, the perception reasoning for the LLMs and the subsequent control decision should be generated in real-time with low latency in order to meet safety requirements for autonomous driving. The number of (Floating-point operations per second)FLOPs of the LLMs has a positive correlation with the latency as well as the power consumption, which should be of consideration if LLMs are hosted in the vehicle. For LLMs deployed remotely, the bandwidth of perception information and control decision transfer will be a great challenge.

Another use case for LLMs in autonomous driving is a navigation planner . Unlike driving planners, the tolerance of response time for the LLMs is much higher, and the number of queries for navigation planners is far less in general. Consequently, the hardware performance demand is easier to meet, and even moving the host to remote servers is a reasonable proposal.

The user-vehicle interaction could also be a use case of LLMs in autonomous driving . LLMs could interpret drivers’ intentions into control commands given to the vehicle. For intentions unrelated to driving, e.g. entertainment control, the high latency of the response from LLMs could be accepted. However, if the intentions involve taking over autonomous driving, then the hardware requirements would be similar to the counterpart of using LLMs as an autonomous driving planner where LLMs are expected to respond with low latency.

LLMs in the applications of autonomous driving could potentially be compressed, which reduces the computation power requirements and the latency and lowers the HW limitation. However, the current effort in this field is still undeveloped.

Using Large Language Models for Understanding HD Maps.

HD maps play a crucial role in autonomous vehicle technology, as they provide essential information about the physical environment in which the vehicle operates. The semantic map layer from the HD map is of utmost importance as it captures the meaning and context of the physical surroundings. To effectively encode this valuable information into the LLMs-powered next-generation autonomous driving, it is important to find a way to represent and comprehend the details of the environment in the language space.

Inspired by transformer-based language models, Tesla proposes a special language that they developed for encoding lanes and their connectivities. In this language of lanes, the words and tokens represent the lane positions in 3D space. The ordering of the tokens and predicted modifiers in the tokens encode the connectivity relationships between these lanes. Producing a lane graph from the model output sentence requires less post-processing than parsing a segmentation mask or a heatmap . Pre-trained models (PTMs) have become a fundamental backbone for downstream tasks in natural language processing and computer vision. Baidu Maps has developed a system called ERNIE-GeoL, which has already been deployed in production. This system applies generic PTMs to geo-related tasks at Baidu Maps since April 2021, resulting in significant performance improvements for various downstream tasks .

Tencent has developed an HD Map AI system called THMA which is an innovative end-to-end, AI-based, active learning HD map labeling system capable of producing and labeling HD maps with a scale of hundreds of thousands of kilometers . To promote the development of this field, they proposed the MAPLM dataset containing over 2 million frames of panoramic 2D images, 3D LiDAR point cloud, and context-based HD map annotations, and a new question-answer benchmark MAPLM-QA.

User-Vehicle Interaction with Large Language Models.

Non-verbal language interpretation is also an important aspect to consider for user-autonomy teaming. Driver distraction poses a critical road safety challenge, including all activities such as smartphone use, eating, and interacting with passengers that divert attention from driving. According to the National Highway Traffic Safety Administration (NHTSA), distractions were a factor in 8.1% of the 38,824 vehicle-related fatalities in the U.S. in 2020 . This issue becomes more pressing as semi-autonomous driving systems, particularly SAE Level 3 systems, gain prominence, requiring drivers to be ready to take control when prompted .

To detect and mitigate driver distraction, driver action recognition strategies are commonly employed. These strategies involve continuous monitoring using sensors like RGB and infrared cameras, coupled with deep learning algorithms to identify and classify driver actions. Significant advancements have been made in this field .

Assessing the driver’s cognitive state is also crucial, as it greatly indicates distraction levels. Physiological monitoring, such as through EEG signals, can provide insights into a driver’s cognitive state , but the intrusiveness of such sensors and their impact on regular driving patterns must be taken into account. Besides, behavior monitoring works such as through facial analysis, gaze, human pose, and motion can also be used to analyze driver’s driving status. Furthermore, current datasets on driver action recognition often lack mental state annotations required to train models in recognizing these states from sensory data, highlighting the need for semi-supervised learning methods to address this relatively unexplored challenge .

Personlized Autonomous Driving.

The integration of LLMs into autonomous vehicles marks a paradigm shift characterized by continuous learning and personalized engagement. LLMs can continuously learn from new data and interactions, adapting to changing driving patterns, user preferences, and evolving road conditions. This adaptability results in a refined and increasingly adept performance over time. Moreover, LLMs have the capability to be precisely fine-tuned or in-context learned to match individual driver preferences, furnishing personalized assistance that significantly improves the driving experience. This personalized approach enriches the driving experience, providing assistance that not only offers information but also aligns closely with the distinct requirements and subtleties of each driver.

Recent studies have indicated the potential for LLMs to enhance real-time personalization in driving simulations, demonstrating their capacity to adapt driving behaviors in response to spoken commands. As the LLM-based personalization in autonomous driving is not well-developed, there are numerous opportunities for further research. Most recent studies focus on utilizing LLMs in the simulation environment instead of real vehicles. Integrating LLMs into actual vehicles is an exciting area of potential, moving beyond simulations to affect real-world driving experiences. Additionally, future investigations could also explore the development of LLM-driven virtual assistants that align with drivers’ individual preferences, the employment of LLMs for the enhancement of safety features like fatigue detection, the application of these models in predictive vehicle maintenance, and the personalization of routing to align with drivers’ unique inclinations. Furthermore, LLMs have the potential for personalizing in-vehicle entertainment, learning from drivers’ behaviors to improve the driving experience.

Trustworthy and Safety for Autonomous Driving.

Another crucial takeaway is enhancing transparency and trust. When the vehicle makes a complex decision, such as overtaking another vehicle on a high-speed, two-lane highway, passengers and drivers might naturally have questions or concerns. In these instances, the LLM doesn’t just execute the task but also articulates the reasoning behind each step of the decision-making process. By providing real-time, detailed explanations in understandable language, the LLM demystifies the vehicle’s actions and underlying logic. This not only satisfies the innate human curiosity about how autonomous systems work but also builds a higher level of trust between the vehicle and its occupants.

Moreover, the advantage of “zero-shotting” was particularly evident during the complex overtaking maneuver on a high-speed Indiana highway. Despite the LLM not having encountered this specific set of circumstances before—varying speeds, distances, and even driver alertness—it was able to use its generalized training to safely and efficiently generate a trajectory for the overtaking action. With some uncertainty estimation techniques , this can ensure that even in dynamic or edge case scenarios, the system can make sound judgments while keeping the user informed, therefore building confidence in autonomous technology.

To sum up, LLMs demonstrate their potential to revolutionize autonomous driving by enhancing safety, transparency, and user experience. Tasked with complex commands like overtaking, the LLM considered real-time data from multiple vehicle modules to make informed decisions, clearly articulating these to the driver. The model also leveraged its zero-shot learning capabilities to adapt to new scenarios, providing personalized, real-time feedback. Overall, the LLM proved effective in building user trust and improving decision-making in autonomous vehicles, emphasizing its utility in future automotive technologies.

Conclusion

In this survey, we explored the pattern of integrating multimodal large language models (MLLMs) into the next generation of autonomous driving systems. Our study began with an overview of the development of both MLLMs and autonomous driving, which have traditionally been considered distinct fields but are now increasingly interconnected. Then, we conducted an extensive literature review on the specific algorithms and applications of multimodal language models for autonomous driving and then focused on the current state of research and benchmarking datasets that apply MLLMs to autonomous driving. A significant highlight of our study was the synthesis of key insights and findings from the first LLVM-AD workshop such as proposing new datasets and improving current MLLMs algorithms on autonomous driving. Finally, we engaged in a forward-looking discussion on vital research themes and the promising potential for enhancing MLLMs in autonomous driving. We discussed both challenges and opportunities that lie ahead, aiming to show the pathway for further exploration. In general, this paper serves as a valuable resource for researchers in the autonomous driving area. It offers a comprehensive understanding of the significant role and vast potential that MLLMs hold in revolutionizing the landscape of autonomous transportation. We hope this paper could facilitate research in integrating MLLMs with autonomous driving in the future.

Acknowledgments

We would like to express our gratitude for the support received from the Purdue University Digital Twin Lab (https://purduedigitaltwin.github.io/), Tencent T Lab, and PediaMed AI (http://pediamedai.github.io/) for their contributions to this survey paper.

References