DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Yang Wang, Zhiyong Zhao, Kun Zhan, Peng Jia, Xianpeng Lang, Hang Zhao
Introduction
Autonomous driving, with its great promise to revolutionize transportation and urban mobility, has been one of the most active areas of research and development over the past two decades. A primary hurdle to a fully autonomous driving system is scene understanding , which involves navigating complex, unpredictable scenarios such as adverse weather, intricate road layouts, and unforeseen human behaviors.
Existing autonomous driving systems, typically comprising 3D perception, motion prediction, and planning, struggle with these scene understanding challenges. Specifically, 3D perception is limited to detecting and tracking familiar objects, omitting rare objects and their unique attributes; motion prediction and planning focus on trajectory-level actions, often neglecting the decision-level interactions between objects and the vehicle.
We introduce DriveVLM, a novel autonomous driving system that aims at the scene understanding challenges, capitalizing on the recent Vision-Language Models (VLMs) which have demonstrated exceptional prowess in visual comprehension and reasoning. Specifically, DriveVLM contains a Chain-of-Though (CoT) process with three key modules: scene description, scene analysis, and hierarchical planning. The scene description module linguistically depicts the driving environment and identifies critical objects in the scene; the scene analysis module delves into the characteristics of the critical objects and their influence on the ego vehicle; the hierarchical planning module formulates plans step-by-step, from meta-actions and decision descriptions to waypoints. These modules respectively correspond to the components of the traditional perception-prediction-planning pipeline, but they differ in that they tackle object perception, intention-level prediction and task-level planning, which were extremely challenging to cope with in the past.
While VLMs excel in visual understanding, they have limitations in spatial grounding and reasoning, and their computational intensity poses challenges for onboard inference speed. Therefore we further propose DriveVLM-Dual, a hybrid system that combines the strengths of both DriveVLM and traditional systems. DriveVLM-Dual optionally integrates DriveVLM with traditional 3D perception and planning modules, such as 3D object detectors, occupancy networks, and motion planners, enabling the system to achieve 3D grounding and high-frequency planning abilities. This dual system design, akin to the human brain’s slow and fast thinking processes, adapts efficiently to varying complexity in driving scenarios.
Meanwhile, we formally define the scene understanding and planning (SUP) task, and propose new evaluation metrics to assess the scene analysis and meta-action planning capabilities of DriveVLM and DriveVLM-Dual. Furthermore, we carry out a comprehensive data mining and annotation pipeline to construct an in-house SUP-AD dataset for the SUP task.
Extensive experiments on the nuScenes dataset and our dataset demonstrate the superiority of DriveVLM, especially in few-shot situations. Moreover, DriveVLM-Dual surpasses state-of-the-art end-to-end motion planning methods.
In summary, the contributions of this paper are fourfold:
We introduce DriveVLM, a novel autonomous driving system that leverages VLMs for effective scene understanding and planning.
We further introduce DriveVLM-Dual, a hybrid system that incorporates DriveVLM and a traditional autonomous pipeline. DriveVLM-Dual achieves improved spatial reasoning and real-time planning capabilities.
We present a comprehensive data mining and annotation pipeline to construct a scene understanding and planning dataset, together with metrics to evaluate the SUP task.
Extensive experiments on the nuScenes dataset and our SUP-AD dataset demonstrate the superior performance of DriveVLM and DriveVLM-Dual in complex driving scenarios.
Related Work
Recently, there has been a surge in research on large Vision-Language Models (VLMs), exemplified by works such as MiniGPT-4 , LLaVA , Qwen-VL , and others . These models integrate pre-trained vision encoders with large language models, enabling large language models to address many tasks involving images as input. In general, these methods align image features with the input embedding space of the language model through Q-former or linear mapping . A crucial step in the training process is supervised fine-tuning using instructional data containing images and text, enhancing the overall performance of vision language models. VLMs can be used in various scenarios, especially robotics . Specifically, given instructions, input images, and robot states, vision language models output corresponding actions that can be high-level instructions or low-level robot actions . DriveVLM focuses on utilizing VLMs to assist in autonomous driving, thereby establishing a novel framework. Concurrent to our work, also shares a similar motivation.
The integration of learning frameworks into motion planning has been an active area of research since Pomerleau pioneering contributions. One promising line of work is Reinforcement learning and imitation learning . These methods can learn an end-to-end planning policy that directly maps raw sensory inputs to control actions . They are particularly suited to high-dimensional state and action spaces, a common challenge in motion planning. However, the direct generation of control outputs from sensor data poses significant challenges in robustness and safety assurance . Several works improve interpretability by explicitly building dense cost maps derived from learning-based modules. While dense cost maps effectively integrate predictions about traffic agents’ future movements and environmental factors, their performance heavily depends on costs tailored through human experience and the trajectory sampling distribution . A recent trend involves training multiple blocks in an end-to-end fashion . These methods enhance overall performance, but rely on backpropagation from future trajectory predictions loss in a less interpretable decision-making process . Our model, DriveVLM, addresses the complexities of long-tailed driving scenarios, often challenging for other methods, by leveraging the generalization and reasoning capabilities of vision-language models. Moreover, users can easily interact with our model through the intuitive language interface provided by the vision-language model, enhancing interpretability.
Recent works argue that language captions are an important medium to connect human knowledge with the driving objective, helping to inform decisions and actions. In support of this trend, some efforts have enhanced mainstream driving scene datasets. Refer-KITTI annotates objects in the KITTI dataset with language prompts that can reference a collection of objects. Talk2Car , NuPrompt and nuScenes-QA introduce free-form captions and QA annotation to the nuScenes dataset . However, these works enrich the datasets that are perception-focused and often contain simple traffic scenes. Instead of augmenting existing datasets, BDD-X and BDD-OIA offer datasets with natural language explanations for the ego vehicle’s actions. HAD employs natural language commands to produce salient maps from drivers’ gaze data. Rank2Tell and DRAMA annotate language explanations and risk localization for traffic scenarios. While these datasets provide scenes tailored for utilizing natural language, there is a lack of enough data that capture scenarios that are crucial for identifying issues that could lead to safety concerns in self-driving systems. Our SUP-AD dataset stands in contrast by purposefully gathering a diverse array of challenging, long-tail scenarios that are essential for addressing complex scene understanding and planning.
DriveVLM
The overall pipeline of DriveVLM is illustrated in Figure 1. A sequence of images is processed by a large Vision Language Model (VLM) to perform a special chain-of-thought (CoT) reasoning to derive the driving planning results. The large VLM involves a vision transformer encoder and a Large Language Model (LLM). The vision encoder produces image tokens; then an attention-based extractor aligns these tokens with the LLM; finally, the LLM performs CoT reasoning. The CoT process can be divided into three modules: scene description (Section 3.2), scene analysis (Section 3.3), and hierarchical planning (Section 3.4).
DriveVLM-Dual is a hybrid system that combines DriveVLM and the traditional autonomous driving pipeline, taking the best of both worlds. It incorporates 3D perception results as language prompts for enhanced 3D scene understanding capability, and further refines the trajectory waypoints with a real-time motion planner. We detail its design and advantages in Section 3.5.
2 Scene Description
The scene description module is composed of environment description and critical object identification.
Environment description. Driving environments, such as weather and road conditions, have a non-negligible impact on driving difficulty. Therefore, the model is first prompted to output a linguistic description of the driving environment, including several conditions: each representing a crucial aspect of the driving environment.
details the weather conditions, ranging from sunny to snowy. Conditions such as rain or fog demand more cautious driving approaches due to reduced visibility and road grip.
encapsulates the time of day, differentiating between daytime and nighttime driving scenarios. For example, nighttime driving, characterized by reduced visibility, necessitates cautious driving strategies.
classifies the type of roads, including urban, suburban, rural, or highway, where each road type presents unique challenges.
gives a description of the lane conditions, identifying the vehicle’s current lane and potential alternatives for maneuvering. This information is vital for lane selection and safe lane changes.
Critical object identification. In addition to environmental conditions, various objects in driving scenarios significantly influence driving behaviors. Unlike traditional autonomous driving perception modules, which detect all objects within a specific range, we solely focus on identifying critical objects that are most likely to influence the current scenario, inspired by human cognitive processes during driving. Each critical object, denoted as , contains two attributes: the object category and its approximate bounding box coordinates on the image. The category and coordinates are mapped to their corresponding language in the language modality, enabling seamless integration into the following modules. Moreover, taking advantage of the pre-trained vision encoder, DriveVLM can identify long-tail critical objects that may elude typical 3D object detectors, such as road debris or unusual animals.
3 Scene Analysis
In the traditional autonomous driving pipeline, the prediction module typically concentrates on forecasting the future trajectories of objects. The emergence of advanced vision-language models has provided us with the ability to perform a more comprehensive analysis of the current scene.
Critical Object Analysis. After identifying the critical objects, we analyze their characteristics and potential influence on the ego vehicle. Characteristics contain three aspects of a critical object: static attributes , motion states , and particular behaviors . Static attributes describe inherent properties of objects, such as a roadside billboard’s visual cues or a truck’s oversized cargo, which are critical in preempting and navigating potential hazards. Motion states describe an object’s dynamics over a period, including position, direction, and action—characteristics that are vital in predicting the object’s future trajectory and potential interactions with the ego vehicle. Particular behaviors refer to special actions or gestures of an object that could directly influence the ego vehicle’s next driving decisions. For instance, a traffic officer’s hand signals are critical in this context, as they can override standard traffic rules and necessitate a corresponding response from the autonomous system. We do not require the model to analyze the three characteristics (, , ) for all the objects. In practice, only one or two apply to a critical object.
Upon analyzing these characteristics, DriveVLM then predicts the potential influence of each critical object on the ego vehicle. For example, a drunken pedestrian on the roadside could potentially step onto the road and block our way. Compared to trajectory-level prediction in the traditional pipeline, the analysis of the potential influence of critical objects is crucial for the system’s adaptability to real-world and long-tail driving scenarios.
Scene-level Summary . The scene-level analysis summarizes all the critical objects together with the environmental description. This summary gives a comprehensive understanding of the scene, linking the following planning module.
4 Hierarchical Planning
We integrate the scene description and scene analysis to form a summary of the driving scenario. The summary is further combined with the route, ego pose and velocity to form a prompt for planning. Finally, DriveVLM progressively generates driving plans, in three stages: meta-actions, decision description, and trajectory waypoints.
Meta-actions . A meta-action, denoted as , represents a short-term decision of the driving strategy. These actions fall into 17 categories, including but not limited to acceleration, deceleration, turning left, changing lanes, minor positional adjustments, and waiting. To plan the ego vehicle’s future maneuver over a certain period, we generate a sequence of meta-actions. Each meta-action in this sequence is pivotal, contributing cumulatively to the strategic navigation of the vehicle in the scene.
Decision description . Decision description articulates the more fine-grained driving strategy the ego vehicle should adopt. It contains three elements: Action , Subject , and Duration . Action pertains to meta actions such as ‘turn’, ‘wait’, or ‘accelerate’. Subject refers to the interacting object, such as a pedestrian, a traffic signal, or a specific lane. Duration indicates the temporal aspect of the action, specifying how long it should be carried out or when it should start. An example of a decision description is: “Wait () for the pedestrian () to cross, then () proceed to accelerate () and merge into the right lane ().". This structured decision description allows for clear, concise, and actionable instructions for the autonomous system.
Trajectory waypoints . Upon establishing the decision description , our next phase involves the generation of corresponding trajectory waypoints. These waypoints, denoted by , , depict the vehicle’s path over a certain future period with predetermined intervals . We map these numerical waypoints into language tokens for auto-regressive generation. In this way, DriveVLM achieves seamless integration of its linguistic processing module with spatial navigation. The trajectory waypoints are the spatial manifestation of the meta-actions and decision descriptions, which can be directly fed into subsequent control modules.
5 DriveVLM-Dual
Although VLMs are adept at recognizing long-tail objects and understanding complex scenarios, they often struggle with precisely comprehending spatial positions and detailed motion states of objects. This shortfall, noted in previous research and our pilot studies, poses a significant challenge. What is worse, the humoungous model size of VLMs leads to high latency, impeding their ability to respond in real-time for autonomous driving. To address these challenges, we propose DriveVLM-Dual, a collaboration between DriveVLM and the traditional autonomous driving system. This novel approach involves two key strategies: incorporating 3D perception for critical object analysis, and high-frequency trajectory refinement.
Integrating 3D Perception. We represent objects detected by a 3D detector as , where denotes the -th bounding box and denotes its category. These 3D bounding boxes are then back-projected onto 2D images to derive corresponding 2D bounding boxes . We conduct IoU matching between these 2D bounding boxes and . are the bounding boxes of previously identified critical objects . We classify critical objects that meet a certain approximate IoU threshold and belong to the same category as matched critical objects , defined as
Those critical objects without a corresponding match in the 3D data are noted as .
In the scene analysis module, for , the center coordinates, orientations, and historical trajectories of the corresponding 3D objects are used as language prompts for the model, assisting in object analysis. Conversely, for , analysis relies solely on the language tokens derived from the image. This novel use of 3D perception results as prompts enables DriveVLM-Dual to understand the locations and motions of critical objects more accurately, enhancing the overall performance.
High-Frequency Trajectory Refinement. Compared to traditional planners, DriveVLM, due to its immense parameter size inherent to Vision-Language Models (VLMs), exhibits significantly slower speeds while generating a trajectory. To achieve real-time, high-frequency inference capabilities, we integrate it with a conventional planner to form a slow-fast dual system, combining the advanced capabilities of DriveVLM with the efficiency of traditional planning methods. After obtaining a trajectory from DriveVLM at low frequency, denoted as , we take it as a reference trajectory for a classical planner for high-frequency trajectory refinement. In the case of an optimization-based planner, serves as the initial solution for the optimization solver. For a neural network-based planner, is used as an input query, combined with additional input features , and then decoded into a new planning trajectory denoted as . The formulation of this process can be described as:
This refinement step ensures that the trajectory produced by DriveVLM-Dual (1) achieves higher trajectory quality, and (2) meets real-time requirements. In practice, the two branches operate asynchronously in a slow-fast manner, where the planner module in the traditional autonomous driving branch can selectively receive trajectory from the VLM branch as additional input.
Task and Dataset
To fully exploit the potential of DriveVLM and DriveVLM-Dual in handling complex and long-tail driving scenarios, we formally define a task called Scene Understanding for Planning (Section 4.1), together with a set of evaluation metrics (Section 4.2). Furthermore, we propose a data mining and annotation protocol to curate a scene understanding and planning dataset (Section 4.3).
The Scene Understanding for Planning task is defined as follows. The input comprises multi-view videos from surrounding cameras and optionally 3D perception results from a perception module. The output includes the following components:
Scene Description : Composed of weather condition , time , road condition , and lane conditions .
Scene Analysis : Including object-level analysis and scene-level summary .
Meta Actions : A sequence of actions representing task-level maneuvers.
Decision Description : A detailed account of the driving decisions.
Trajectory Waypoints : The waypoints outlining the planned trajectory of the ego vehicle.
2 Evaluation Metrics
To comprehensively evaluate a model’s performance, we care about its interpretation of the driving scene and the decisions made. Therefore, our evaluation has two aspects: scene description/analysis evaluation and meta-action evaluation.
Scene description/analysis evaluation. Given the subjective nature of human evaluation in scene description, we adopt a structured approach using a pre-trained LLM. This method entails comparing the generated scene description with a human-annotated ground truth description. The ground truth description encompasses structured data such as environmental conditions, navigation, lane information, and critical events with specific objects, verbs, and their influences. The LLM assesses and scores the generated descriptions based on their consistency with the ground truth.
Meta-action evaluation. Meta-actions are a predefined set of decision-making options. A driving decision is formulated as a sequence of meta-actions. Our evaluation method employs a dynamic programming algorithm to compare the model-generated sequences with a manually annotated ground truth sequence. The evaluation should also weigh the relative importance of various meta-actions, designating some as ‘conservative actions’ with a lower impact on the sequence’s overall context. To increase robustness, we first use the LLM to generate semantically equivalent alternatives to the ground truth sequence to enhance robustness. The sequence with the highest similarity to these alternatives calculates the final driving decision score. More details of the proposed metric are available in the Appendix B.
3 Dataset Construction
We propose a comprehensive data mining and annotation pipeline, shown in Figure 3, to construct a Scene Understanding for Planning (SUP-AD) Dataset for the proposed task. Specifically, we perform long-tail object mining and challenging scenario mining from a large database to collect samples, then we select a keyframe from each sample and further perform scene annotation. Dataset statistics are available in the Appendix A.
Long-tail object mining. According to real-world road object distribution, we first define a list of long-tail object categories, such as weird-shaped vehicles, road debris, and animals crossing the road. Next, we mine these long-tail scenarios using a CLIP-based search engine, capable of mining driving data using language queries from a large collection of logs. Following that, we perform a manual inspection to filter out scenes inconsistent with the specified categories.
Challenging scenario mining. In addition to long-tail objects, we are also interested in challenging driving scenarios, where the driving strategy of the ego vehicle needs to be adapted according to the changing driving conditions. These scenarios are mined according to the variance of the recorded driving maneuvers.
Keyframe selection. Each scene is a video clip, it is essential to identify the ‘keyframe’ to annotate. In most challenging scenarios, a keyframe is the moment before a significant change in speed or direction is required. We select this keyframe 0.5s to 1s earlier than the actual maneuver, based on comprehensive testing, to guarantee an optimal reaction time for decision-making. For scenes that do not involve changes in driving behavior, we select a frame that is relevant to the current driving scenario as the keyframe.
Scene annotation. We employ a group of annotators to perform the scene annotation, including scene description, scene analysis, and planning, except for waypoints, which can be auto-labeled from the vehicle’s IMU recordings. To facilitate scene annotation, we make a video annotation tool with the following features: (1) the annotators can slide the progress bar back and forth to replay any part of a video; (2) while annotating a keyframe, the annotator can draw bounding boxes on the image together with language descriptions; (3) annotators can select from a list of action and decision candidates while annotating driving plans. Each annotation is meticulously verified by 3 annotators for accuracy and consistency, ensuring a reliable dataset for model training. Figure 2 illustrates a sample scenario with detailed annotations.
Experiments
SUP-AD dataset. The SUP-AD dataset is a dataset built by our proposed data mining and annotation pipeline. It is divided into train, validation, and test splits with a ratio of . We train models on the training split and use our proposed scene description and meta-action metrics to evaluate model performance on the validation/test split.
nuScenes dataset. The nuScenes dataset is a large-scale driving dataset of urban scenarios with 1000 scenes, where each scene lasts about 20 seconds. Keyframes are evenly annotated at a frequency of 2Hz over the entire dataset. Following previous works , we adopt Displacement Error (DE) and Collision Rate (CR) as metrics to evaluate models’ performance on the validation split.
1.2 Base Model
We use Qwen-VL as our default large vision-language model, which exhibits remarkable performance in tasks like question answering, visual localization, and text recognition. It contains a total of 9.6 billion parameters, including a visual encoder (1.9 billion), a vision-language adapter (0.08 billion), and a large language model (Qwen, 7.7 billion). Images are resized to a resolution of before being encoded by the vision encoder. During training, we randomly select a sequence of images at the current time s, s, s, and s as input. The selected images ensure the inclusion of the current time frame and follow an ascending chronological order.
2 Main Results
SUP-AD. We present the performance of our proposed DriveVLM with several large vision-language models and compare them with GPT-4V, as shown in Table 1. DriveVLM, utilizing Qwen-VL as its backbone, achieves the best performance due to its strong capabilities in question answering and flexible interaction compared to the other open-source VLMs. Although GPT-4V exhibits robust capabilities in vision and language processing, its inability to undergo fine-tuning, restricting it solely to in-context learning, often results in the generation of extraneous information during scene description tasks. Under our evaluation metric, the additional information is frequently classified as hallucination, consequently leading to lower scores.
nuScenes. As shown in Table 2, DriveVLM-Dual achieves state-of-the-art performance on the nuScenes planning task when cooperating with VAD. It demonstrates that our method, although tailored for understanding complex scenes, also excels in ordinary scenarios. Note that DriveVLM-Dual significantly improves over UniAD: it achieves a reduction of 0.64 meters in terms of average planning displacement error, and a 51% reduction of collision rate.
3 Ablation Study
Model Design. To better understand the significance of our designed modules in DriveVLM, we conduct ablations on different combinations of modules, as shown in Table 3. The inclusion of critical object analysis enables our model to identify and prioritize important elements in the driving environment, enhancing the decision-making accuracy for safer navigation. Integrating 3D perception data, our model gains a refined understanding of the surroundings, which is crucial for capturing the motion dynamics and improving trajectory predictions.
Inference speed. The inference speed of DriveVLM and DriveVLM-Dual are measured on the NVIDIA Orin platform, shown in Table 4. Due to the huge number of parameters of LLM, the inference speed of DriveVLM is an order slower than the conventional autonomous driving method similar to VAD, preventing it from running onboard. However, after cooperating with the traditional autonomous driving pipeline in a slow-fast cooperation pattern, the overall latency depends on the speed of the fast branch, making DriveVLM-Dual an ideal solution for real-world deployment.
4 Qualitative Results
Qualitative results of DriveVLM are shown in Figure 4. In Figure 4(a), DriveVLM accurately predicts the current scene conditions and incorporates well-considered planning decisions regarding the cyclist approaching us. In Figure 4(b), DriveVLM effectively comprehends the gesture of the traffic police ahead, signaling the ego vehicle to proceed, and also considers the person riding a tricycle on the right side, thereby making sensible driving decisions. These qualitative results demonstrate our model’s exceptional ability to understand complex scenarios and make suitable driving plans. More visualization of our model’s output is shown in the Appendix C.
Conclusion
In summary, we introduce DriveVLM and DriveVLM-Dual. DriveVLM leverages VLMs, significantly progressing in interpreting complex driving environments. The DriveVLM-Dual further enhances these capabilities by synergizing existing 3D perception and planning approaches, effectively addressing the spatial reasoning and computational challenges inherent in VLMs. Moreover, we define a scene understanding for planning task for autonomous driving, together with evaluation metrics and dataset construction protocol. Through rigorous evaluation, DriveVLM and DriveVLM-Dual have demonstrated their ability to surpass state-of-the-art methods in autonomous driving, especially in handling intricate and dynamic scenarios. We believe this research offers a roadmap for the development of safe and interpretable autonomous vehicles in the future.
References
Appendix A SUP-AD Dataset
We use the meta-action sequence to formally represent the driving strategy. Meta actions are classified into 17 categories. We show the distribution of each meta-action being the first/second/third place in the meta-action sequence, as shown in Figure 5. It indicates that the meta-actions are quite diverse in the SUP-AD dataset. We also show the distribution of the length of meta-actions per scene in Figure 6. Most scenes contain two or three meta-actions, and a few scenes with complex driving strategies contain four or more meta-actions.
The meta-action sequence for each driving scene is manually annotated based on the actual driving strategy in the future frames. These meta-actions are designed to encompass a complete driving strategy and are structured to be consistent with the future trajectory of the ego vehicle. They can be divided into three primary classes:
Speed-control actions. Discerned from acceleration and braking signals within the ego state data, these actions include These actions can be discerned from acceleration and braking signals within the ego state data. They include speed up, slow down, slow down rapidly, go straight slowly, go straight at a constant speed, stop, wait, and reverse.
Turning actions. Deduced from steering wheel signals, these actions consist of turn left, turn right, and turn around.
Lane-control actions. Encompassing lane selection decisions, these actions are derived from a combination of steering wheel signals and either map or perception data. They involve change lane to the left, change lane to the right, shift slightly to the left, and shift slightly to the right.
A.2 Scenario Categories
As shown in Figure 7, the SUP-AD Dataset encompasses diverse driving scenarios, spanning over 40 categories. Detailed explanations for certain scenario categories are provided below:
AEB Data: Automatic Emergency Braking (AEB) data.
Road Construction: A temporary work zone with caution signs, barriers, and construction equipment ahead.
Close-range Cut-ins: A sudden intrusion into the lane of the ego vehicle by another vehicle.
Roundabout: A type of traffic intersection where vehicles travel in a continuous loop.
Animals Crossing Road: Animals crossing the road in front of the ego vehicle.
Braking: Brake is pressed by human driver of the ego vehicle.
Traffic Police Officers: Traffic police officers managing and guiding traffic.
Blocking Traffic Lights: A massive vehicle obscuring the visibility of the traffic signal.
Cutting into Other Vehicle: Intruding into the lane of another vehicle ahead.
Ramp: A curved roadway that connects the main road to the branch road in highway.
Debris on the Road: Road with different kinds of debris.
Narrow Roads: Narrow roads that require cautious navigation.
Pedestrians Popping Out: Pedestrians popping out in front of the ego vehicle, requiring slowing down or braking.
People on Bus Posters: Buses with posters, which may interfere the perception system.
Merging into High Speed: Driving from a low-speed road into a high-speed road, requiring speeding up.
Barrier Gate: Barrier gate that can be raised obstructing the road.
Fallen Trees: Fallen trees on the road, requiring cautious navigation to avoid potential hazards.
Complex Environments: Complex driving environments that requiring cautious navigation.
Mixed Traffic: A congested scenario where cars, pedestrians, and bicycles appear on the same or adjacent roadway.
Crossing Rivers: Crossing rivers by driving on the bridge.
Screen: Roads with screens on one side, which may interfere the perception system.
Herds of Cattle and Sheep: A rural road with herds of cattle and sheep, requiring careful driving to avoid causing distress to these animals.
Vulnerable Road Users: Road users which are more susceptible to injuries while using roads, such as pedestrians, cyclists, and motorcyclists.
Road with Gallet: A dusty road with gallet scattered across the surface.
The remaining scenario categories are: Motorcycles and Trikes, Intersection, People carrying Umbrella, Vehicles Carrying Cars, Vehicles Carrying Branches, Vehicles with Pipes, Strollers, Children, Tunnel, Down Ramp, Sidewalk Stalls, Rainy Day, Crossing Train Tracks, Unprotected U-turns, Snowfall, Large Vehicles Invading, Falling Leaves, Fireworks, Water Sprinklers, Potholes, Overturned Motorcycles, Self-ignition and Fire, Kites, Agricultural Machinery.
A.3 Annotation Examples
We provide more examples of annotation contents in Figure 8, 9, 10, 11, 12, and 13. The scenario categories of these examples are overturned bicycles and motorcycles, herds of cattle and sheep, collapsed trees, crossing rivers, barrier gate, and snowfall respectively.
Appendix B Evaluation Method
The ability of an autonomous driving system to accurately interpret driving scenes and make logical, suitable decisions is of paramount importance. As presented in this paper, the evaluation of VLMs in autonomous driving concentrates on two primary components: the evaluation of scene description/analysis and the evaluation of meta-actions.
In terms of scene description/analysis evaluation, the process of interpreting and articulating driving scenes is subject to inherent subjectivity, as there are numerous valid ways to express similar descriptions textually, which makes it difficult to effectively evaluate the scene description using a fixed metric. To overcome this challenge, we utilize GPT-4 to evaluate the similarity between the scene descriptions generated by the model and the manually annotated ground truth. Initially, we prompt GPT-4 to extract individual pieces of information from each scene description. Subsequently, we score and aggregate the results based on the matching status of each extracted piece of information.
The ground truth labels for scene descriptions encompass both environment descriptions and event summaries. Environmental condition description includes weather conditions, time conditions, road environment, and lane conditions. Event summaries are the characteristics and influence of critical objects. We employ GPT-4 to extract unique key information from both environment descriptions and event summaries. The extracted information is then compared and quantified. Each matched pair is assigned a score, which is estimated based on the extent of the matching, whether complete, partial, or absent. Instances of hallucinated information incur a penalty, detracting from the overall score. The aggregate of these scores constitutes the scene description score.
The prompt for GPT-4 in evaluating scene descriptions is carefully designed, as shown in Table 5. Initially, a role prompt is employed to establish as an intelligent and logical evaluator, possessing a comprehensive understanding of appropriate driving styles. This is followed by specifying the input format, which informs GPT-4 that its task involves comparing an output description with a ground truth description. This comparison is based on the extraction and analysis of key information from both descriptions. Lastly, the prompt outlines the criteria for scoring, as well as the format for the evaluation output, ensuring a structured and systematic approach to the evaluation process.
B.2 Meta-action Evaluation
The evaluation process for the meta-action sequence must consider both the quantity and the sequential arrangement of the matched meta-actions. We employ dynamic programming to compare the model’s output and the annotated ground truth. Our dynamic programming approach is similar to the method utilized in identifying the longest common subsequence, albeit with two supplementary considerations.
The first consideration acknowledges the unequal weighting of different meta-actions. For instance, certain meta actions such as “Slow Down", “Wait", and “Go Straight Slowly" exhibit a greater emphasis on attitude rather than action. The presence or absence of these actions from a meta-action sequence does not alter the basic semantic essence of driving decisions but rather modifies the driving strategy to be either more assertive or more cautious. For example, a meta action sequence of “Slow Down -> Stop -> Wait" conveys a similar driving decision as a sequence with only the meta action “Stop". Consequently, these sequences should not incur a penalty comparable to other meta actions such as “Turn Left" or “Change Lane to the Right". Therefore, these are designated as “conservative actions”, and a reduced penalty is applied when they do not match during sequence evaluation.
The second consideration addresses the potential semantic equality among different meta-action sequences. For example, the sequences “Change Lane to the Left -> Speed Up -> Go Straight At a Constant Speed -> Change Lane to the Right" and “Change Lane to the Left -> Speed Up Rapidly -> Go Straight At a Constant Speed -> Change Lane to the Right" might both represent valid approaches to overtaking a slow-speed vehicle ahead. Recognizing that different meta-action sequences might convey similar meanings, we initially use GPT-4 to generate variant sequences that have comparable semantic meanings, in addition to the unique ground truth meta-action sequence, as shown in Table 6. In the subsequent sequence-matching phase of the evaluation, all these variations, together with the manually annotated ground truth, are taken into consideration. The highest-scoring matching is then adopted as the definitive score for the final decision evaluation.
The state of dynamic programming is saved in a 2D matrix, wherein each row corresponds to a meta action in the ground truth action sequence, and each column corresponds to a meta action in the model output action sequence, noted as . The dynamic programming initiates recursive calculations beginning from the first meta action of both sequences. Each element of the 2D matrix encompasses the optimal total score at the current matching position, as well as the preceding matching condition that yielded the optimal matching. In our dynamic programming algorithm, three transition equations govern distinct cases: for missing matching, for redundant matching, and for successful matching. Successful matching occurs when the meta action is identical at the position in the reference sequence and the position in the model-generated sequence. In the case of missing matching, the meta action at the position in the reference sequence is unmatched, prompting a comparison with the position in the reference sequence and the position in the model-generated sequence. Conversely, redundant matching implies that the meta action at the position in the model-generated sequence is unmatched, leading to further examination of the position in the reference and the position in the model-generated sequence. The transformation equations for these cases are as follows:
where represents the reward score after a successful matching. If an action considered missing or redundant is classified as a conservative action, the penalties and are quantified as half of , i.e., 0.5. Conversely, if an action is not conservative, both penalties are assigned the same magnitude as , i.e., 1.0. This approach is based on the premise that omitting a crucial meta action or inaccurately introducing a non-existent one equally hampers the effectiveness of the action sequence. The final score should be divided by the length of the selected reference meta-action sequence, formulated as follow:
Appendix C Qualitative Results
To further demonstrate the effectiveness and robustness of our proposed DriveVLM, we provide additional visualization results in Figure 14, 15, 16, 17, and 18. In Figure 14, DriveVLM recognizes the slowly moving vehicle ahead and provides a driving decision to change lanes for overtaking. In Figures 15 and 16, DriveVLM accurately identifies the type of unconventional vehicles and a fallen tree, demonstrating its capability in recognizing long-tail objects. In Figure 17, the traffic police signaling to proceed with hand gestures has been accurately captured by DriveVLM. In Figure 18, DriveVLM successfully recognizes the road environment of a roundabout and generates a planned trajectory with a curved path.