NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets using Markup Annotations

Yuichi Inoue, Yuki Yada, Kotaro Tanahashi, Yu Yamaguchi

Introduction

Generating accurate text explanations for visual scenes has become crucial in the progress of Vision Language Models (VLMs) integrated with Large Language Models (LLMs). It is believed that many challenges can be more effectively tackled by describing visuals in natural language using VLMs. Specifically, autonomous driving requires accurate recognition and complex situation evaluations. This context underscores the potential benefits of VLMs, which can harness the superior logical reasoning capabilities of LLM. As a result, the pursuit of VLMs tailored for autonomous driving is currently a topic of intense research interest.

The trend of using autonomous driving datasets enriched with natural language annotations has been growing steadily. By using the data described in natural language, we can align with LLMs that possess general knowledge and high logical reasoning capabilities. By incorporating LLMs, there is potential to create a more intelligent autonomous driving system. To achieve this, it might be necessary to formulate a VLM and curate a relevant dataset for its training. However, datasets annotated in a QA format, which ensures precise language generation and scene recognition from driving scenes, are still a challenge to obtain.

Interpreting visual content accurately is not only essential for autonomous driving research, but is also crucial across diverse tasks. Among them, VQA, which provides accurate descriptions of the visual scenes, is particularly important. Various VQA datasets have been proposed, continuously demonstrating their significant value in training and evaluating state-of-the-art VLMs. Traditional QA tasks have predominantly focused on predicting a singular word. However, with the recent proliferation of sophisticated high-performance LLMs, predicting just one word may not fully harness their potential and may even suppress their inherent linguistic generative abilities. To holistically validate the model’s comprehension of the visual content, LLaVA-RLHF employed RLHF (Reinforcement Learning from Human Feedback) to counteract the vision language model’s hallucination. However, this proved to be both time-consuming and costly.

To address this, we introduced ”Markup-QA”, wherein the QA segment within a naturally composed text is enclosed by our unique markups. Post-processing extracts this markup-wrapped segment, enabling the evaluation of the accuracy of QAs embedded in the text. By removing the markup, the text retains its completeness, allowing us to assess the model’s text generation capabilities using standard evaluation metrics. Another advantage of Markup-QA is its flexibility, allowing for the extraction of QAs from any text segment and embedding multiple QAs within a single sentence, marking an innovative departure in QA tasks.

Using the rich annotations of nuScenes concerning spatial object information, we systematically generated natural language annotations embedded with Markup-QA in a rule-based manner. This dataset, named NuScenes-MQA, comprises 1,459,9331,459,933 annotations, covering aspects such as object presence, counts, proximity, and relative positions. Using NuScenes-MQA ensures simultaneous evaluations of accurate QA capabilities and natural language proficiency.

Our contributions can be summarized as follows:

We introduced Markup-QA, a novel dataset annotation technique in which QAs are enclosed within markups. By using data annotated with Markup-QA, QA tasks can be embedded within natural sentences, allowing a concurrent evaluation of textual quality and QA accuracy.

We proposed and publicly released the NuScenes-MQA dataset annotated in the Markup-QA style, along with its evaluation methodology.

Using VLMs capable of handling multiple images, we established a baseline for the NuScenes-MQA dataset.

Related Work

Datasets for autonomous driving are inherently multimodal, curated from a variety of sensors . Recently, several existing autonomous driving datasets collected from these sensors have been augmented with textual annotations. For example, the BDD-X supplements driving conditions with textual descriptions that explain the underlying reasons. In addition to describing driving scenarios, DriveGPT4 uses off-the-shelf object detection models in conjunction with GPT, improving BDD-X with recognition tasks and textual captions. The Honda DRAMA dataset introduces the challenge of localizing risk objects and explicating their risks. In pursuing recognition-specific datasets, NuScenes-QA utilizes pre-annotated object information to propose a 3D VQA task. Moreover, DriveLM constructs a dataset for nuScenes that encapsulates perception, prediction, and planning, all described in text. The trend underscores the growing attention towards leveraging rich sensor data in autonomous driving to recognize spatial object information and articulate it through natural language.

2 VQA

The task of VQA involves processing an image and a natural language question to produce a concise natural language response. A wide variety of datasets, such as VQA, VQA v2.0, GOA, and Visual Genome, have been introduced. In particular, in the field of autonomous driving, NuScenes-QA, which incorporates position information from surrounding objects, is noteworthy. Early research predominantly combined CNN-based image feature extractor with RNNs . However, with the emergence of Transformer architecture, transformer-based language models have become the dominant choice for performance. This trend is evident with the introduction of models such as the encoder-decoder-based PALI, decoder-only LLaVA and Mini-GPT4, BLIP-2 with the resampling qformer module, and Flamingo with gated cross-attentions.

3 Special Words in Vision Language Prompts

Special words improve the efficiency of vision language tasks by better linking visual information to text. The token, for instance, ties image data directly to specific sentences . QwenVL further innovates using tokens such as and to address the visual grounding task, specifying which textual phrases correspond to the regions of the image. Kosmos-2 expands on this by associating multiple bounding boxes with phrases using tokens like , , and location tokens. These tokens enrich prompts by conveying details not captured by natural language alone.

Proposed Dataset

We introduce a novel QA dataset, called NuScenes-MQA, based on nuScenes. Moving away from the conventional short-answer paradigm, our method emphasizes full-sentence responses, enriching both the content and structure of the answers. The dataset targets key aspects of autonomous driving, such as the presence of objects and relative positioning. Using unique markups, we can highlight and evaluate specific information within the answers. With this methodology, we generated 1,459,9331,459,933 QA pairs derived from 34,14934,149 driving scenarios. Example annotations are shown in Fig. 1.

In order to create a QA dataset, we employed the annotations provided as ground truth in nuScenes. Contrary to traditional QA datasets that typically structure their response sections with one word, we decided to make our answers as full sentences. To enrich the diversity of our QA templates, we used GPT-4 and crafted 50 expressions per template, ensuring semantic consistency. Human reviewers then curated and adjusted a subset of 20 to 30 from these generated expressions.

Our dataset is based on four core concepts.

Specific Object Presence: Questions that ask about the existence and number of specific objects.

Objects in Specific Direction: Questions asking for the number and category of objects in a specific direction.

Relative Distance to Ego Vehicle: Questions about the relative distance to vehicles. For simplicity, we identify the object closest to the vehicle and its corresponding distance.

Relative location to Ego vehicle: Questions about the location of objects. Similarly, we simplified the task to identify the closest object and its coordinates.

Using the extensive information available in the nuScenes annotations, such as the class and location of recognized objects, and the cameras that capture them, we were able to automate the QA creation process.

2 Markup Implementation

Traditional QA evaluation often revolves around predicting a single word, a method that tends to compromise sentence generation capabilities. To address this, we incorporated special markups in our dataset. These markups were differentiated for each QA type as follows:

Represents an object, restricted to a single word.

Represents a count, restricted to a single word.

Represents a binary response, a single word.

By enveloping the target objects with these markups, we can easily extract the relevant words from the answers. For example:

In the back, 3 trucksare detected.

The closest object to the ego-car is a car located at coordinates (3.43, 1.41).

This method allows for the simultaneous evaluation of multiple detections. For example, in the phrase ”33 trucks”, it is necessary to recognize both the class ”trucks” and its count ”33”. In this particular regard, conventional methods are insufficient. Using our markup methodology, we can simultaneously evaluate both elements. Similarly, this framework facilitates the recognition of two or more classes at the same time. In the example provided above, it is possible to accurately respond to two distinct questions regarding the object category and the location of the object. Hence, the usage of markups enables the design of tasks that can simultaneously answer multiple queries.

3 Dataset Statistics

In this section, we discuss the statistics of the NuScenes-MQA dataset. An overview of this dataset is provided through word clouds representing the most frequent terms found within the questions and answers. These visualizations are depicted in Fig. 2 (a) and (b), showing the results for the questions and answers, respectively. The word clouds display a particularly diverse range of phrases, especially in the answers.

Specific Object Presence and Objects in Specific Direction tasks facilitate recognition-based QA from images, essentially assessing the presence and count of objects. In Fig. 2 (c), simple QA tasks seeking binary ”Yes” or ”No” answers show a slight dominance of ”Yes” responses, but the bias does not significantly influence model outcomes. To maintain task feasibility and prevent undue complexity, we excluded questions that required counting more than 20 objects. The distribution of the number of objects to be counted is depicted in Fig. 2 (d).

Our dataset harnesses the markup, enabling the embedding of multiple QAs within a single statement. Fig. 2 (e) illustrates the number of targets encapsulated in a single statement, highlighting tasks that predict counting and category. Although tasks identifying a single target are common, tasks requiring multiple target responses are also prevalent. This variety allows for an extended evaluation beyond typical simple QA.

3.2 Distance and position distribution

The nuScenes annotations provide valuable information on the location of objects. Using positional relationships between objects and the ego vehicle, we created QA tasks designated as Relative Distance to Ego vehicle and Relative Location to Ego vehicle. Drawing inspiration from the range of recognition tasks, such as the occupancy prediction , we delimited our focus to objects situated within a 4040-meter radius. Fig. 2 (f) and (g), which plot the distance and positional relationship of objects, reveal that most are within a 2020-meter range in all directions. To the best of our knowledge, our tasks are the first to require textual answers regarding the spatial information of objects.

Methodology

For our Markup-QA tasks, we introduce a model that combines a vision transformer (ViT) and a decoder-only language model via a simple linear module. The efficiency of this architecture is supported by previous works such as GIT, LLaVA, and MiniGPTv2. The architecture of our model is illustrated in Fig. 3.

The visual input undergoes feature extraction using a ViT pre-trained in CLIP, subsequently extracting patch features from the final layer. Considering the nuScenes dataset accommodates images from six distinct camera orientations, we use six ViTs, each dedicated to extracting features from its corresponding camera’s image. The ViT shares the parameters. These extracted features are subsequently added with six trainable positional embeddings, in a manner similar to . The ViT patch features are then projected, using a single layer adapter module, to match the size of the text embeddings, thus formulating the visual embeddings. These visual embeddings, similar to text embeddings, are fed into the language model for both training and inference. During training, a causal mask is applied to text embeddings to negate the influence of future information during self-attention computations. However, this constraint is not imposed on visual embeddings, as shown in Fig. 4.

Words used for markup adopt a format rarely seen in conventional text. Based on previous studies , these markups were typically incorporated as additional tokens. In the results section, we provide a comparison between tokenization using conventional tokenizers and tokenization that treats these markups as additional tokens.

Experiments

To evaluate the quality of the generated text, we used n-gram-based standard metrics, notably BLEU-1 , BLEU-4 , METEOR , and ROGUE-1 F-score . QA tasks were evaluated in the accuracy metric, while distance measurement tasks were evaluated using the mean absolute error (MAE).

Due to the inherent nature of sentence generation as an evaluation method, inference requires considerable time. To efficiently assess the model’s performance, we chose a smaller subset for the model evaluation. From our test set, we carefully extracted 2,0002,000 samples, ensuring a balanced representation of the tasks. Then, our evaluations were executed on this curated subset.

2 Training Details

For the ViT, we utilized OpenAI’s ViT large patch14, which is pre-trained using CLIP. We used an image resolution of 224224x224224 during training and evaluation. The language model was OPT , with parameter sizes of 125M, 1.3B, and 6.7B. Both the ViT and the language model were initialized using pre-trained parameters, while the adapter module started with random initialization. All parameters spanning the ViT, the language model, and the adapter module were trained.

Our training data combined random QAs from a given scene, ensuring that all types of question were present in a sample. The maximum sequence length for the text segment, excluded from visual embeddings, was set at 256256. AdamW optimizer was used, with a learning schedule outlined by a one-cycle scheduler starting at 1e−61e-6, peaking at 1e−41e-4, and dropping back to 1e−61e-6. The other parameters retained their default values. Training was carried out for 1010 epochs using the standard cross-entropy loss. The comprehensive training regimen was orchestrated on 88 Nvidia A100 GPUs or 88 Nvidia H100 GPUs.

3 Results

Table 1 shows the performance differences in the Sentence Generation Performance (SGP) and VQA models, according to the size of the model and to the use of markup tokens as additional tokens.

When evaluating models without the special token, the OPT-1.3B model consistently outperformed its counterparts in most metrics. It achieved the highest scores of BLEU-1, BLEU-4, and ROUGE, with values of 0.6980.698, 0.4040.404, and 0.6260.626, respectively. Although the METEOR scores were comparably high across all three models, the OPT-125M slightly surpassed the others with a score of 0.6790.679. Furthermore, the OPT-1.3B model achieved a peak Avg. SGP score of 0.6010.601.

In contrast, using markups as a special token led to a marked decline in METEOR scores across all models. Specifically, the OPT-125M model showed the most significant drop, falling to 0.4550.455. Furthermore, not only METEOR, but other scores also experienced a general decline. Interestingly, for BLEU-1, the OPT-1.3B model with the special token retained a slight advantage, recording a score of 0.6810.681. The Avg. SGP also decreased with the incorporation of the special token. These observations suggest that integrating markup as a special token may adversely impact the language generation capabilities of the models.

3.2 VQA Performance

Focusing on models without additional tokens, the performance in the Yes/No metric remained consistent across all models. In the accuracy of categories, Cat., the OPT-125M model had a slight advantage, achieving a score of 0.7100.710. Interestingly, for Loc. x and Loc. y, none of the models showed impressive results, indicating that the task was particularly challenging.

When the special token was integrated, the OPT-125M model showed notable improvements in several metrics such as Yes/No, Cat., and Cat. & Count, with scores of 0.8200.820, 0.7630.763, and 0.3310.331, respectively. Both the OPT-125M and OPT-1.3B models showed improvement in the RD and Cat. RD metrics. Remarkably, the OPT-125M model exhibited significant gains, with the scores for RD and Cat. RD to 2.6442.644 and 0.5040.504, respectively. The OPT-1.3B model exhibited a slight improvement in most of the VQA metrics upon the addition of markup tokens as a special token.

3.3 Quantifying the difficulty of multiple QAs

In our dataset, following the criteria defined under Objects in Specific Direction, a single sentence encompasses multiple QA tasks that require identification of both the categories of objects and their counts. Table 2 delineates the accuracy rates based on the number of QAs present in a single sentence (n-QA). As the number of QAs increases, tasks become more complex, leading to lower accuracy rates. This decrease is more noticeable in Cat. & count than in Cat.. Given these challenges, there is a pressing need for further research to effectively address multiple QAs in a single sentence.

Conclusion

In this work, we introduced the NuScenes-MQA dataset, which employs a Markup-QA approach where QA is encapsulated within the text using markup. Through the implementation of the Markup-QA scheme, we established a framework that facilitates the simultaneous evaluation of a model’s capabilities in sentence generation and VQA. This dataset empowers the development of vision language models, especially for autonomous driving tasks, by focusing on both descriptive capabilities and precise QA. We have also established a baseline model that provides a starting point to demonstrate the practical value of our approach.

Limitation

While our research offers significant insight, it comes with certain constraints worth noting. First, the dataset has been constructed using a rule-based approach. As a consequence, it may lack the rich diversity often inherent in natural language. This limitation could potentially make it less ideal for training larger models, such as OPT-6.7B, due to potential insufficiency. Furthermore, the limited variety of tasks that address spatial information raises concerns. Specifically, it remains uncertain whether the model truly captures expressions pertaining to positional data. These limitations underscore areas for deeper future research. We believe that future work will overcome these challenges and further advance the field.

References