ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Yucheng Han, Chi Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, Hanwang Zhang
Introduction
In the past year, the field of artificial intelligence has undergone remarkable advancements. A key highlight is the emergence of large language models (LLMs) like GPT-4 OpenAI 2023. These models Ouyang et al. 2022; Zeng et al. 2022; Team 2023; Baichuan 2023; Touvron et al. 2023a; Touvron et al. 2023b have demonstrated a remarkable capability to comprehend and generate intricate textual data, opening doors to myriads of applications in both academia and industry. Taking this progress a step further, the introduction of GPT-4V Yang et al. 2023 marked another milestone. It endows LLMs with the ability to interpret visual information, essentially providing them with a vision. As a result, they can now extract and analyze data from images, marking a significant evolution in the capacities of these models.
However, despite the achievements and potentials of models like GPT-4V, the details behind GPT-4V’s architecture remain a mystery. This opacity has given rise to questions within the academic world about the best practices for designing multi-modal LLMs. Notably, pioneering research initiatives, like LLaVA Liu et al. 2023c; Liu et al. 2023b and MiniGPT Zhu et al. 2023; Chen et al. 2023, provide insightful directions in this regard. Their findings suggest that by incorporating visual encoders into existing LLMs and then fine-tuning them using multi-modal instruction-tuning datasets, LLMs can be effectively transformed into multi-modal LLMs. It’s noteworthy that these multi-modal datasets are typically derived from established benchmarks, presenting a cost-effective method for accumulating data required for instruction tuning.
Datasets grounded on established benchmarks, such as COCO Lin et al. 2014, have significantly enhanced the abilities of multi-modal LLMs to interpret everyday photographs adeptly. However, when confronted with specialized visual representations, such as charts, they reveal a noticeable limitation Yang et al. 2023; Liu et al. 2023a. Charts are important visual instruments that translate complex data sets into digestible visual narratives, playing a crucial role in facilitating understanding, shaping insights, and efficiently conveying information. Their pervasive presence, from academic publications to corporate presentations, underscores the essentiality of enhancing the capability of multi-modal LLMs in interpreting charts. Indeed, gathering data specifically to refine instructions for understanding charts presents several challenges. These typically stem from two areas: understanding and generation. An effective chart understanding model should be capable of extracting and summarizing data from various types of charts and making predictions based on this information.
However, most existing datasets Masry et al. 2022; Kantharaj et al. 2022; Methani et al. 2020; Masry et al. 2023 only provide support for simple question-answering or captioning, primarily due to the absence of detailed chart information and annotations that provide a high-level understanding of raw data. The high dependency on manually annotated charts gathered by web crawlers negatively affects the quality of these datasets. Thus, the previous annotating methods could only result in chart datasets with lower quality and less comprehensive annotations. Compared with chart understanding, generating chart figures is a more challenging task for the model because existing deep-learning-based generation methods Ramesh et al. 2021; Rombach et al. 2021 struggle to accurately create images based on instructions. Using Python code to generate charts seems promising which needs the corresponding annotations to supervise models. Most charts obtained from the web are devoid of detailed annotations, making it challenging to annotate the generation code. The absence of code annotations makes it challenging to supervise models in code generation. These issues combined impede the model’s ability to understand charts and learn generation jointly.
To address this, we introduce an adaptive and innovative data collection approach exclusively tailored to chart understanding and generation. At the heart of our methodology is the strategic employment of GPT-4’s robust linguistic and coding capabilities, which facilitate the creation of rich multi-modal datasets. This innovative integration not only optimizes data accuracy but also ensures its wide-ranging diversity. Specifically, our method comprises three main phases: 1) Chart Data Generation. Our strategy for data collection stands out for its flexibility. Rather than limiting data collection to conventional data sources such as the web or existing datasets, we harness the power of GPT-4 to produce synthesized data. By providing specific characteristics such as topics, distributions, and trends, we guide GPT-4 to produce data that is both diverse and precise. 2) Chart Figure Generation. Subsequently, GPT-4’s commendable coding skills are utilized to script chart plots using the open-sourced library, like Matplotlib, given the data and function documentation. The result is a collection of meticulously rendered charts that span various forms, each accurately representing its underlying data. 3) Instruction data generation. Beyond chart rendering, GPT-4 is further employed to interpret and narrate chart content, ensuring a holistic understanding. It is prompted to construct relevant question-answer pairs correlating with the charts. This results in a comprehensive instruction-tuning corpus, amalgamating the narrative texts, question-answer pairs, and source or modified codes of the charts.
A standout feature of our methodology is its flexibility, which diminishes the potential for bias while simultaneously offering scalability. Building on this robust methodology, we’ve crafted a benchmark dataset, which is made available for public access. This dataset stands out, not only for its superior quality but also its unparalleled diversity. A comparative analysis of our benchmark against existing datasets can be viewed in Table 1. To showcase the superiority of our benchmark, we introduced a multi-modal Large Language Model (LLM) named ChartLlama trained with our established benchmarks. Our extensive experiments evaluated on multiple existing benchmark datasets show that our model outperforms previous methods with remarkable advantages and considerably less training data. Additionally, ChartLlama is equipped with several unique capabilities, including the ability to support a wider range of chart types, infer across multiple charts, undertake chart de-rendering tasks, and even edit chart figures.
Our main contributions are summarized as follows:
We introduce a novel multi-modal data collection approach specifically designed for chart understanding and generation. The proposed data collection method boasts superior flexibility and scalability, enabling easy migration to different types of charts and various tasks.
Through our innovative data collection approach, we create a benchmark dataset that stands out in terms of both quality and diversity. We make this dataset publicly available to catalyze further advancements in the field.
We develop ChartLlama, a multi-modal LLM that not only surpasses existing models on various existing benchmarks but also possesses a diverse range of unique chart understanding and generation capabilities.
Related work
The series of LLM models, such as GPT-3.5 Ouyang et al. 2022 and GPT-4 OpenAI 2023, have demonstrated remarkable reasoning and conversational capabilities, which have garnered widespread attention in the academic community. Following closely, a number of open-source LLM Baichuan 2023; Touvron et al. 2023a; Touvron et al. 2023b; Zeng et al. 2022; Bai et al. 2023a models emerged, among which Llama Touvron et al. 2023a and Llama 2 Touvron et al. 2023b are notable representatives. With extensive pre-training on large-scale datasets and carefully designed instruction datasets, these models have also showcased similar understanding and conversational abilities. Subsequently, a series of works have been developed, aiming to achieve specific functionalities by leveraging efficient supervised fine-tuning algorithms based on the Llama series. Among these influential works, Alpaca Taori et al. 2023 and Vicuna Zheng et al. 2023 stand out, with Vicuna’s framework serving as the cornerstone for subsequent multi-modal works.
2 Multi-modal Large Language Model
Concurrently, the academic community has witnessed a surge of development in multi-modal LLMs Li et al. 2023a; Ye et al. 2023; Li et al. 2023b; Li et al. 2023c; Zhang et al. 2023b; Hu et al. 2023; Zhao et al. 2023; Bai et al. 2023b; Zhang et al. 2023a; Liu et al. 2023b; Liu et al. 2023c; Chen et al. 2023; Zhu et al. 2023 built upon existing open-source models. Earlier efforts in this domain, such as LLaVA Liu et al. 2023c, MiniGPT Chen et al. 2023, BLIP2 Li et al. 2023b, and mPLUG-Owl Ye et al. 2023, have shown significant room for improvement in both performance and functionality. With further exploration of training strategies and an increase in dataset scale, the performance of these new models has steadily improved, reaching comparable levels to GPT-4V in specific evaluation metrics. Notably, LLaVA-1.5 Liu et al. 2023b, an iterative version of LLaVA, has gained popularity as a baseline due to its user-friendly training framework, superior performance, and data efficiency. Our work is also based on LLaVA-1.5.
3 Chart Understanding
In evaluations such as the report of GPT-4V Yang et al. 2023 and HallusionBench Liu et al. 2023a, it is evident that current multi-modal LLMs still struggle with complex chart-related problems. There are already some datasets Methani et al. 2020; Kantharaj et al. 2022; Masry et al. 2022 available for evaluating models’ chart understanding capabilities, mainly divided into two categories, each with its own advantages and disadvantages. One category measures through simple question-and-answer tasks, such as ChartQA Masry et al. 2022, which has high-quality questions and answers annotated by humans, and PlotQA Methani et al. 2020, which has lower-quality questions and answers generated through templates. The advantage of these datasets lies in their large scales and the ability to generate them through templates. However, their limitations include the difficulty in ensuring the quality of questions and answers, as well as a tendency to focus too much on simple questions about the data in the charts. The other category converts charts into textual descriptions, with Chart-to-text Kantharaj et al. 2022 being a representative work in this field. The charts and annotations in these datasets are derived from the real world, ensuring higher quality, and encouraging models to delve deeper into the trends and meanings behind the charts. However, the corresponding drawbacks include the presence of more noises in the textual annotations and the over-reliance on BLEU-4. Previous works focusing on chart understanding tasks can be divided into two main kinds of approaches. One kind of approach is using a single model to understand the charts and answer questions in natural language, for example, Masry et al. 2023; Liu et al. 2022b. The other kind of approach, such as Liu et al. 2022a; Xia et al. 2023, is to first utilize the model to convert the charts into structured data and then analyze and answer questions based on the structured data using existing large models. In our work, we primarily explore the former kind, aiming to leverage a single model to complete the entire process of chart understanding.
Method
In this section, we detail our unique approach to chart understanding and generation. Our method involves three interconnected steps: data collection, chart figure generation, and instruction data generation. We illustrate this process in Fig. 3. These steps are detailed in the following subsections.
Our primary goal in chart data collection is to collect diverse and high-quality data. We employ two main strategies for this purpose: 1) Data Generation from Scratch Using GPT-4: To collect a diverse and high-quality dataset, we initially generate tabular data from scratch using GPT-4. We instruct GPT-4 to create data tables based on specific themes, distributions, and other characteristics like the size of the dataset in terms of rows and columns. This process ensured the creation of data with known and controlled characteristics, which can be essential for generating reliable instruction-answer pairs. Moreover, by managing these characteristics, we can intentionally minimize bias, leading to a more balanced dataset. 2) Synthesizing Data from Existing Chart Datasets. Our second strategy is to synthesize data by referencing existing chart datasets. These datasets already encompass a range of topics and characteristics, providing a solid base for data generation. By prompting GPT-4 with these datasets, we guide it to generate reasonable data that complements its existing knowledge base. This method added variety to our dataset and improved its overall quality.
Generating diverse data at scale using the LLM is not an easy task. When the prompt is designed improperly, the model tends to generate repetitive and meaningless data that deviates from the distribution of real-world data and thus lacks valuable insights that could be important for designing meaningful question-and-answer tasks. If we simply provide a set of data and require the model to imitate without any additional guidance, the model will probably just repeat the reference data. Therefore, in this step, it is necessary to provide the model with additional information, such as the topic and distribution, to ensure that it can be properly guided to generate meaningful data. We will now explain these pieces of information in detail.
Chart theme: We first generate hundreds of possible themes, which are all short phrases. When we generate data, we randomly select one from all those themes, which makes the data meaningful and diverse. This also makes it much more easy to generate questions and responses for instruction tuning.
Data trends: Another important characteristic of the data is the trends. We first generate several typical trend descriptions, like steadily increasing and suddenly dropping, then randomly select a few trends and require the model to generate data following them. If lacking such characteristics, the model will tend to generate several sets of data with meaningless distributions.
Column and row lengths: The lengths of columns and lengths are also necessary for data generation. Without specific constraints, LLMs tend to generate excessively long or even repetitive data, which is difficult to present in a meaningful way through charts.
Chart types: Charts of different types usually share different characteristics. For example, the sum of the values in pie charts should be 100%. If not specify the type of chart, we might end up generating data that doesn’t comply with the corresponding chart standards.
2 Chart Figure Generation
The next step is to transform our dataset into visual charts using GPT-4’s coding capabilities. We used popular chart plotting libraries, such as Matplotlib, as our primary tools. When prompting GPT-4, we provide the collected data, relevant function documentation, and in-context examples. We also give detailed instructions on diversifying aspects like color schemes and line types to enhance the visual appeal of the charts. To increase the diversity and success rate of our chart generation, we randomly sample successfully generated codes as in-context examples in the prompts. Compared with previous automated chart generation efforts that relied on templates, our approach offers greater variety and better visual appeal. It also enables us to generalize across different chart types effectively. The result was a collection of meticulously crafted charts, each accurately representing its data and visually appealing, showcasing the effectiveness of our method. The necessary input for the prompts in this stage is listed below.
Chart data: This is the most essential input for the task. The chart data is the information that will be visualized in the chart. Without it, no meaningful chart can be made.
Related function documentation: This is an important reference for generating the Python code. It provides information about the available functions and features that can be used to create the chart. With the documentation, the model could even create charts in new styles that are not in the in-context examples.
In context example: These in-context examples are sampled from pre-selected high-quality code. This helps to facilitate the construction of the Python code. When there is new generated code in high quality, we can save and sample it, which is used as in-context examples later.
Other requirements: To ensure that the final generated code is suitable for batch processing and execution, we also need to include several requirements in the prompt. For example, the data is required to be listed in the code to make the generated code self-contained and executable without the need for external files. We also set the requirements for the title, axis labels, legend, and text annotations. They provide context about what the chart represents and make it easier to understand the data. Without them, the chart can be confusing and difficult to interpret.
3 Instruction data generation
After completing the first two stages, we gathered comprehensive information about each chart, including precise tabular data, various characteristics from various perspectives, and the chart plotting code. Leveraging this rich information, we move on to generating a wide range of instruction-answer data with the assistance of GPT-4, significantly enhancing the capabilities of models trained on this dataset. In addition to fundamental chart understanding functionalities such as Q&A and summarization, our approach allows us to construct instructions and answers for more complex tasks, such as accurate data extraction, detailed chart descriptions, chart code generation, and even chart editing. Compared to previous pipelines for instruction data generation that often rely on human annotation, our methods yield significant time savings while enhancing diversity and quality in the resulting dataset.
Here are more details about the data that needs to be filled into the prompt.
Chart descriptions and raw data: providing these descriptions helps the model understand the context better. The first description helps the model to understand the nature of the data, and the second description assists in understanding the visual representation of the data. The raw data feeds the model with the actual values to base its responses on. All the descriptions and raw data are generated in the first and second stages.
Characteristics to be asked about: This requirement ensures that the model asks diverse and relevant questions about the chart. It prompts the model to explore different features of the data and its representation.
Experiment
Implementation details. We train ChartLlama based on LLaVA-1.5 which provides fundamental abilities crucial for chart understanding and generation, including the OCR functionality. The projection layer and LLM are trained on our proposed dataset. Details of the model architecture and training hyper-parameters can be referred to in our appendix.
Dataset statistics. We show the statistics of our generated dataset in Table 1 and Figure 2. In our instruction-tuning data, Q&A dominates while the other tasks correspond to similar proportions of data. This is mainly because a single chart could be utilized to construct multiple Q&A data. Previous datasets usually gather only three types of charts: bar charts, line charts, and pie charts. Unlike them, we support a wide range of chart types. This is mainly due to the strong flexibility of our data construction method. It’s worth noting that we can continue to expand on more data and chart types in the future.
2 Evaluation Benchmark and Metrics
We evaluate possible models on seven tasks, including both the traditional tasks and novel tasks which verifies that our data generation pipeline has good scalability towards various tasks and chart types.
Traditional Tasks. Three traditional tasks are evaluated, namely ChartQA, Chart-to-text, and Chart-extraction. 1) For ChartQA Masry et al. 2022, we evaluate relaxed accuracy on human and augmentation split, respectively. The question-and-answer data on the human split is more challenging because it includes more questions that require mathematical reasoning. 2) Chart-to-text contains two separate datasets for training and evaluation. BLEU-4 and GPT-4 serve as metrics for evaluation. BLEU-4 is widely used in many NLP tasks. However, when it comes to Chart-to-text, there’s a critical issue. The Chart-to-text datasets contain too few ground-truth references, which means that the results must be very close to the reference targets to achieve high scores. Thus, we have to train ChartLlama on the train split when evaluating using BLEU-4. To facilitate more reasonable evaluations, we propose a new evaluation metric based on GPT-4, referring to the GPTScore Fu et al. 2023. We designed scoring criteria that require the ground-truth reference and raw data as input conditions. Details can be found in the appendix. 3) Chart extraction aims to extract the tabular data from the given chart figure. We follow the evaluation framework of DePlot Liu et al. 2022a and report the Precision and F1 scores on the challenging ChartQA dataset, which also provides the tabular data for each chart figure.
New tasks. In addition to traditional tasks, we have devised four additional innovative tasks, three of which are targeted at chart generation to verify the scalability of novel tasks. 1) Detailed description. This task necessitates a comprehensive description of the given chart figure in a detailed manner, rather than summarizing it briefly. The evaluation metric for detailed description is similar to the evaluation metric in Chart-to-text using GPT-4. We include detailed descriptions of the data and chart figures as conditions for GPT-4 to assist evaluation. Additionally, our evaluation criteria are more exhaustive, outlining various elements that the model under evaluation should generate. These elements include the data characteristics and visual attributes of the chart figures. 2) Chart-to-chart. This task aims to reconstruct the given chart figure. We design comprehensive evaluation metrics for code generation and utilize GPT-4 to measure the quality of the code. For the chart-to-chart task, we evaluated the precision of data, axes, colors, chart types, and titles, rating from 0 to 5. Then we average them as the score for each sample. Finally, we normalize it to a range of 0 to 100 for easier analysis and report the average score across the entire test set. 3) Text-to-chart. The task aims at generating chart figures according to instructions and tabular data. We provide the input instructions and the generated code as conditions for evaluation criteria. The evaluation focuses mainly on visual similarity, completeness, accuracy, and aesthetics. Each standard is equally rated from 1 to 5 points. After averaging and normalization, we get the final score. 4) Chart-editing. The input condition for this task is a chart figure and an instruction describing how to edit the chart. It is expected to create a new figure that has been modified according to instructions based on the given chart figure. The evaluation method for chart-editing uses a similar process to previous chart generation-related tasks. The input conditions include the code of the chart to be modified, instructions, and the generated code of the model. The data accuracy, completeness, aesthetics, and instruction following performance are scored on a scale from 0 to 5. After averaging and normalization, the final result is obtained. For further details, please refer to the appendix.
3 Results
We first compare our methods with existing chart understanding models, such as Pix2Struct Lee et al. 2023, Matcha Liu et al. 2022b, unichart Masry et al. 2023. Then we further construct Baseline* using the same model architecture Liu et al. 2023b as ours, but is trained on the training split of each dataset separately. On traditional tasks, we have also tried to compare with existing multimodal large language models such as InternLM-XComposer Zhang et al. 2023a, MiniGPT-v2 Chen et al. 2023, and vanilla LLaVA Liu et al. 2023b. However, we found the limitation of their instruction-following ability makes it hard to be evaluated by existing metrics.
ChartQA. ChartLlama achieves the best performance on both human and augmented splits of ChartQA Masry et al. 2022 as listed in Table 2. Previous methods typically involved pretraining on larger datasets and then finetuning on the training split of the same datasets to achieve better results, while ChartLlama does it in a zero-shot way after training on our dataset. Notably, although previous methods are trained on the ChartQA’s training split, our method achieves significant advantages using much less data as shown in Table 1. Besides, we also evaluate our model on charts of novel types as shown in Table 4. Our model gains significant improvement towards Unichart and the Baseline*. This shows the superiority of ChartLlama in the ability to understand charts in novel charts.
Chart-to-text. As shown in Table 2 and Table 3, our method consistently outperforms the previous state-of-the-art approaches under different evaluation metrics and splits in Chart-to-text Kantharaj et al. 2022. The improvement in our performance primarily stems from the model’s ability to handle long texts. Previous works often encountered meaningless repetitions at the end of sentences when dealing with relatively longer texts.
Chart extraction. Our model performed the best in this task on ChartQA Masry et al. 2022 as listed in Table 2. ChartLlama has been trained on a variety of instruction-tuning data, which greatly improved its ability to understand chart figures. This is the reason why it can significantly outperform LLaVA-1.5 in terms of performance.
Detailed description. ChartLlama gains significant performance improvement over LLaVA-1.5 which is shown in Table 3. The detailed description task requires the model’s ability to understand image details, which can be significantly improved during the training for tasks related to chart figures.
Chart generation and modification. In Table 3, we compare our method with the original LLaVA-1.5, and we can see that our model gains consistent improvement over three tasks. LLaVA-1.5, which is the base model of ChartLlama, processes strong abilities to follow instructions and generate Python code, and thus also gains reasonable performances on chart generation and modification tasks.
4 Qualitative results
Figure 4 visualizes the chart-to-chart and chart-editing results of ChartLlama and our baseline model LLaVA-1.5. ChartLlama plots with the correct color and chart type, while LLaVA-1.5 cannot guarantee the correctness of color, data value, or chart type. Figure 5 shows the text-to-chart results of ChartLlama and LLaVA-1.5. In the first example, ChartLlama successfully generates a funnel chart following the instructions and plots correct values. But LLaVA-1.5 even cannot draw funnel charts. In the second example, it is obvious that the result of ChartLlama contains more details and adds data values for human convenience. Both two examples show the strong ability of chart-generating and editing abilities of ChartLlama. Our diverse dataset and rich instruction-tuning data have endowed our model with a wide range of practical capabilities.
Conclusion
In this paper, we propose a flexible and robust approach for synthesizing chart images and instruction-tuning data and train a multimodal LLM on the proposed dataset. Our synthesis process consists of three steps: chart data generation, chart figure generation, and instruction data generation. The data generation flow we propose greatly reduces the difficulty of generating chart-related data for models and improves the controllability and diversity of the generated data. Experiments conducted on both traditional datasets and our newly constructed dataset validate the outstanding performance of the multimodal LLM. Thanks to the diverse instruction-tuning data in our dataset, the trained multimodal language model possesses various capabilities that were absent in previous models. Moreover, its ability to comprehend both instructions and figures can easily extend to new categories of chart figures or tasks. We believe that our data generation process can make significant contributions to multimodal LLM in tasks related to chart understanding. Furthermore, it will facilitate the application of similar data generation processes in other domains.
Limitations. The current version of ChartLlama’s vision encoder lacks the ability to handle multilingual OCR tasks, restricting the model’s utility for charts containing non-English text. To overcome this limitation, we are contemplating the creation of a novel vision encoder that boasts proficiency in multilingual OCR tasks.
References
Appendix A Model architecture
To elucidate our training strategies, we provide some clarification about the modifications in LLaVA-1.5 Liu et al. 2023b, and introduce its essential model architectures.
LLaVA-1.5 incorporates CLIP’s vision encoder Radford et al. 2021. The primary distinction is that LLaVA-1.5 employs ViT-L/14@336px, while LLaVA uses ViT-L/14@224px. Another notable alteration concerns the image processor. Eschewing traditional center cropping, LLaVA-1.5 adopts padding as an image pre-processing technique, ensuring that all information in the provided image can be apprehended.
Projection layer:
In LLaVA-1.5, the initial single linear layer is substituted with a two-layer MLP, resulting in improved performance.
Lora Layer:
Based on experiments in Lu et al. 2023; Liu et al. 2023b, implementing Lora Hu et al. 2022 layers is sufficient to achieve performance comparable to full fine-tuning strategies. For the original LLaVA Liu et al. 2023c, Lora layers with a Lora rank of 64 suffice, whereas for LLaVA-1.5 Liu et al. 2023b, the Lora rank needs to exceed 128.
Appendix B Dataset Scale
The model’s training process is broken down into two critical stages: pretraining and fine-tuning. The primary objective of pretraining is to effectively initialize the vision projector while fine-tuning steers the Language Learning Model (LLM) to adhere to the provided instructions.
In the pretraining phase, LLaVA-1.5 utilizes approximately 558k image-caption pairs to train the projection layer. It is anticipated that the vision features will align with the language features to a certain extent. This dataset originates from a subset of around 558K image-text pairs from LAION-CC-SBU, each paired with a BLIP caption.
The fine-tuning phase involves further training of the model on 665k instruction-following data pairs. LLaVA-1.5 manifests an array of capabilities during this stage. The instruction-following data pairs are meticulously generated to encompass the required abilities. To enhance the model’s capacities in varied contexts, additional academic-task-focused Visual Question-Answering (VQA) datasets for VQA, Optical Character Recognition (OCR), and region-level perception are incorporated. The final compilation includes several datasets: OpenKnowledge VQA (OKVQA, A-OKVQA), Region-level VQA (Visual Genome, RefCOCO), and OCR (OCRVQA, TextCaps). A-OKVQA is transformed into multiple-choice questions, employing a specific response formatting prompt: answer by directly specifying the option’s letter from the provided choices.
Appendix C Generation Prompt for ChartLlama
As listed in Figure 9, Figure 10, and Figure 11, we have provided standard prompts for data generation in three stages. The text in black color in the figure denotes the fixed prompt template, while the text in red color brackets requires filling in, which serves to enhance the diversity and controllability of the generated results. The detailed meanings of the different variables have already been discussed in the main text, thus we will not elaborate further.
Appendix D Ablation Study on the Conditions of Generation Prompt for ChartLlama
In order to verify the impact of our proposed generation process on the results, we designed an ablation experiment on the prompt for the second step, diagram construction, which is shown in Table 6. Specifically, we removed the in-context examples and the description of the function, then retested the probability of successful generation. The results show that combining both in-context examples and documentation could significantly improve the successful rate of plotting figures. Also, we observe that the diversity could also improve a lot, which is hard to quantify.
Appendix E Filtering Mechanism
The data generation process may produce some erroneous samples, but filtering and correcting these samples can be challenging because the samples contain figures that cannot be processed by GPT-4. We only performed basic error correction, including checking the data generation format and verifying the correct execution of the code. The data generation format check involves confirming whether the model has separated different data results with different markers according to our requirements. The check for correct code execution involves running the generated plotting script. If this script fails to run, we no longer use the training sample corresponding to that plot. Such basic data screening is sufficient to ensure the quality of the generated dataset. We are also considering incorporating more effective automatic screening mechanisms to avoid contamination of the dataset by poor-quality samples.
Appendix F Evaluation Prompt for ChartLlama
We have prepared five evaluation prompts in total, each tailored for a specific task: chart-to-text in Figure 16, detailed description in Figure 12, chart-to-chart in Figure 13, text-to-chart in Figure 14, and chart-editing in Figure 15. We have designed distinctive scoring criteria for different tasks and provided reference information based on the additional annotations in the dataset. Ultimately, we employed GPT-4 for scoring purposes.
Appendix G Comparison with Multi-modal LLMs
The Table 5 includes a comparison of existing state-of-the-art (SOTA) models, illustrating their respective performances. Interestingly, some models Zhang et al. 2023a show unexpectedly low performance. This outcome is not a consequence of our experimental configuration. Rather, it derives from the fact that these models have not been trained on corresponding instruction-following tasks, which results in outputs that are incompatible with the evaluation framework. We argue that training these models specifically on instruction-following tasks using specific datasets would likely yield improved performance. Another notable observation is the performance gap of Qwen-VL between the ChartQA test splits and the ChartQA on our specially generated charts. Despite being trained on ChartQA, Qwen-VL underperforms on the specially generated charts, underscoring the effectiveness and need for our proposed benchmark. However, the lack of general training scripts provided by many models poses a challenge to our fine-tuning efforts. Nonetheless, our hypothesis finds support in the model LLaVA-1.5. Initially, LLaVA-1.5 performed poorly on the dataset but showed significant improvement when trained on the designated dataset.
Novel Tasks:
We also conducted tests on the newly proposed tasks. However, most of the given dataset cannot generate executable Python code except LLaVA-1.5 Liu et al. 2023b. We speculate that this is because these large multimodal models have been overtrained on visual language datasets, resulting in the loss of their code generation capabilities in Language Learning Models (LLMs); while LLaVA-1.5 adopted a series of optimization measures during its training process. For instance, compared to other large multimodal language models, LLaVA-1.5 has a shorter training time, fewer training parameters, a more moderate dataset scale, and incorporates pure text data during training to maintain the basic capabilities of LLMs. This experiment also suggests that if we expect the model to have a certain level of generalization ability, we should avoid making excessive adjustments to the LLMs. This is also why our ChartLlama model chose to train with fewer parameters.
Appendix H More Qualitative Results
As shown in Figure 6, we compare our ChartLLaMA with Unichart and LLaVA-1.5. The given examples are both related to longer questions and calculations, which is hard for Unichart. What’s more, without the language understanding ability, Unichart even cannot follow complex instructions. In Example 2, the answer of Unichart is even not a percentage. Although LLaVA-1.5 has the ability of OCR and instruction-tuning, it cannot identify which part of the image is related to the question because it has not been trained on chart figures. Thus, it fails in both examples, either.
Chart Extraction.
As depicted in Figure 7, ChartLLaMA also possesses the capability to convert charts into structured data. Both the output results of Unichart and ChartLLaMA are a string of characters and we visualize it as tables for convenience. The first mistake of Unichart is reversing the order of years. Another mistake in Unichart is the persistent output of repetitive and meaningless characters at the end. Meanwhile, our proposed model, ChartLLaMA, benefits from strong language comprehension and output capabilities, which prevent the occurrence of such errors.
Chart Description.
In Figure 8, we visualize the results of Unichart, LLaVA-1.5, and ChartLLaMA on the Chart-to-text task. The results from Unichart contain incorrect values and meaningless repetitions when generating long texts. LLaVA-1.5 performs better for long output sequences due to the strong language understanding and generation abilities of the LLM backbone. However, it suffers from wrong OCR recognition results and hallucinations. Our proposed ChartLLaMA performs best among these three models.