DePlot: One-shot visual language reasoning by plot-to-table translation

Fangyu Liu, Julian Martin Eisenschlos, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Wenhu Chen, Nigel Collier, Yasemin Altun

Introduction

Multimodal reasoning on visual language such as plots and charts is an extremely complex task. For downstream tasks such as question answering (QA) on plots/charts, a model needs to first extract relevant information from the image, organize them in a sensible manner, then perform reasoning over the entries extracted. Previous studies have proposed end-to-end solutions to such methods (Lee et al., 2023; Liu et al., 2023a). Whilst being an effective solution, end-to-end methods need to be finetuned on large amounts of task data and they still lag behind on queries that require complex reasoning even after finetuning. As an example, the current SOTA model MatCha (Liu et al., 2023a) achieves only 38.2% accuracy on ChartQA (Masry et al., 2022) (human written queries).

In the meantime, large language models (LLMs) such as GPT-3 (Brown et al., 2020) and PaLM (Chowdhery et al., 2022) have demonstrated exceptional few-shot reasoning skills, without requiring expensive human annotations. However, it is an open question how multimodal reasoning tasks could benefit from LLMs. In this work, we propose to decompose the multimodal visual language reasoning problem into: (1) converting the input plot image to a linearized table and (2) passing the linearized table to LLMs for one-shot reasoning.

The key of the method is a modality conversion module called DePlot that maps charts and plots to the underlying data table. While there has been prior works in chart information extraction, they are usually hybrid systems combining complex hand-designed rules, OCR, keypoint detection, and object segmentation modules (Siegel et al., 2016; Luo et al., 2021; Masry et al., 2022). For different types of charts, distinct approaches have been used (Rane et al., 2021; Kato et al., 2022). Besides, there does not exist a unified, consistent, and accurate framework for evaluating chart information extraction – metrics specific to certain types of charts (Siegel et al., 2016) or overly-simplified number matching metrics (Luo et al., 2021) have been used. Our proposed DePlot is an end-to-end image-to-text Transformer model trained with the task of plot-to-table translation. A combination of synthetic and web-crawled charts and plots and their underlying data table are collected as the training corpus. We demonstrate that DePlot significantly outperforms hybrid systems and can uniformly handle all types of charts. To accurately capture plot-to-table systems’ effectiveness (and avoid error propagation to downstream tasks), we propose a novel table matching metric that considers both textual and numeric entries with relative error tolerance, and is invariant to transpositions, row and column permutations.

After accurately translating plot images to texts (as linearized tables), we can pass the output from DePlot in conjunction with a query to LLMs to compute the answer. We take advantage of novel prompting techniques such as Chain of Thoughts (CoT) Wei et al. (2022), Self-Consistency (SC) Wang et al. (2023), and Program of Thoughts (PoT) Chen et al. (2022) to elicit more accurate answers. An illustration of the whole process can be seen in Figure 1.

To summarize, this work has the following contributions: (1) We standardize the plot-to-table task and propose a unified and informative metric for table comparison. (2) We propose a highly-effective modality conversion model DePlot to translate a multimodal task into a language-only task and then leverage LLMs to solve it with just one shot. (3) DePlot+LLM achieves SOTA on ChartQA with just one-shot supervision, outperforming the second best method (which is fully supervised) by 29.4% on human-written queries.

Background

Numerous large pretrained models, either for cross-modal tasks such as CLIP (Radford et al., 2021), or single-modal tasks, such as GPT-3 and PaLM, have been introduced in the past few years. These pretrained models’ strong zero/few-shot inference capabilities have enabled creative solutions to more complex multimodal tasks. Socratic Models (Zeng et al., 2023) combine multimodal pretrained models using multimodal prompts for tasks such as multimodal assistive dialogue and robot perception & planning. Minds’ Eyes (Liu et al., 2023b) converts physical reasoning queries into code that could be executed in physical engines. MAGIC (Su et al., 2022) inserts visual control using CLIP in text generation models for unsupervised image captioning. Similar our work, Yang et al. (2022) also translates natural images into texts and leverage GPT-3 for knowledge-based VQA.

However, all above approaches focus on natural images and the tasks of interest usually only require capturing very basic visual information such as types of objects. Visual language reasoning poses a different set of challenges from natural image reasoning – it requires, first, accurate and detailed information extraction (IE) from complex visual language data (plots and charts in this work); and secondly very strong numerical reasoning skills to answer queries based on information extracted. While end-to-end fully-supervised models struggle to answer complex human-written queries, DePlot when combined with LLMs can outperform the supervised SOTA by 29.4%. This is achieved by decomposing the two key challenges in visual language reasoning into leveraging two strong pretrained models that excel at their own respective tasks.

Zero & few-shot reasoning over tables.

Traditionally, table reasoning tasks are dominated by end-to-end neural models with table-specific architectural designs (Herzig et al., 2020; Yin et al., 2020; Andrejczuk et al., 2022). Recently, there has been a surge in using LLMs to process tables for downstream tasks such as QA. Chen (2023) shows that with just one-shot in-context demonstration, GPT-3 could reach near SOTA performance on table QA datasets, on par with end-to-end models trained with at least thousands of training examples. Beyond pure LLM approaches, Binder (Cheng et al., 2023), Program of Thoughts (Chen et al., 2022), and Program-Aided Language models (Gao et al., 2022) all combine LLMs with compilers/program executors for table reasoning tasks and have achieved SOTA performance. DePlot can be combined with pure LLMs and also any of the aforementioned neural-symbolic methods in a plug-and-play style.

Information extraction from plots and charts.

Prior works on plot/chart IE is usually pipeline-based, combining OCR, object detection/segmentation systems, and hand-crafted rules. Such specialized systems are frequently designed for specific types of graphs, e.g., Kato et al. (2022) for line graphs, and Rane et al. (2021) for bar plots. ChartBERT (Akhtar et al., 2023) adopts an OCR-based method for text extraction from charts and uses two more stages of neural methods for processing the extracted texts. ChartOCR (Luo et al., 2021) is a hybrid system that accepts all types of chart inputs and has been adopted by downstream task models for chart QA (Masry et al., 2022) and summarization (Kantharaj et al., 2022). DePlot, as an end-to-end neural model, outperforms ChartOCR by very large margins on plot-to-table conversion.

Beyond methodology, the evaluation of plot data extraction tasks has traditionally been ununified. Siegel et al. (2016); Luo et al. (2021); Kato et al. (2022) design different metrics for different types of charts and the metrics can be defined upon coordinates, bounding boxes, or keypoints of the graphs’ objects. However, this measures only the intermediate steps of the data extraction process rather than the quality of data extraction itself. We formulate chart data extraction as a plot-to-table translation task since the ultimate goal of chart IE is obtaining the underlying data table. Besides our work, Masry et al. (2022) also considers chart IE as plot-to-table conversion. However, the metric used in Masry et al. (2022) is a number set matching metric, ignoring table structure (i.e., correct organization of the extracted numbers). We propose a better table comparison metric and discuss more in Section 3.

Standardizing the Plot-to-table Task

To perform visual language reasoning, we propose to decompose a visual language reasoning task on plots into two steps: (1) converting plots to texts (in the form of linearized tables) using DePlot and (2) inputing the linearized table to LLMs for reasoning. Accurately performing plot-to-table translation is essential for the downstream visual language reasoning tasks. Plot-to-table is also an important task standalone as it addresses IE from plots/charts, which can benefit applications such as automatic reports and documents digitization. We will standardize the plot-to-table conversion task in Section 3.1 and propose a new metric for evaluating plot-to-table conversion quality. Then in Section 3.2, we introduce the DePlot model and training procedure for performing plot-to-table conversion.

Prior research in table similarity metric is limited. Masry et al. (2022) has introduced a metric based on the graph IE metric proposed in Luo et al. (2021), which we denote Relative Number Set Similarity or RNSS. The metric looks only at the unordered set of numeric entries predicted and measures how the predicted set matches the target set of numbers. In the following, we first introduce RNSS more formally and then discuss our rationales of proposing a more well-rounded metric Relative Mapping Similarity or RMS.

Let the model predicted numbers in table be P={pi}1≤i≤N\mathcal{P}=\{p_{i}\}_{1\leq i\leq N} and numbers in target tables be T={tj}1≤j≤M\mathcal{T}=\{t_{j}\}_{1\leq j\leq M}. We compute the pairwise set of relative distances between them:

However, RNSS has several key limitations: it does not distinguish the position of numbers within the table; it completely ignores all non numeric content; it gives credit to very high relative errors; and it does not distinguish precision versus recall losses in table reconstruction.

In contrast, we argue that a metric to measure similarity between tables should satisfy the following desiderata:

Be invariant to transpositions, as well as permutations of column and rows.

Allow but penalize small errors in numeric or textual values up to a certain threshold.

Clearly reflect losses in precision or recall.

Relative Mapping Similarity (RMS).

In order to address all of these requirements, we propose RMS, which views tables not as sets of numbers but as unordered collection of mappings from row and column headers (r,c)(r,c) to a single value vv, which we write pi=(pir,pic,piv)p_{i}=(p^{r}_{i},p^{c}_{i},p^{v}_{i}) and tj=(tjr,tjc,tjv)t_{j}=(t^{r}_{j},t^{c}_{j},t^{v}_{j}) for each entry in the predicted table P={pi}1≤i≤N\mathcal{P}=\{p_{i}\}_{1\leq i\leq N} and the target table T={tj}1≤j≤M\mathcal{T}=\{t_{j}\}_{1\leq j\leq M} respectively.

Following Biten et al. (2019), the distance between textual entries can be measured with Normalized Levenshtein Distance, or NLτ\texttt{NL}_{\tau} where the variable τ\tau is such that values above τ\tau are set to the maximum of 11, in order to prevent partial credit for very dissimilar texts. Therefore the distance of two keys pip_{i} and tjt_{j} is NLτ(pr∣∣pc,tr∣∣tc)\texttt{NL}_{\tau}\left(p^{r}||p^{c},t^{r}||t^{c}\right) where ∣∣|| denotes string concatenation. The distance between numeric entries is computed using relative distance Dθ(p,t)=min⁡(1,∥p−t∥/∥t∥)\text{D}_{\theta}(p,t)=\min(1,\|p-t\|/\|t\|) and distances above θ\theta are set to the maximum of 11. Combining this two distances we can compute the similarity between two entries in a mapping Dτ,θ(p,t)\text{D}_{\tau,\theta}(p,t) as (1−NLτ(pr∣∣pc,tr∣∣tc))(1−Dθ(pv,tv))\left(1-\texttt{NL}_{\tau}\left(p^{r}||p^{c},t^{r}||t^{c}\right)\right)\left(1-\text{D}_{\theta}\left(p^{v},t^{v}\right)\right). When both the keys and values are similar, the similarity Dτ,θ\text{D}_{\tau,\theta} is close to 11 (close to when dissimilar).

The RMSF1\texttt{RMS}_{\text{F1}} score can be computed the harmonic mean of the precision and recall. Because permutations of columns and rows yield the same set of column header, row header, value entries, the resulting metric is invariant to them. In order to allow for table transpositions, we just consider both the table and its transposed version and return the one that corresponds to highest RMSF1\texttt{RMS}_{\text{F1}} score.

2 Training Plot-to-table Conversion Models

Unlike prior works that combine rule-based heuristics, OCR systems, and object/keypoint segmentation/detection systems (Siegel et al., 2016; Luo et al., 2021; Kato et al., 2022), we propose DePlot as an end-to-end solution to plot information extraction. DePlot is conceptually simple yet can robustly work for all types of charts (line, dot, bar, and pie charts) without requiring type-specific engineering and hybrid components. Specifically, we initialize an image-to-text encode-decoder Transformer model with the architecture and weights of the SOTA visual language model MatCha (Liu et al., 2023a). We continue finetuning the MatCha checkpoint with the task of mapping plots to their underlying data tables. The table is linearized as a textual sequence (markdown format) with | separating cells and \n separating rows. DePlot is trained to generate the table from left to right autoregressively.

The training corpus is a set of parallel plot-table pairs collected similar to Liu et al. (2023a) – both synthetic data and real world plot-table pairs are combined to form a finetuning corpus. Specifically, three sources of plot-table pairs are used: (1) synthetic data generated by Liu et al. (2023a); (2) synthetic data generated by Methani et al. (2020) (also used in PlotQA dataset); (3) real-world data crawled by Masry et al. (2022) (also used in ChartQA). (3) is sourced from four websites, they are statista.com, pewresearch.com, ourworldindata.org, and oecd.org. The three corpora are mixed with the rate of 1:1:1. The size of each can be seen in Liu et al. (2023a). To avoid data leackage in downstream evaluation, only training set charts from the above datasets are used. We call our finetuned checkpoint DePlot.Note that the original MatCha model is also pretrained with the task of plot derendering (which includes plot-to-table), however for a different purpose – i.e., transferring knowledge to downstream finetuning tasks. Our continue finetuning focuses solely on the task of plot-to-table conversion. We also use a much longer sequence length (512 vs. 192) to accommodate long tables.

3 Human Eval of Plot-to-table Metrics

To verify that RMS is indeed more sensitive and robust than previously proposed table comparison metric, we conduct human evaluation to compare RMSF1{}_{\text{F1}} with the previously used RNSS metric. Specifically, we sample 50 plot-table pairs where the tables are predictions of the plot-to-table conversion models (to be introduced in more details in Section 5.2). We score the 50 pairs with RNSS and RMSF1{}_{\text{F1}}. Then we collect human judgment of the table prediction quality from 6 human annotators on the 50 examples.The 6 annotators are all experienced NLP researchers in information extraction with at least a Master’s degree. For each instance, the human annotators are given a plot, the model’s predicted table, and three questions regarding different aspects of the quality of the predicted table. The three questions are (1) “Does the model overgenerate columns/rows or some rows/columns are missing?”, (2) “Are the x, y label/index names, and title correct?”, and (3) “Are numbers close to the true values and associated with the correct column, row labels/indexes?”. Annotators should rate the table from 1–5 (the higher the better). We attach the full annotation form in Appendix B. The final human score for a plot-table pair is the average of the scores across the three questions across all human annotators. We compute the Pearson’s rr and Spearman’s ρ\rho correlations between metric scores and human scores. As shown in Table 1, under both correlation metrics, we can observe a great improvement of RMSF1{}_{\text{F1}} over the baseline RNSS, suggesting that RMSF1{}_{\text{F1}} is a much more sensitive and informative metric for evaluating the model generated tables.

Prompting LLMs for Reasoning

With DePlot introduced in Section 3, we can convert a given chart/plot into its textual form (as a linearized table). We can then construct textual prompts by concatenating the linearized tables and the questions for QA tasks. We follow the typical in-context learning paradigm to prepend a one-shot example before the current prompt.

The full prompts use either Chain-of-Thoughts (CoT) (Wei et al., 2022) or Program-of-Thoughts (PoT) (Chen et al., 2022) and can be seen in Appendix C. They are slightly modified versions of the ones used by Chen (2023) and Chen et al. (2022) for evaluating reasoning on tabular data. Besides CoT prompting, we also explore combining DePlot+LLM with self-consistency (SC) (Wang et al., 2023), which samples a diverse set of reasoning paths and choose the majority-voted answer instead of relying on one greedily-decoded answer as in CoT. In order to simplify performing arithmetic on large numbers, we also tested prompting the models to generate python code that can be passed through an interpreter. In order to do that, we adapt the paradigm from Chen et al. (2022); Gao et al. (2022) to the context of tables. Future work could alternatively take advantage of finetuned tabular QA models such as Herzig et al. (2020) or use LLMs that generate SQL programs Cheng et al. (2023) and might require multiple LLM iterative invocations to perform different atomic operations.

Experiment

We introduce the experimental setup in Section 5.1 and then the results in Section 5.2 including both plot-to-table translation and downstream QA tasks.

DePlot is trained for 10k steps with a maximum sequence length of 512. The other hyperparameters are identical to MatCha pretraining as introduced in Liu et al. (2023a). In DePlot inference we set temperature to (so the output is deterministic). For LLM prompting, in all cases we use temperature of 0.40.4.

Datasets and metrics.

We evaluate on two chart/plot question answering benchmarks ChartQA (Masry et al., 2022) and PlotQA (Methani et al., 2020). ChartQA contains two sets: augmented (aug.) and human where the augmented set is synthetically generated and the human set is human written. Human written queries usually are more diverse and complex, requiring more reasoning while synthetic questions are usually highly templatic. PlotQA is purely synthetic. It contains v1 & v2 sets where v1 is mostly extractive questions and v2 focuses more on numerical reasoning. Both RNSS and RMSF1{}_{\text{F1}} are used for evaluating plot-to-table translation (though we have argued that RMSF1{}_{\text{F1}} is the more informative metric). Following Masry et al. (2022); Methani et al. (2020), exact match accuracy with 5% tolerance on numerical error is used to report all QA numbers.

We list data statistics of plot-to-table training in Table 2. Note that the plot-table pairs are only from ChartQA and PlotQA training sets (not their validation/test sets). The statistics of PlotQA and ChartQA test data are listed in Table 3. Note that we are also using plot-table pairs from the PlotQA test set for evaluating the plot-to-table task (plot-table pairs from v1 and v2 are identical).

Hardware.

We train and evaluate our models using 64 GCP-TPUv3. The training of DePlot can be completed in roughly 5 hours.

Parameters.

DePlot has 282M parameters. FlanPaLM has 540B parameters. Codex and GPT3 have roughly 175B parameters.

2 Main Results

We evaluate plot-to-table conversion against an OCR and keypoint detection based system proposed by Luo et al. (2021) called ChartOCR. This system also relies on multiple hand-crafted rules that depend on the type of chart. We also compare against two PaLI models Chen et al. (2023) (of different input resolutions) finetuned with the same plot-to-table corpus as DePlot. Finally, we compare with the MatCha base model off-the-shelf. The results are shown in Table 4.

On both metrics, DePlot outperforms the baseline ChartOCR by very significant margins. The gap is especially large on RMSF1{}_{\text{F1}} since ChartOCR might suffice to extract numbers from the plot but can struggle to organize the extracted numbers into a structured table with the correct row and column labels. When compared against PaLI and MatCha, DePlot is also better, suggesting that a visual-language-specific architecture/initialization and task-specific finetuning can both boost plot-to-table accuracy. It is also worth noting that PaLI-17B (res. 588) performs much better than the 224-resolution variant, indicating that high input resolution is a key ingredient for chart information extraction.

Downstream tasks.

We list the main results on ChartQA (Masry et al., 2022) and PlotQA (Methani et al., 2020) in Table 5. We evaluate different DePlot+LLM setups. We evaluate chain-of-thoughts (CoT) (Wei et al., 2022) prompts for GPT-3 Brown et al. (2020) (text-davinci-003) and FlanPaLM Chung et al. (2022) (540B). In addition, we use self-consistency (SC) Wang et al. (2023) across 10 predictions. Finally, we use program-of-thoughts (PoT) Chen et al. (2022) to prompt Codex Chen et al. (2021) (code-davinci-002) to generate python snippets that can be subsequently executed by an interpreter to extract an answer.We also evaluated PaLM and FlanPaLM for code generation but found Codex to be more likely to write correct code instead of do the reasoning in comment blocks. Since some reasoning operations are better done by plain language (like computing an argmax) and some by code snippets (like floating point arithmetic), we find optimal results by doing self-consistency across both CoT and PoT predictions.

DePlot+LLM performs especially strong on the ChartQA human set (denoted with “ ”) which contains complex human written queries. Compared with prior SOTA MatCha, DePlot+LLM when combined with FlanPaLM and Codex and Self-Consistency (SC) achieves an improvement of 29.4% (38.2%→\rightarrow67.6%). This is also the best setup for PlotQA. On the heavily synthetic queries from PlotQA v1 and v2 (denoted with “ ”), DePlot+LLM models underperform the end-to-end SOTA MatCha.

In summary, DePlot+LLM significantly outperforms finetuned SOTA on human-written chart QA queries and overall underperforms finetuned SOTA on synthetic QA queries. We believe it is especially important to achieve good performance on the human set as it is much more diverse and reflects the real-world challenges. The results suggest DePlot+LLM’s strong capability in solving novel human queries unseen in demonstration. It is also worth emphasizing again that DePlot+LLM requires much less supervision than the finetuned SOTA methods (one shot vs. tens of thousands of training examples). We will discuss why DePlot+LLM underperforms on PlotQA in error analysis (Section 6.1).

Besides one-shot learning, we also experimented with zero- and few-shot inference. We found the models generally fail without demonstration and few-shot has similar performance as one-shot. After the submission of this paper, we experimented with RLHF-ed LLMs such as ChatGPTopenai.com/blog/chatgpt, GPT-4 (OpenAI, 2023), and Bardbard.google.com, finding that such aligned conversational models are capable of processing the DePlot-generated tables in a zero-shot manner. This can potentially further improve DePlot+LLM’s performance on academic benchmarks by large margins.

Analyses and Discussions

In this section, we first conduct case studies and error analysis in Section 6.1 to better understand DePlot’ wins and losses against end-to-end methods. Then in Section 6.2 we study the performance of DePlot when exposed to out-of-distribution web charts and plots.

To more concretely demonstrate the strengths and weaknesses of DePlot+LLM, we present two case studies for the downstream task ChartQA. We compare DePlot+FlanPaLM using either PoT or CoT.

First, in Table 6 we show an example demonstrating the benefit of using LLM and prompting techniques for stronger numerical reasoning. While the finetuned SOTA MatCha wrongly predicts the answer, DePlot+FlanPaLM (using either CoT or PoT) produces accurately the answer.

Second, we show an example where the DePlot+LLM framework fails in Table 7. The LLMs are unable to accurately identify the “highest value of the gray bar” since they do not have information about the color of bars. In Table 7, though DePlot+FlanPaLM correctly predicted “Yes”, it is correct for the wrong reason – FlanPaLM randomly chose the highest value in light blue bars which also happens to be smaller than the average of “identity theft”. This is a typical failure mode where the query refers to a visual attribute but such attribute is lost in plot-to-table translation. In future work, we plan to develop a table encoding scheme that also considers visual attributes to avoid such errors.

While DePlot+LLM has surpassed finetuned SOTA on ChartQA, we notice that the picture on PlotQA is different – DePlot underperforms finetuned SOTA MatCha by a large margin (66.6% vs. 91.5%). Through error analysis, we observe that there are two major reasons. First, synthetic queries are highly templatic and covers only restricted types of questions. Models finetuned with thousands of examples can learn to solve such templatic questions, even better than humans do (human ceiling on PlotQA is just 80.5% compared with MatCha performance of 91.5%). However, DePlot+LLMs only learn from one-shot in-context example and thus cannot exploit such bias encoded in the training set. The second reason is the loss of information in plot-to-table translation. Synthetic queries are usually highly extractive and include questions asking visual attributes such as color, shape, or direction of objects in a plot. When plots are converted to tables, such information is lost. We plan to also decode visual attributes in future work when training the plot-to-table model.

More successful and failure case analyses are available in Appendix D.

2 Out-of-distribution Analysis

One limitation of our evaluation setup is that the kind and style of charts that are part of DePlot’s training corpus are in the same domain as those in the evaluation sets from ChartQA and PlotQA. This raises the question of whether DePlot will generalize to charts sourced from different websites or built using completely different tools. However, few public resources exist containing both charts and their associated tables.

In order to estimate the out-of-distribution capabilities of DePlot we annotate 10 charts from the recently released TaTa dataset Gehrmann et al. (2022), sourced from dhsprogram.com. We skip choropleth maps since none have been seen during training. We find DePlot obtains an average 78% RMSF1\texttt{RMS}_{\text{F1}} score in reconstructing the underlying tables. We observed two limitations in DePlot which we outline below and can attributed to the nature of the training datasets used. First the model could get distracted by adjacent text, such as references to external sources, and it benefited from cropping the chart in advance. Secondly, DePlot struggled to understand labels or values linked to their corresponding bar/pie section by an arrow. We will address these issues in future work by making the synthetic data creation pipeline more robust.

Conclusion

We have proposed DePlot+LLM, a method for visual language reasoning by decomposing the task into two steps. The first step is converting a plot into linearized table using an image-to-text Transformer model finetuned for the conversion. The second step is combining the plot-to-text model with an off-the-shelf LLM to reason on the linearized table with just one-shot supervision.

We standardize the plot-to-table conversion task by proposing a new table similarity comparison metric that considers the structure of the table and the numeric values but is invariant to column/row permutation. With the new metric, we compare our image-to-text model DePlot’s performance with an OCR-based baseline and three end-to-end baselines, achieving the best improvement. The conversion model is then used for downstream tasks of ChartQA and PlotQA. On ChartQA human-query set, the one-shot DePlot+LLM model achieves +29.4% performance compared with end-to-end SOTA finetuned with thousands of examples. We have also conducted comprehensive analyses to understand the wins and loses of the DePlot+LLM framework and highlight that encoding visual attributes can be a fruitful direction for future exploration.

Limitations

DePlot’s strength is highly dependent on the accuracy of plot-to-text(table) conversion. To obtain effective plot-to-text conversion, large amounts of diverse and in-domain plot-table parallel data are usually needed. It is unknown to which extent DePlot can work for out-of-domain (OOD) plot-to-text conversion. We investigated this in section Section 6.2 but in the future a wider range of web charts can be used to gain a deeper understanding into DePlot’s robustness for OOD plots.

Beyond, DePlot does not work for visual language that does not have a clear latent textual representation such as textbook figures where the visual illustrations are created using specialized software and do not have clear structured representations.

Another limitation of the current DePlot approach is that we ignore any layout information such as orientation and color of the visual elements/objects. In future work, we can incorporate such attributional information by including them in the decoding target.

Ethics Statement

To the best of our knowledge, DePlot is of low risk to the society since it is an information extraction model that converts graphics information from image to textual information in the form of table. That said, when combined with LLMs, DePlot+LLM can demonstrate potential risk such as generating toxic content similar to when LLMs are used standalone. As a result, we should proceed with caution when deploying DePlot+LLM in the real-world and take necessary precautions such as having a filtering stage after the generation.

In terms of data used, all training and evaluation data are either synthetically created using rules or publicly available data on the web with appropriate permissive licenses.

References

Appendix A Details of Baselines

We introduce below the details of the baselines used in Table 5.

T5 is an encode-decoder Transformer model proposed by Raffel et al. (2020). The baseline model T5 takes the concatenation of a linearized table (and a query, when the task is QA) as input, and aims to decode the target (answer or summarization). When the gold table is availible, the gold table is used as the input and the chart image is not used directly. VL-T5 proposed by Cho et al. (2021) is similar to T5 but also takes a visual input (i.e., the chart image) on the encoder side. VisionTaPas (Masry et al., 2022) is modified from TaPas (Herzig et al., 2020) to incorporate the visual modality by adding a ViT model (Dosovitskiy et al., 2021) and cross-modal fusion layers. T5-OCR, VL-T5-OCR, and VisionTaPas-OCR are the same model as T5, VL-T5, and VisionTaPas, respectively. However, they do not assume the existence of gold table but use an OCR-based system to extract the data table from the chart image. The above mentioned models and their performance numbers are all extracted from Masry et al. (2022) and Kantharaj et al. (2022). Please see the original paper for more details. Classification - Regression Chart Transformer (CRCT) (Levy et al., 2022) is the best performing model on PlotQA according to the PlotQA benchmark on paperswithcode.com. It uses a detector that extracts all textual and visual elements of chart then processes these elements with a multimodal Transformer. PaLI (Chen et al., 2023) with 17B parameters is a SOTA on multiple vision-language tasks in the natural image domain however fails significantly on chart understanding tasks. MatCha Liu et al. (2023a) is the strongest supervised baseline and uses a mixture of image-to-text tasks as pretraining to inject math reasoning and chart layout understanding knowledge to the base model. In downstream tasks ChartQA and PlotQA, the full-supervised models are finetuned with the corresponding training sets (ChartQA has ∼\sim33k data points and PlotQA has ∼\sim37M). The fully supervised results are collected from Liu et al. (2023a).

Appendix B Human Evaluation Questions

We list below (Figure 2) the annotation form of the three questions asked when producing the human judgment scores of plot-table pairs. Each question asks one aspect regarding the quality of the generated table and the annotator needs to rate the table from 1 to 5. The final table score is the average score from the three questions.

Appendix C Chain-of-thoughts and Program-of-thoughts Prompt

In Figure 3 we show the one-shot prompt used across all experiments CoT. It is taken from development set examples in combination with prompts used by Chen (2023). We also modified this prompt to output Python code when needing to compute arithmetic operations in Figure 4.

Appendix D More Case Study

In Table 8 and Table 9 we demonstrate two more cases where DePlot+LLM are successful due to its stronger numerical reasoning capabilities.

Failures.

For questions concerning color or other visual attributes of the graph, the DePlot+LLM framework is unable to handle since such information is lost in modality translation and not considered in the current textual table encoding scheme. We show an additional examples in Table 10.

Besides color, plot-to-table conversion can ignore other visual attributes such as the example in Table 11. There does not exist a one-to-one alignment between dots on the line graphs and x labels. The DePlot model produces a table with only x labels and the extrapolated y values and ignore the dots in the graph.