TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
Weichao Zhao, Hao Feng, Qi Liu, Jingqun Tang, Shu Wei, Binghong Wu, Lei Liao, Yongjie Ye, Hao Liu, Wengang Zhou, Houqiang Li, Can Huang
Introduction
With the rapid advancement of digital technology, numerous paper documents must be converted into electronic formats for efficient storage and utilization. Tables, as indispensable components of documents, play a vital role in summarizing facts and quantitative data . The compact yet informative nature of tables makes them advantageous for various applications, thereby attracting widespread research attention toward Visual Table Understanding (VTU). VTU generally encompasses four subtasks: Table Detection (TD), which locates tables within document images; Table Structure Recognition (TSR), which parses the structure of tables in table-centric images; Table Querying (TQ), which recognizes the structure of a table from an entire image at a given location, a task that remains underexplored in the previous works; and Table Question Answering (TQA), which answers questions based on table contents. These tasks pose challenges from various perspectives due to the need for representations at different visual-semantic granularities and hierarchies.
Given the success achieved, many pioneering works have mainly centered on the specific subtask with various task-specific architectures, as shown in Fig. 1 (a). For visual table perception tasks such as TD and TSR, one of most adopted approaches is in the detection manner . In contrast, generative vision-language models are often employed to generate answers conditioned on the semantic content of tables for TQA task. Specifically, Vision Transformers (ViT) pretrained on CLIP or EVA-CLIP , Swin-Transformer , and similar models serve as vision encoders, while language models operate in either encoder-decoder or decoder-only frameworks . Besides, recent fast-growing Large Vision Language Models (LVLMs) have shown their powerful capabilities to perceive and understand visual clues by integrating instruction following of Large Language Models (LLMs) . Despite impressive progress, the status quo begs for a question: “Can we leverage the advantages of LVLMs to solve all the VTU tasks once and for all?”
A straightforward solution would be to train the LVLM directly using all the VTU data. However, aside from the diverse table structure and the various relations of table contents, it remains a nontrivial issue due to two cruxes of table parsing and understanding: (i) discrepancy between the representation formats (two-dimensional structure VS. one-dimensional sequence); (ii) required image resolutions. Although some works represent table structure in markup formats like HTML, XML, Markdown, or LATEX. However, they neglect spatial coordinates for cells and only encode logical relationships implicitly. The generated code contains extensive formatted information from different markup languages, increasing output length and potentially causing parsing issues with illegal grammars.
To attack above issues, we in this paper propose a novel LVLM tailored for comprehensive VTU, TabPedia, to effectively solve all VTU tasks in a unified framework, as shown in Fig. 1 (b). More concretely, we employ dual vision encoders, namely ViT-L and Swin-B , to encode the global and fine-grained local information in the low- and high-resolution formats of the input image respectively, acquiring multi-source visual embeddings. Here, all the involved VTU tasks and multi-source visual embeddings are abstracted as concepts and concept synergy mechanism is implemented by introducing the mediative tokens to the LLM in our model. Thanks to this mechanism, all the concepts in TabPedia can work in synergy flexibly. Quantitative and qualitative experimental results on both table perception and comprehension tasks across various public benchmarks confirm the effectiveness of our proposed TabPedia. To further investigate the potential of our model in more challenging and realistic scenarios, we establish a new and comprehensive table VQA benchmark, ComTQA, featuring round 1,500 images and 9,000 QA pairs.
Our contributions are summarized as follows,
We propose a novel large vision-language model, TabPedia, to integrate various VTU tasks into a unified framework, including TD, TSR, TQ and TQA. Specifically, TabPedia fully leverages the comprehensive capabilities of LLMs to fertilize complex table understanding.
We design a concept synergy mechanism to harmonize both table perception and comprehension tasks. Through introducing the meditative tokens into our framework, TabPedia adaptively enables useful information in multi-source visual embeddings and task instructions, generating accurate and plausible responses.
Extensive quantitative and qualitative experiments validate the effectiveness of our proposed TabPedia across various tasks and benchmarks. To further exploit the potential of our model in more complex scenarios, we build a new table VQA benchmark, ComTQA, involving multiple answers, mathematical calculation and logical reasoning, etc.
Related Work
Table recognition is generally divided into table detection, table structure recognition and table content recognition In our work, table content recognition is beyond our scope.
For TD task, the earliest approaches are rule-based methods for locating tables inside documents . With the rapid advances in deep learning, numerous CNN-based methods show impressive performance. Most of these methods directly adopt top-down object detection frameworks to solve this problem . For instance, Sun et al. adopt Faster R-CNN to detect table boxes and the corresponding corner boxes simultaneously, and then adjust table boundaries according to the detected corners. Some other methods model each document image as a graph and formulate TD as a graph labeling problem . In addition, TATR first applies the transformer-based detector, DETR , to improve the detection accuracy without special customization.
For TSR task, one of the most common modeling approaches is still to regard it as some form of object detection . Among them, DeepDeSRT and TableNet are both representative works exploring semantic segmentation to obtain table cell boundaries. TATR first proposes to utilize DETR for this task. TSRFormer introduces a cross-attention module into the DETR framework to improve the localization accuracy of row/column separators. Some other methods attempt to parse table structure via modeling relationship among different table elements . As the most relevant to our approach, markup generation-based methods directly generate markup (HTML or LaTeX) sequences from raw table images . EDD introduces a cell decoder and a structures decoder to generate HTML codes. OmniParser further integrates three task-specific decoders to enhance the table structure representation.
While the previous methods have achieved promising results on table perceptive tasks, they are still limited in table intricate content understanding. In our work, we jointly exploit table perception and comprehension tasks in a unified framework, concurrently enriching visual table understanding.
2 Large Vision-Language Models
LVLMs aim to equip LLMs with visual comprehension capability. The mainstream approaches attempt to connect visual encoders and LLMs with intermediate modules such as simple Projectors , QFormer , Perceiver Resamplers , achieving visual language understanding through pre-training alignment and instruction fine-tuning. For text-rich document scene, several works propose to enhance the LVLMs’ capabilities in understanding textual elements (text-centric VQA, OCR, text spotting, etc.). Among them, TextMonkey employs shifted window attention and token resampler module to improve the training process. DocOwl-1.5 collects a comprehensive dataset DocStruct4M to support unified structure learning.
Despite achieving extraordinary progress on visual understanding, existing LVLMs still face challenges in two-dimensional table parsing and understanding. In this paper, we propose a unified framework to concurrently achieve table perception and comprehension with the support of LLMs.
3 Additional Tokens
In the trend of Transformer-based approaches, extending the input sequence with special tokens is popularized for various intentions, such as extracting task-specific information , providing extra information or improving model performance . For instance, ViT utilizes [CLS] token for classification. Similarly, DETR proposes object queries for detection. ATR adopts tape tokens to obtain useful information from a memory bank. In addition, the Memory Transformer presents a simple approach to improve translation performance by attaching trainable memory tokens after the token sequence. Darcet et al., further attempt to add extra tokens in ViT-based frameworks, e.g., CLIP and DINOv2 , thus improving visual tasks. In our work, we inherit this spirit and design meditative tokens to enhance TabPedia’s perceptive and comprehensive capability for visual tables.
Method
As shown in Fig 2, we present an overview of TabPedia. The overall training pipeline consists of two phases. Concretely, the pre-training stage aims to align the visual features to the large language model, and the fine-tuning stage focuses on visual table-aware understanding. In the following, we elaborate on the architecture of TabPedia, followed by the exposition of its two training phases.
Low-Resolution Vision Encoder. To keep the overall layout information, the raw image is also resized to a low-resolution one denoted as . We choose the pre-trained CLIP visual encoder ViT-L/14 to encode the low-resolution image with . The output sequence is composed of 256 tokens, each with 512 dimension.
Objective. Since TabPedia is trained to predict the next tokens like other LLMs, it is optimized by maximizing the likelihood of prediction loss at training time.
2 Pre-training
To enable the capable of vision encoders to capture text-rich information from high-resolution images and aligning embedding space with the large language model , we first perform extensive text-aware pre-training. As shown in Fig. 2, we jointly optimize the high-resolution visual encoder with both projectors, while freezing the large language model and low-resolution vision encoder. Specifically, followed by , our pre-training procedure involves a variety of perception tasks, i.e., text detection , recognition , spotting , long-text reading and image captioning . The first four tasks focuses on the various document images, while the last one targets natural scene images. These comprehensive tasks endow the vision encoders of TabPedia to effectively perceive textual and visual information from both document and natural scene images. More detailed pre-training settings about dataset and experiment could be referred to .
3 Table-aware Fine-tuning
Through pre-training, TabPedia could well understand text and structure of diverse document images but cannot follow instructions to perform different table understanding tasks. In order to enhance the model capability of instruction following, we first construct a large-scale dataset for visual table understanding. We will elaborate on the dataset construction in the Sec. 4. Based on this dataset, we introduce four table-related tasks, i.e., TD , TSR , TQ and TQA to simultaneously cultivate the perception and comprehension capabilities. In this stage, we further unfreeze the LLM and fine-tune the entire framework except the low-resolution vision encoder.
Dataset Construction
In this section, we aim to introduce the collected instruction following dataset. The entire data is derived from five public datasets, including PubTab1M , FinTabNet , PubTabNet , WikiTableQuestions (WTQ) and TabFact . Among them, PubTab1M contains two subsets, i.e., PubTab1M-Detection (PubTab1M-Det) and PubTab1M-Structure (PubTab1M-Str). Moreover, since the table images in PubTab1M-Str are cropped from PubTab1M-Det, we transform the annotations of the table structure in PubTab1M-Str into the original images and synthesize a new subset PubTab1M-Syn, which could be utilized for TQ task. The statistical data are summarized in Tab. 2. To ensure the instruction diversity, we generate multiple instructions for each task using GPT3.5 . In Tab. 2, we display one exemplar about user’s question for each table task. We will provide a detailed exposition of them in the following.
Table Detection (TD). As a fundamental task, TD task targets to detect all table locations in a document image. Previous methods mainly utilize DETR or variants of R-CNN to predict numerous overlapping bboxes, that inevitably needs complex post-processing, such as non-maximization suppression (NMS), to generate final results. In contrast, we employ LLM to directly generate the locations of instance tables in the format of “[x1, y1, x2, y2]”, where x1, y1, x2, y2 represent the normalized coordinates of the top-left and bottom-right of the corresponding bbox. Moreover, to facilitate detection results for multiple tables, we split multiple table positions with the special symbol “\n” in the output response. We adopt PubTab1M-Det to perform TD task, where images are collected from PDF documents with different scale and rotation types of tables.
Table Structure Recognition (TSR). The TSR targets to parse table structure in terms of rows, columns and cells. HTML and Markdown codes are mainly two kinds of text sequences used to represent a table. HTML could represent all kinds of tables, with or without cells spanning multiple rows and grids, but they contain massive markup grammars i.e., “
” and “We select the PubTab1M-Str , FinTabNet and PubTabNet to support the TSR task, where tables are collected from scientific and financial articles. These datasets contain pairs of table images and HTML annotations. We convert HTML codes into our designed annotation format using the pre-processing tool offered by .
Table Querying (TQ). Different from recognizing table structure from the cropped table-centric images in TSR task, the TQ task directly parses the table from the original document image based on the given table location. This task is more challenging due to the degradation of the table’s resolution and the interference of other document contents around it. Moreover, this task could potentially be combined with TD task to enable automatic parsing of all table structure information in original images. Therefore, we introduce this task to fully unlock the comprehension capabilities of large language models for visual table understanding. For the annotation of table parsing, we adopt the same format as TSR. Since there is no readily available dataset, we synthesize a large amount of available data based on the annotations from PubTab1M , namely PubTab1M-Syn.
Table Question Answering (TQA). TQA aims to provide precise answers through table understanding and reasoning. For both public TQA datasets, i.e., WTQ and TabFact , the table images are collected from wikipedia tables with pairs of content-related question and answer. Thus, we could directly apply these available data to support this task. However, the images of current TQA data are rendered from text-based tables with variations in background color and font size, resulting in poor generalization in real-world tables. In addition, the TQA data volume lags far behind other tasks. To alleviate these obstacles, we generate numerous TQA data with partial images in FinTabNet and PubTab1M by employing the powerful multi-modal understanding capabilities of Gemini Pro . We provide more detailed descriptions of the procedure in the Appendix A.1
To better evaluate TQA performance of various models on real-world table images, we build a complex TQA dataset (ComTQA) based on test set of FinTabNet and PubTab1M . Compared to WTQ and TabFact, ComTQA has more challenging questions, such as multiple answers, mathematical calculations, and logical reasoning. In total, we annotate 9k high-quality QA pairs from 1.5k images by expert annotation. More statistics about ComTQA could be found in the Appendix A.2.
Experiment
Parameter Settings. For the hyper-parameters in model design, the number of meditative tokens is set to 256. The max length of text sequence is set to 4000 to satisfy task requirements. To implement TabPedia, we adopt a cosine schedule with one-cycle learning rate strategy . In the pre-training phase, the learning rate warms up in the first 2% of the training process and then decreases from the peak rate (1e-3) with batch sizes of 64. In the fine-tuning phase, we set the peak learning rate as 5e-6 with batch sizes of 16. We employ the AdamW optimizer in both phases. All experiments are implemented by PyTorch and trained on 16 A100 GPUs.
Datasets. In order to comprehensively evaluate the capability of TabPedia, we employ multiple benchmarks for each task. For performance assessment, we set the temperature parameter as 0.2 in both quantitative and qualitative evaluations. For TD task, PubTab1M-Det contains 57,125 images for testing. For TSR task, FinTabNet , PubTabNet and PubTab1M-Str are adopted for evaluation with 9,289, 9,115 and 93,834 testing samples, respectively. For TQ task, the synthetic dataset PubTab1M-Syn also provides 47,186 samples for testing. For TQA task, WTQ , TabFact and our annotated ComTQA contain 4,343, 12,722 and 9,070 QA pairs, respectively.
2 Quantitative Results
We conduct quantitative evaluations of current state-of-the-art methods for specific tasks in perception and comprehension, comparing them to our proposed TabPedia.
Evaluation on TD. In Tab. 4, we compare TabPedia with the previous state-of-the-art method, TATR . TATR performs the table detection with two classic visual detection backbones, i.e, DETR and Faster R-CNN . Compared with them, TabPedia outperforms Faster R-CNN with a notable margin and achieves competitive performance with DETR. Notably, since TabPedia directly generates the independent locations of instance tables without densely overlapped bboxes, there are no extra post-processing operations involved, i.e., Non-Maximum Suppression (NMS). This advantage could enable TabPedia to perform more complex table understanding, such as parsing all tables by combining TD and TQ tasks.
Evaluation on TSR. Tab. 4 reports the performance of TSR task compared to end-to-end TSR models on PubTabNet and FinTabNet datasets. Specifically, the OCR-free model Donut is fine-tuned for TSR with the official default training configuration. Although OmniParser integrates multiple visually-situated text parsing tasks into a unified framework, it adopts three isolated decoders to perform different tasks. Compared with OmniParser, TabPedia consistently surpasses it with 4.96% and 3.56% S-TEDS on both datasets, respectively. In Tab. 6, TATR as the task-specific method, shows high performance with the DETR architecture. Our proposed TabPedia, a generic model for tasks involving both perception and comprehension, still achieves comparable performance without the need for complex post-processing. These results highlight the exceptional capability of TabPedia.
Evaluation on TQA. Due to the complex structure of tables and the dense text, the understanding of the table contents remains a challenging issue. To thoroughly evaluate the performance of the understanding of table content and structure, we adopt two public benchmarks, i.e., WTQ and TabFact , and our collected dataset ComTQA, as shown in Tab. 8. On the WTQ and TabFact, TabPedia achieves promising performance among the open and close sources LVLMs. In contrast to existing benchmarks, ComTQA contains real-world table images with more complex questions. It is observed that current LVLMs show poor performance due to the incomplete understanding of real-world table structures. Compared with them, TabPedia achieves the optimal result with a notable margin, which demonstrates the effectiveness of jointly learning perception and comprehension tasks.
3 Qualitative Results
We further conduct qualitative evaluation on TabPedia’s perception and comprehension capabilities. Firstly, we show the perception capability of TabPedia with solely TD and TSR tasks, as illustrated in the first row of Fig. 3. TabPedia accurately generates reliable and formatted results, which are rendered to the original image for better observation. Secondly, TabPedia performs a complex task to directly parsing all table structure information in a document image by integrating instructions of TD and TQ tasks within a multi-round dialogue. As shown in the second row of Fig. 3, the example indicates that TabPedia is capable of exploring more holistic visual table understanding. In the last row, we display the table comprehensive capability of TabPedia. It is observed that the response not only contains concise and reliable answer, but also provides the specific contents in the table to support its answer. Especially, TabPedia even acquires certain math calculation ability to capture the connections among table contents, as shown in the bottom right example in Fig. 3. These results demonstrate Tabpedia’s powerful multimodal comprehension capabilities. We also display more visualization results in the Appendix D.
4 Ablation Studies
In this section, we conduct ablation studies to validate the effectiveness of core settings and components in TabPedia. All experiments are conducted on three datasets across three tasks: PubTab1M-Det , FinTabNet and WTQ .
Impact of meditative tokens. In Tab. 10, we conduct the experiment to investigate the impact of adding meditative tokens in TabPedia. It is observed that adding meditative tokens significantly improves TabPedia’s capabilities of table perception and comprehension. We also provide a detailed analysis of the attention map of meditative tokens in Fig. D5 of Appendix. D.
Impact of vision encoders. As shown in Tab. 10, we explore the impact of different vision encoders. In TabPedia, we propose both vision encoders to capture global and local information of the input image with different resolutions. For the high-resolution encoder, it could extract more intricate information from the text-rich image and achieve better performance than solely utilizing the low-resolution encoder. Furthermore, the low-resolution encoder plays a crucial role in furnishing comprehensive layout information, addressing the constrained receptive field of the high-resolution encoder. The results clearly indicate that the synergy between both vision encoders enhances the extraction of structural and content-related details from tables, which effectively improves perception and comprehension tasks.
Limitation
In this section, we discuss the limitations of our TabPedia. Firstly, since we represent the table structure with regular rectangular boxes, TabPedia is currently not capable of accurately parsing structural information for twisted or distorted tables. Secondly, all images in TQA datasets, including WTQ , TabFact and ComTQA are dominated by tables. Therefore, TabPedia still lacks the capability to directly answer the table question with original document image. In addition, it also exhibits a deficiency in table cell recognition.
Conclusion
In this paper, we propose a novel large vision-language model to unify diverse visual table understanding tasks, namely TabPedia. Specifically, we present a concept synergy mechanism to seamlessly integrate diverse tasks and multi-source visual tokens embedded from dual vision encoders as concepts. This mechanism is implemented by introducing the meditative tokens into the LLM. Then, we fully leverage the capability of LLMs to effectively understand these concepts and generate accurate and plausible responses. Extensive quantitative and qualitative experiments across various public benchmarks validate the effectiveness of our TabPedia. To further investigate the potential of TabPedia, we establish a challenging table VQA dataset, ComTQA, featuring round 9,000 QA pairs.
References
Appendix A More details about TQA datasets
We depict the procedure of collecting QA pairs with an example in Fig. A1. For input image, Gemini Pro is prompted to first recognize the table structure with OCR results in the image, then generate several question and answer pairs according to OCR results. In order to improve the reliability of the generated answers, we leverage various prompting techniques, i.e, Chain-of-Thought and few-shot prompting. According to the specific prompt, Gemini Pro will generate multiple QA pairs for each input image and return them in an agreed-upon format. After obtaining raw responses generated by Gemini Pro, we utilize the regularized matching algorithm and the special character filter in turn to extract available question and answer pairs.
A.2 ComTQA Benchmark
In Tab. A.2, we present the distribution of both data sources within the ComTQA dataset. Concretely, ComTQA comprises a total of 9,070 QA pairs across 1,591 images, averaging 5 questions per image. Different from existing TQA benchmarks , ComTQA contains more complex table questions in real-world table images to assess the robustness of various models. As shown in Fig. A2, we showcase several representative examples, including multiple answers, mathematical calculation and logical inference, which are the question types lacking in previous benchmarks. To this end, we hope that ComTQA could fill this gap and serve as a reasonable benchmark for community development.
Appendix B Annotation in TSR task
We illustrate the object classes utilized in TSR and TQ tasks as shown in Fig. B3. The intersection of each pair of table column and table row objects can be regarded as table grid cells. These objects construct a table’s hierarchical structure through physical overlapped rectangle boxes.
Appendix C Broader Impact
Our proposed model targets to unify multiple visual form comprehension tasks. This technology could help more people with visual impairments access tabular data through cooperating with improved screen readers and other assistive technologies. Moreover, automating table understanding technology could reduce the need for time-consuming manual data entry and correction, freeing up human resources for more complex and creative tasks. To be honest, this technology also brings some negative societal impacts. As more table data is extracted and processed with automatic visual table understanding, there is a heightened risk of sensitive information being mishandled or exposed. It is crucial to ensure robust data privacy measures.
Appendix D More Qualitative Results
Results on in-the-wild cases. For better investigating the generalization of our proposed TabPedia, we randomly select some document images from a document website and illustrate the generation results in Fig. D4. For perception and comprehension tasks, TabPedia generates accurate and reasonable responses in TD, TSR and TQA tasks, which sufficiently proves the robustness of our method for visual table understanding.
Attention map of meditative tokens. In order to analyze the information extraction of meditative tokens for different tasks, we visualized the attention maps of meditative tokens for input instructions with different granularity of visual feature tokens, as shown in Fig. D5. For each task, we select the shallow and deep four-layer attention maps in the LLM for visualization, respectively. The y-axis represents the meditative tokens, while the x-axis represents the sequence of instruction tokens and different granular visual tokens. For perceptive tasks, meditative tokens are densely attentive to most of the input information in the shallow layers, while they showcase diverse attention regions in the deeper layers.This phenomenon illustrates that meditative tokens could adaptively capture task-related information with respect to diverse tasks. For the comprehension task (TQA), meditative tokens show a different attention pattern from perception tasks, which maintain sparse attention with input tokens in the shallow layers. These results validate that our proposed meditative tokens adaptively enable different regions of visual tokens and understand the intention of specific task questions.