mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, Jingren Zhou
Introduction
Leveraging the strong language understanding and generation ability of Large Language Models (LLM) , some recent works have developed Multimodal Large Language Models (MLLMs) for general vision-and-language understanding. By aligning a pre-trained visual encoder (e.g. the ViT/L-14 from CLIP ) and the LLM with a Vision-to-Text (V2T) module, these models present promising performance on understanding general images. However, they still face great challenges with images with rich text information, such as documents, webpages, tables, and charts . This is mainly because the visual encoder and V2T module are trained on general image-text pairs and not specifically optimized to represent the textual and structural information in text-rich images.
Textual information in images manifests with a multitude of visual structures, spanning the simplicity of plain text to the systematic grid layouts of tables and incorporating a spectrum of graphical representations such as pie, line, and bar charts. These elements may appear in isolation or be intricately interwoven within the framework of documents and webpages, reflecting a rich diversity of informational architecture across posters, invoices, infographics, scientific reports, academic and news websites, etc. As shown in Fig. 2, besides the basic textual content, structure information also plays a big role in Visual Document Understanding . With basic abilities to understand general images and comprehend structured texts through the LLM decoder, MLLM has the potential to achieve unified structure learning on text-rich images. For better Visual Document Understanding with MLLMs, some works attempt to design text-reading tasks to strengthen the text recognition ability, but either ignore the structure comprehension or only cover limited domains of text-rich images, such as just webpages or documents . In this work, we first propose to perform unified structure learning on text-rich images for MLLMs across 5 domains: document, webpage, table, chart, and natural image.
For better structural understanding, we first design a simple and effective vision-to-text module, namely H-Reducer. Unlike the Resampler or Q-former which fuses visual features with learnable queries but affects spatial information, the H-Reducer accumulates neighborhood visual features through convolution to keep the relative positional relationships. Compared with V2T modules with only linear layers , it produces much fewer visual features, which is more efficient for LLM to understand high-resolution document images. Considering texts in document images are most organized from left to right, H-Reducer merges visual features at the horizontal level. Our Unified Structure Learning comprises structure-aware parsing tasks and multi-grained text localization tasks. To learn the organization of text contents, the former mainly teaches the model to parse the texts in the image in a structure-aware style, such as using line feeds and spaces to represent the structure of documents or webpages, and using extended Markdown syntax to represent the structure of tables and charts. Multi-grained text localization tasks further enhance the ability to correlate visually situated texts and concrete positions in the image. To support unified structure learning, based on publicly available datasets, we carefully build a comprehensive training set DocStruct4M by constructing structure-aware sequences and multi-grained pairs of text and bounding boxes. The DocOwl 1.5 is trained in a two-stage framework, starting with the Unified Structure Learning and then followed by the Multi-task Tuning among downstream tasks. Finally, to trigger the reasoning ability of MLLM in Visual Document Understanding, we construct a high-quality instruction tuning dataset DocReason25K. By performing joint training on DocReason25K and downstream datasets, DocOwl 1.5-Chat well balance giving a simple answer or detailed explanations.
Our contributions in this work are four-fold:
We first propose Unified Structure Learning on text-rich images for MLLMs and design both structure-aware parsing tasks and multi-grained text localization tasks across 5 domains. A comprehensive dataset DocStruct4M is carefully built to support Unified Structure Learning.
We design a simple and effective vision-to-text module for structure learning and perform extensive experiments to validate its effectiveness.
We construct a high-quality instruction tuning set to trigger the reasoning ability of MLLMs on Visual Document Understanding.
DocOwl 1.5 and DocOwl 1.5-Chat achieves state-of-the-art OCR-free performance on 10 Visual Document Understanding tasks, achieving improvement of more than 10 points on 5/10 tasks among similar-sized models.
Related Work
Visual Document Understanding(VDU), also known as Visually-situated Language Understanding , aims to comprehend images with rich text information. Such images range from documents , tables , charts , natural images to webpage screenshots , where diverse composition of text and visual objects contains a wealth of information. To evaluate the multimodal document understanding performance, the task formats include low-level recognition, e.g. information extraction , and high-level semantic understanding, such as visual question answering , image captioning , and natural language inference . According to whether relying on an off-the-shelf OCR system to recognize texts in the image, models for Visual Document Understanding can be categorized into OCR-dependent models and OCR-free ones . To leverage recognized texts from an OCR system, OCR-dependent models are always trained to align textual and visual inputs. For example, UDOP is pre-trained to recover masked text and layout information given image and retained text as inputs. As for OCR-free methods, training with tasks about text recognition is indispensable. Dount design the text reading task to output continuous text sequences that ignore structure information. To leverage structure information, Pix2Struct designs a Screenshot Parsing Task to generate the HTML DOM tree for webpage screenshots but is hard to apply to other types of images. In this work, we first propose Unified Structure Learning for all image types and carefully build a comprehensive dataset to support layout learning.
Multimodal Large Language Models(MLLM) have shown strong vision understanding and open-ended conversation abilities for natural images. They follow the architecture paradigm of connecting a vision encoder,e.g. ViT , with a Large Language Model(LLM) by a vision-to-text module, such as simple linear layers or a Q-Former /Resampler /Abstractor with learnable queries. To enable MLLMs to comprehend images with rich texts, there are major two challenges: how to encode high-resolution images and how to understand visually-situated texts. To tackle high-resolution images, most works choose to further train or extraly add a high-resolution vision encoder . UReader first proposes to keep the low-resolution vision encoder and use a shape-adaptive cropping module to crop raw images into multiple sub-images with low resolution. To enhance the visually-situated text understanding, some work design tasks of reading texts from top-left to bottom-right without taking into account the importance of structure . CogAgent and DocPedia further try strengthening the layout understanding for documents, webpages, and natural images with text grounding tasks. However, the comprehension of the overall structure is ignored, and tables and charts are not covered. In this work, we follow UReader to process high-resolution images. To strengthen structure understanding, we design structure-aware praising and multi-grained text localization tasks for all types of images, covering documents, tables, charts, webpages, and natural images. We propose a vision-to-text architecture to better maintain spatial information of visual features by convolution. Finally, to support unified structure learning, we build a comprehensive training dataset DocStruct4M and greatly improve the visual document understanding performance.
DocOwl 1.5
DocOwl 1.5 follows the typical architecture of Multimodal Large Language Models, which consists of a visual encoder, a vision-to-text module, and a large language model as the decoder. To better keep the textual and layout information in text-rich images of high resolution, we design an H-Reducer as the vision-to-text module to ensemble horizontal visual features. As shown in Fig. 3(a), to enhance the text recognition and structure understanding abilities, we first perform Unified Structure Learning with structure-aware parsing and multi-grained text localization tasks for all types of images. Then, the model is jointly tuned on multiple downstream tasks of Visual Document understanding.
High-resolution Image Encoding. As proved by previous works , the ability to encode high-resolution images is critical to ensuring that the decoder can use rich text information from document images. As shown in Fig. 3(b), following UReader , we utilize a parameter-free Shape-adaptive Cropping Module to crop a shape-variable high-resolution image into multiple fixed-size sub-images , where is the number of crops. To keep the overall layout information, the raw image is also resized to a low-resolution one as the global image . Then, each image in is independently encoded to a sequence of visual features by a transformer-based Visual Encoder, where is a -dimension vector, is the length of visual features for each image.
Spatial-aware Vision-to-Text Module: H-Reducer. There are two kinds of popular vision-to-text modules for Multimodal Large Language Models: a MLP or a cross-attention module with learnable queries . Both two are not quite suitable for representing high-resolution text-rich images. The former projects complete visual features into the language embedding space. It maintains all spatial information in the document image but keeps the sequence length of raw visual features, which is too long when processing high-resolution images. For example, encoding a 1,344x1,344 image with the ViT/L-14 results in 9,216 visual tokens. The cross-attention module could greatly reduce the length of the visual sequence to the number of learnable queries, but may lose spatial information during semantic fusion.
In this work, we design a more appropriate vision-to-text module for Visual Document Understanding, namely H-Reducer, which not only reduces visual sequence length but also keeps the spatial information. As shown in Fig. 3(b), the H-Reducer is comprised of a convolution layer to reduce sequence length and a fully-connected layer to project visual features to language embedding space. Since most textual information in document images is arranged from left to right, the horizontal text information is usually semantically coherent. Thus, the kernel size and stride size in the convolution layer are set as 1x4 to ensemble horizontal 4 visual features. The output channel is set equal to the input channel . The convolution calculation is as follows:
where represents the dot product with kernel weights on multiple channels. After the convolution layer, the visual features of image are converted to the , the feature length of which is .
Then, with a fully connected layer to align visual features to the language embedding space, the are transferred to .
Multimodal Modeling with LLM. As the decoder of MLLM, large language models should understand both the visual features of images and the textual features of language instructions. Following mPLUG-Owl2 , we apply the Modality-adaptive Module(MAM) in LLM to better distinguish visual and textual inputs. During self-attention, MAM utilizes two sets of linear projection layers to separately perform the key/value projection for visual features and textual features. To help the LLM correlate multiple cropped sub-images, UReader designs learnable crop position embeddings to denote the row and column position in the raw image. In this work, we simply add special textual tokens ‘
where means the concatenation operation, is the crop number of the image, is the textual embeddings of the special textual indicator for the global image or positions of cropped images, is the visual features of a global or cropped image, is the textual embeddings of the instruction, is the predicted answer.
2 Unified Structure Learning
Most Multimodal Large Language Models are trained with image-text pairs of natural images to align the visual encoder with the LLM, such as Conceptual Captions , LAION and COYO . Initializing from such models could inherit the shallow text recognition ability, but is far from understanding complex textual and structural information in various text-rich images. In this work, to empower the comprehensive document understanding abilities of MLLM, we design a Unified Structure Learning across 5 domains, including natural images, documents, tables, charts, and webpages. It involves both structure-aware parsing tasks and multi-grained text localization tasks, as shown in Fig. 4.
Document Parsing. For representing the structure information, Pix2Struct parses webpage screenshots with condensed HTML DOM trees, which are built based on the HTML source codes and are not available for other formats of documents or webpage screenshots, e.g. PDF. In documents or webpages, horizontal and vertical distances between texts form the main layout information. Therefore, to make the structure-aware parsing task applicable to most documents and webpage screenshots, we choose to add extra line feeds(‘’) and spaces into the text sequence to denote different lines and horizontal distances. The greater the horizontal distance, the more space characters.
We choose CCpdf , RVL-CDIP , VisualMRC and datasets encapsulated in DUE-Benchmark (DocVQA , InfoVQA , DeepForm , KLC , WTQ , TabFact ) to support the Document Parsing task. CCpdf is a multi-lingual PDF dataset built upon webpages from Common Cramwlhttps://commoncrawl.org, covering diverse domains of documents, such as industry, academic, and medical. In this work, we mainly focus on English Document Understanding and drop PDFs detected as other languages. RVL-CDIP contains 16 categories of industry documents, such as ‘letter’, ‘email’, and ‘scientific reports’. We further remove some categories with flipping and blurring texts, such as ‘handwritten’ and ‘form’. DUE-Benchmark is a collection of available and reformulated datasets over various document domains and layouts featuring tables, graphs, lists, and infographics. VisualMRC is a webpage screenshot dataset across 35 websites. OCR annotations in VisualMRC are aligned with local regions, thus, we follow them to utilize crops of a screenshot as input for this parsing task. For CCpdf and DUE-Benchmark, a PDF-parsing tool pdfplumberhttps://github.com/jsvine/pdfplumber can be directly used to generate structure-aware text sequence with a PDF page as the input. For RVL-CDIP and VisualMRC, there are no PDF files, just annotations of bounding boxes of texts. As an alternative, akin to the LATIN-Prompt , we insert the line feeds and spaces by calculating and comparing the horizontal and vertical distances of bounding boxes. To avoid too many space characters resulting in sparse texts, we further limit the maximum number of consecutive spaces to 4. This strategy allows us to construct structure-aware text sequences in the same style as pdfplumber.
Table Parsing. Different from documents or webpages, tables are structured in a more standardized way, where row and column correspondences represent key-value pairs. HTML and Markdown codes are mainly two kinds of text sequences used to represent a table. HTML codes can represent all kinds of tables, with or without cells spanning multiple rows and grids, but they contain too many paired labels (e.g. ‘
We choose TURL and PubTabNet to do the structure-aware table parsing task, where tables are collected from Wikipedia pages and scientific articles, respectively. Without cells across rows and columns, tables in TURL can be directly represented with Markdown codes. Due to lacking table images in TURL, we transfer tables into HTML codes and render table images with variations in background color and font size. PubTabNet contains pairs of table images and HTML codes. We convert HTML codes into Markdown style and add ‘ Chart Parsing. Unlike documents and tables, organizing texts in reading order cannot represent the structure of charts. Considering that the chart is a visualization form of the table, parsing charts to tables could best maintain the mathematical characteristics of the chart. This requires the model to understand the structure of the chart and the alignment of the x/y axis. Besides, to keep consistent with the Table Parsing task, we also use Markdown codes to represent the data tables of charts, as shown in Fig. 4. We adopt PlotQA , FigureQA , DVQA , and ChartQA to support the structure-aware chart parsing task. These datasets cover charts on both synthetic data and data from real-world sources . Chart types include vertical bar, horizontal bar, line, dot line, and pie chart. Source data of the chart is provided in the JSON or CSV format , both can be conveniently converted to Markdown codes. However, some raw values are not suitable as standard answers for parsing because there are too many significant digits to be represented on the chart. Therefore, to reduce the difficulty of estimating values and make the model focus more on structural understanding, we keep 4 significant digits for all values. Natural Image Parsing. Quite different from text-dominant images mentioned above, the semantics of natural images is a combination of natural objects and scene texts. Thus, parsing natural images is necessary to organize scene texts and mention the main image content. Manually annotating captions to describe the relationship between objects and scene texts is labour- and financial-intensive. Like TAP , we concatenate the general caption with OCR texts to form the target parsing sequence. We utilize OCR-CC to support the Natural Image Parsing task. OCR-CC is a subset of Conceptual Caption , which contains images with scene texts detected by the Microsoft Azure OCR system. Multi-grained Text Localization. As proved in previous works on general image understanding, semantic comprehension and object grounding tasks can be well unified in a single model. For Visual Document Understanding, structure-aware parsing tasks mainly focus on organizing texts according to the overall structure, while neglecting the correspondence between specific texts and local positions. Correlating texts with the concrete position in images is another basic structure understanding ability for visual documents. To support text position learning, we design two symmetrical tasks, namely Multi-grained Text Grounding and Multi-grained Text Recognition. The former aims to predict the bounding box given the visually-situated texts, while the latter does the opposite. We set four granularities of texts for these two tasks: word, phrase, line, and block. The ‘word’ is the smallest granularity of the bounding box, referring to only 1 word. To ensure that the word is visible and the answer is unique, words that are too small (normalized area < 0.001) and words that appear multiple times in the same image are excluded from candidates. The ‘line’ consists of texts that are judged to be horizontally parallel by vertical distance, and the ‘phrase’ is comprised of multiple adjacent words within the same line. The ‘block’ is a combination of multiple successive lines, ranging from 2 to half of the total lines. The text sequences of word-level and phrase-level question answering are much shorter than the other two. Therefore, in order to learn localization more efficiently, each word-level or phrase-level sample consists of up to 5 question-answer pairs for the same image. As for the representation of bounding boxes, we transfer each continuous value in the normalized bounding box into a discrete position token, ranging from 0 to 999. The bounding box annotation is necessary for constructing samples for Multi-grained Text Localization tasks. Therefore, we take DocVQA, InfoVQA, WTQ, TabFact, DeepForm, KLC, ChartQA, VisualMRC, and TextVQA for this task, across domains of the document, table, chart, webpage, and natural image. Overall, to support the unified structure learning for text-rich images, we build a DocStruct4M dataset by ensembling multiple training sets of publicly available datasets and constructing structure-aware text sequences or text-position pairs as the targets. The form of instructions for each task is very diverse for developing the general instruction-following ability of the model. Fig. 5 shows the detailed statistics of DocStruct4M. Through Unified Structure Learning, models could well understand the structure of diverse document images but cannot follow users’ instructions to do different types of tasks, such as information extraction or image captioning. So, we further perform multi-task fine-tuning to train a generalist of visual document understanding as UReader . As shown in Fig. 3(a), DocOwl 1.5 is trained in a two-stage framework. Considering the LLM has strong comprehension abilities for structured text , we argue that the main limitation of MLLM in visual document understanding is the representation ability of the Visual Encoder and Vision-to-Text module for visually-situated text and structure information. Thus, during the Unified Structure Learning, we freeze the LLM parameters and tune the Visual Encoder and H-Reducer. The MAM is also optimized to help the LLM better distinguish visual features and texts parsed from the image. During the stage of Multi-task Fine-tuning, the model mainly learns how to follow the user’s instructions to give answers based on visually-situated text and structure understanding capabilities acquired in the first stage. Therefore, the Visual Encoder is frozen and other modules are tuned. Existing benchmarks mainly evaluate the document understanding ability by answering the question with simple phrases and neglect detailed explanations. In this work, to better leverage the strong language reasoning ability of Large Language Models on Visual Document Understanding, we build a small instruction-tuning set with detailed explanations on text-rich image understanding, namely DocReason25K. Based on raw questions from DocVQA , InfoVQA , WTQ , VisualMRC , ChartQA and TextVQA , we collect detailed explanations with ChatGPThttps://openai.com/chatgpt. Text contents are dominant information on documents, tables or webpage screenshots. Therefore, for DocVQA, InfoVQA, WTQ, and VisualMRC, we take the structure-aware text sequence of the image as the input to gpt-3.5-turbo-0301 and prompt it to answer the question with simple answers and detailed explanations. As for ChartQA and TextVQA, we take the image as the input and utilize the gpt-4-vision-preview to answer the question with detailed explanations. In order to filter out samples where ChartGPT answers incorrectly, we further prompt gpt-3.5-turbo-0301 to judge whether the answer given by ChartGPT is consistent with the concise human-annotated ground-truth answer. Compared with raw questions in benchmark datasets, questions in DocReason25K are added with a prompt ‘Answer the question with detailed explanation’. Detailed statistics of DocReason25K are presented in Table 1. DocOwl 1.5-Chat is trained by combining downstream datasets with DocReason25K and performing multi-task tuning after Unified Structure Learning. DocOwl 1.5 is initialized from mPLUG-Owl2 , which utilizes the ViT/L-14 as the Visual Encoder and a 7B Large Langauge Model with the Modality Adaptive Module as the language decoder. According to the aspect ratio and resolution, each image is cropped into up to 9 sub-images with a fixed resolution of 448x448. Each sub-image is encoded to 1,024 features by the ViT/L-14 and then reduced to 256 features by the H-Reducer. The model is trained with 12,000 iterations on DocStruct4M, with the learning rate and batch size set as 1e-4 and 1,024. It costs about 128 A100 days. During the Multi-task finetuning, the model is trained for 6,500 iterations with the batch size set as 256 and the learning rate set as 2e-5. This further costs about 24 A100 days. We evaluate the Visual Document Understanding performance on 10 text-rich image benchmarks, covering documents (DocVQA , InfoVQA , DeepForm , KLC ), tables (WTQ , TabFact ), charts (ChartQA ), natural images (TextVQA , TextCaps ), and webpage screenshots (VisualMRC ). We compare DocOwl 1.5 with state-of-the-art OCR-free models, including both Multimodal Large Language Models adapted for recognizing texts and much smaller models trained only for document understanding. The detailed comparison of model settings can be found in Table 2. As shown in Table 3, previous MLLMs with more than 7B parameters underperform domain-specific models with less than 1B parameters, showing that the document understanding is still a shortcoming for existing MLLMs. Our DocOwl 1.5 outperforms both domain-specific models and MLLMs with similar sizes on all 10 benchmarks. This validates that DocOwl 1.5 is much stronger on visual document understanding across 5 domains, covering visual question answering, information retrieval, natural language inference, and image captioning tasks. Besides, with much fewer unnatural data (3M vs 9M) and parameters (8.1B vs 17.3B), DocOwl 1.5 outperforms CogAgent on InfoVQA and ChartQA, and achieves comparable performance on DocVQA. This suggests that our unified structure learning with DocStruct4M is more efficient in learning printed text recognition and how to analyze documents. However, our model still underperforms CogAgent on TextVQA, which requires the ability of scene text recognition and general knowledge about natural objects. The primary reason is that scene texts are more diverse in shapes than printed texts and CogAgent is trained on 98M samples of scene text recognition from LAION-2B and COYO-700M , much more than the natural images (1M) in DocStruct4M. In this work, we mainly focus on improving the unified structure comprehension of visual documents and leave further scaling up data on natural scenes as future work. Finally, DocOwl 1.5-Chat can also be evaluated on these concise-answer benchmarks by removing the prompt of detailed explanation. It achieves comparable or slightly better performance than DocOwl 1.5, showing that a small amount of detailed explanatory data may better help the model understand the semantics of text-rich images. As shown in Table 4, we further perform a comprehensive ablation study to validate the effectiveness of our H-Reducer and Unified Structure Learning. Firstly, initializing from a stronger general MLLMs brings better performance on text-rich images (r2 vs r1), showing general vision-and-language knowledge benefits visual document understanding. Tuning the visual encoder during multi-task fine-tuning significantly improves the document understanding performance (r3 vs r2). This suggests that the visual representation of document images may be the main shortcoming of MLLMs and inspires us to design Unified Structure Learning to enhance the representation ability of the visual encoder for visually situated texts and structure. Effectiveness of H-Reducer. When using the Shape-adaptive Cropping Module, the image resolution supported by the MLLM is the product of the cropping number and basic resolution of each crop. With the Abstractor as the vision-to-text module, reducing the cropping number causes an obvious performance decrease (r4 vs r3) on documents. However, with a smaller cropping number, the H-Reducer achieves better performance than the Abstractor (r5 vs r3), showing that is an acceptable resolution for existing benchmarks and the H-Reducer is stronger on maintaining rich text information during vision-and-language feature alignment. Besides, we further compare different settings of the merging shape in the convolution layer. With the same number of merged tokens, the model with the 1x4 merging shape achieves better performance than the one with the 2x2 merging shape on document and table datasets but slightly worse performance on chart understanding (r6 vs r5). This is consistent with the common sense that documents and tables mainly organize texts in the left-to-right order while the semantic structures of charts are much more flexible. A square merging shape is more suited to encode visual features in the form of bars, lines, or pies while the 1x4 merging shape is more appropriate for general document understanding. As shown in r7-r9, further extending the 1x4 merging shape horizontally and vertically decreases the length of visual features but at the cost of performance degradation. Considering the overall performance on all text-rich images, we finally choose the 1x4 as the merging shape in H-Reducer. Effectiveness of Unified Structure Learning. After determining the vision-to-text module, we perform two-stage training with Unified Structure Learning. With only the structure-aware parsing tasks, there is significant improvement across different domains (r10 vs r5). This validates that fine-tuning the visual encoder and H-Reducer with structure-aware parsing tasks greatly helps MLLMs understand text-rich images. Further tuning the parameters of LLM brings slight improvement (r11 vs r10), suggesting that general language knowledge is not the main obstacle to visual document understanding. By replacing the learnable crop position embeddings with special textual tokens, the model achieves better performance (r12 vs r11), showing that the LLM can well understand the relative positions of multiple cropped images with just simple textual indicators. Finally, by introducing Multi-grained Text Localization tasks, DocOwl 1.5 achieves the best performance, validating that correlating visually situated texts with concrete positions helps comprehend documents more accurately. Effectiveness of the Two-stage Training. As shown in Table 5, instead of two-stage training, we also try one-stage joint training of the structure learning and downstream tasks and gradually increase the samples from DocStruct4M. The epoch is gradually reduced because we didn’t observe performance improvements with more iterations. For joint training, the model improves significantly on DocVQA as the samples of Unified Structure Learning increase when it is below 1M. However, as the Unified Structure Learning samples are further increased, the improvement of the model becomes subtle and its performance is not as good as the one using two-stage training. This shows that the two-stage training could better enhance basic text recognition and structure parsing abilities and is more beneficial and efficient for downstream document understanding. Besides proving the effectiveness of H-Reducer through downstream text-rich image understanding performance in Table 4, we further directly compare the text localization performance after the Unified Structure Learning to validate its superiority in preserving spatial features. We build a text localization evaluation set DocLocal4K with 4,250 samples balanced on 4 granularities and covering both text recognition and text grounding tasks. The detailed statistics of DocLocal4K are shown in Table 6. Considering that document images are much more diverse and complex than other images, there are more samples in this domain than others. The IOU@0.5 is used to evaluate the text grounding performance. As for text recognition, the word, phrase, line, and block granularity is evaluated with BLEU1, BLEU2, BLEU3, and BLEU4 , respectively. As shown in Table 7, when trained with the same iterations, the H-Reducer achieves much better performance on both Text Recognition and Text Grounding tasks, showing that H-Reducer with the 1x4 merging shape helps the LLM better understand concrete positions in images. Question Answering with Simple Phrases. Besides quantitative results, we further present some qualitative results of visual document understanding on different domains of images. As shown in Fig. 6(a) and (b), both models answer the question with texts in the image. DocOwl 1.5 can better understand the structure of two documents and give correct answers. In Fig. 6(c), due to the learning of parsing chart with Markdown codes, DocOwl 1.5 can better understand the chart and successfully correlate the x/y axis. Fig. 6(d) shows that although inconsistent with the ground truth, DocOwl 1.5 gives another correct answer with the help of stronger structure understanding on tables. Question Answering with Detailed Explanations. Fig. 7 and Fig. 8 present qualitative results of detailed explanations. Through a small amount of reasoning training, DocOwl 1.5-Chat can well inherit the reasoning ability of LLM and provide detailed explanations about the answer. However, as presented in Fig. 8(c), like most general Multimoal large Language Models , DocOwl 1.5-Chat may also suffer from the hallucination problem in Visual Document Understanding. In this work, we mainly focus on enhancing the unified structure understanding ability of MLLMs and leave how to resolve the hallucination problem in OCR-free document understanding as future work. Structure-aware Parsing. As shown in Fig. 9, DocOwl 1.5 could parse a document image by using line feeds and spaces to represent the structure of text contents. Besides parsing the whole document, as shown in Fig. 10, it could also parse texts from the middle of the image according to human instruction. Fig. 11 presents qualitative results of structure-aware table parsing through extended Markdown syntax on tables with cells spanning multiple columns or not. Furthermore, Fig. 12 shows some cases of parsing different types of charts into Markdown codes, including vertical bar, horizontal bar, pie, and line charts. When all data points are presented in the chart, DocOwl 1.5 can accurately align statistic objects with corresponding numbers. It makes some mistakes in Fig. 12(d) because estimating the concrete numbers is quite challenging when no data points are provided. Finally, as shown in Fig. 13, DocOwl 1.5 can both describe the content of natural images and read scene texts. Multi-grained Text Localization. Fig. 14 and Fig. 15 show qualitative results of text grounding and text recognition at granularities of word, phrase, line and block. The image domains range from documents, webpages, charts, and tables to natural images. To enhance the Visual Document Understanding performance of Multimodal Large Language Models, we first propose Unified Structure Learning across 5 domains of text-rich images, including both structure-aware parsing tasks and multi-grained text localization tasks. To better maintain structure and spatial information during vision-and-language feature alignment, we design a simple and effective vision-to-text module, named H-Reducer. It mainly utilizes a convolution layer to aggregate horizontally neighboring visual features. To support the Unified Structure Learning, we build a training dataset DocStruct4M by collecting publicly available images and carefully constructing structure-aware text sequences and multi-grained pairs of texts and bounding boxes. With Unified Structure Learning, our model DocOwl 1.5 achieves state-of-the-art OCR-free performance on 10 visual document understanding benchmarks.’ label. 3 Multi-task Fine-tuning
4 Training Paradigm
DocOwl 1.5-Chat
Experiments
2 Main Results
3 Ablation Study
4 Text Localization Evaluation
5 Qualitative Results
Conclusion
References