DocLLM: A layout-aware generative language model for multimodal document understanding
Dongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, Xiaomo Liu
Introduction
Documents with rich layouts, including invoices, receipts, contracts, orders, and forms, constitute a significant portion of enterprise corpora. The automatic interpretation and analysis of these documents offer considerable advantages , which has spurred the development of AI-driven solutions. These visually rich documents feature complex layouts, bespoke type-setting, and often exhibit variations in templates, formats and quality. Although Document AI (DocAI) has made tremendous progress in various tasks including extraction, classification and question answering, there remains a significant performance gap in real-world applications. In particular, accuracy, reliability, contextual understanding and generalization to previously unseen domains continues to be a challenge .
Document intelligence is inherently a multi-modal problem with both the text content and visual layout cues being critical to understanding the documents. It requires solutions distinct from conventional large language models such as GPT-3.5 , Llama , Falcon or PaLM that primarily accept text-only inputs and assume that the documents exhibit simple layouts and uniform formatting, which may not be suitable for handling visual documents. Numerous vision-language frameworks that can process documents as images and capture the interactions between textual and visual modalities are available. However, these frameworks necessitate the use of complex vision backbone architectures to encode image information, and they often make use of spatial information as an auxiliary contextual signal .
In this paper we present DocLLM, a light-weight extension to standard LLMs that excels in several visually rich form understanding tasks. Unlike traditional LLMs, it models both spatial layouts and text semantics, and therefore is intrinsically multi-modal. The spatial layout information is incorporated through bounding box coordinates of the text tokens obtained typically using optical character recognition (OCR), and does not rely on any vision encoder component. Consequently, our solution preserves the causal decoder architecture, introduces only a marginal increase in the model size, and has reduced processing times, as it does not rely on a complex vision encoder. We demonstrate that merely including the spatial layout structure is sufficient for various document intelligence tasks such as form understanding, table alignment and visual question answering.
Existing efforts to incorporate spatial layout information typically involve either concatenating spatial and textual embeddings or summing the two . In contrast, we treat the spatial information as a distinct modality and compute its inter-dependency with the text modality in a disentangled manner . In detail, we extend the self-attention mechanism of standard transformers to include new attention scores that capture cross-modal relationships. This is motivated by the observation that there is often a correlation between the content, position and size of the fields in a form. Representing their alignments at various abstraction levels across the transformer layers can enhance document understanding.
A common characteristic of visual documents is their heterogeneous content, irregular layouts, and disjointed text segments. When working with such documents, employing a classical next token prediction objective during the self-supervised pre-training phase can be restrictive. In particular, the preceding tokens may not always be relevant due to the diverse arrangements of text, which can be positioned horizontally, vertically, or even in a staggered manner. To tackle this issue, we propose two modifications to the pre-training objective: (a) adopting cohesive blocks of text that account for broader contexts, and (b) implementing an infilling approach by conditioning the prediction on both preceding and succeeding tokens. Due to these modifications, the model is better equipped to address misaligned text, contextual completions, intricate layouts, and mixed data types. Although text spans and infilling tasks have been studied before , our solution is tailored for visual documents with an emphasis on semantically coherent blocks.
We adapt the pre-trained knowledge of DocLLM for several document intelligence tasks by fine-tuning it on instruction data curated from several datasets. These tasks encompass key information extraction, natural language inference, visual question-answering and document classification. Our instruction-tuning data covers both single and multi-page documents. Layout hints such as field separators, titles and captions can be integrated during instruction-tuning to facilitate learning the logical structure of the documents. We observe that the modifications introduced by DocLLM result in a performance improvement ranging from 15% to 61% for the Llama2-7B model in four out of five previously unseen datasets.
Fig. 1 summarizes the framework. Our contributions include:
A light-weight extension to LLMs designed for understanding visual documents.
A disentangled spatial attention mechanism that captures cross-alignment between text and layout modalities.
An infilling pre-training objective tailored to address irregular layouts effectively.
An instruction-tuning dataset specially curated towards visual document intelligence tasks.
Comprehensive experiments and valuable insights into the model behavior.
Related Work
The remarkable success of ChatGPT has generated substantial research interest in LLMs across academia and industry. Subsequently, numerous LLMs have been introduced starting from text-based LLMs to multimodal LLMs . In this section, we review these recent advances in LLMs and discuss their connection to and distinctions from our work.
Text-based LLMs. The introduction of the transformer model in 2017 has been foundational for the pre-trained models such as BERT , GPT , and T5 , each designed with specific pre-training objectives. The emergence of ChatGPT and GPT-4 marked a notable shift, characterized by a substantial increase in both model parameters and training data size. This enhancement has resulted in remarkable zero-shot generalization capabilities, allowing these models to excel in tasks previously unseen. Such success of LLMs has prompted the development of additional LLMs such as OPT , BLOOM , PaLM , and Llama . Particularly, Llama2 is an open-source LLM that achieves comparable or better performance to both open and closed-sourced models, including ChatGPT, PaLM and Falcon, with enhanced safety strategies. Llama2 employs the standard Transformer architecture with pre-normalization , SwiGLU activation function , and rotary positional embeddings . The pre-training data consists of two trillion tokens from publicly available sources.
Multimodal LLMs. Multimodal LLMs extend the scope of text to diverse modalities, with a focus on visual input. These models can be categorized into two tropes: general-purpose multimodal LLMs and models that are tailored for visually-rich document understanding . The general-purpose multimodal LLMs exhibit promising performance in identifying and reasoning with image information. However, they have not yet been vigorously evaluated on VRDU tasks. As an example, the GPT-4 Technical Report highlights diverse multimodal test cases, such as explaining meme picture distinctiveness, but very few examples are included for visual document use cases. Prior to the advent of large language models, fine-tune-based models relying on vision only were less effective than layout (and vision) modality models in processing visual documents. For example, models like UDOP and LayoutLM outperform vision-only models such as Donut and Pix2Struct in VRDU tasks. But such models require task- and dataset-specific fine-tuning, and are thus excluded in our analysis. The more recent mPLUG-DocOwl and UReader , built upon LLMs, undergo instruction finetuning on a diverse set of VRDU, visual, and textual datasets, and exhibit impressive zero-shot generalization capabilities. Hence, we include those as baselines in our evaluation in Section 4.
Despite the remarkable performance of LLMs, unimodal models aren’t equipped to process multimodal input, and multimodal LLMs rely on complex and memory intensive open-domain vision encoders. Our proposed model, DocLLM, addresses these challenges by explicitly modeling spatial layouts and text semantics, enabling effective comprehension of visual documents. Notably, DocLLM offers an extension to the unimodal architecture by adding the spatial signal to text semantics, avoiding the expensive vision encoder, resulting in a more compact model and efficient processing time.
2 LLM Architectures
Autoregressive Infilling. There are two main autoregressive infilling approaches: “fill-in-the-middle” (FIM) where a single span is sampled, and “blank infilling” with multiple spans.
The OpenAI FIM approach uses the template (prefix, middle, suffix) to divide a document into three segments. Next, these segments are reorganized into (prefix, suffix, middle), enabling the model to predict the middle segment. This process relies on three special tokens, [PRE], [SUF], and [MID], which structure a document as: [PRE] prefix [SUF] suffix [MID] middle. The [MID] token denotes the start for prediction, while the other two special tokens guide the model on where to infill. This method demonstrates that autoregressive models can learn to infill text where the middle part is missing. Fill-in Language Model (FiLM) is a subsequent development that enables flexible generation at arbitrary positions, unconstrained by a predefined generation order. In contrast, approaches like GLM sample multiple spans for infilling. For each blank to be infilled, a pair of special tokens is used: [blank_mask] and [start_to_fill]. The multiple spans not only require special tokens but also global indicators to distinguish which middle span the model should infill. This global indicator is implemented with 1D token positions, ensuring that each pair of the two special tokens, i.e., [blank_mask] and [start_to_fill], share the same positions. We adopt a similar infilling object with the goal to prevent disconnected next-token predictions while avoiding breaking sparse documents into very short segments, e.g., word pieces and/or phrase pieces.
Disentangled attention. Disentangled attention is introduced in the DeBERTa model , where token embeddings and relative positional encodings were kept separate rather than summed together, and each used independently when computing attention weights using disentangled matrices. The motivation behind this was to facilitate the learning of decoupled attention alignments based on content and position separately. This innovation proved effective as it allowed DeBERTa to outperform RoBERTA-large and T5 on NLU benchmarks, as well as to surpass the human baseline on SuperGLUE . In our work, given considerably more complex position encodings used in visually rich documents, disentanglement becomes ever more important to our model’s performance.
DocLLM Framework
In this section, we discuss the architecture of DocLLM and outline the pre-training and instruction tuning procedures. Figure 2 presents an overview of the model architecture.
DocLLM is constructed upon the foundation of an auto-regressive transformer language model following a causal decoder structure. It is composed of stacked transformer blocks, where each block contains a multi-head self-attention layer and a fully connected feed forward network. Standard language models are typically unimodal, accepting only a sequence of text tokens as input. In contrast, DocLLM is a multi-modal system that integrates lightweight visual information by utilizing the spatial positions and dimensions of text tokens obtained using OCR. Simply augmenting the text with bounding box information via additive positional encoding may not capture the intricate relationships between text semantics and spatial layout, especially for visually rich documents . Consequently, we treat the spatial information about the text tokens as a distinct modality. In particular, we use separate vectors to represent these two modalities and extend the self-attention mechanism of the transformer architecture to compute their inter-dependencies in a disentangled manner, as explained in the following section. Furthermore, instead of the traditional left-to-right next token prediction during self-supervised training, we employ a text infilling objective that better leverages contextual information.
2 Disentangled Spatial Attention
It is important to mention that the hidden vectors are reused across different layers, while each layer retains the flexibility to employ different projection matrices. We also note that the number of extra parameters required to encode the bounding box information is significantly lower compared to the overhead introduced by image based models . By simply adding to similar to , we could have avoided using matrices altogether and further reduced the number of parameters. However, it would have irreversibly coupled the layout information with the text semantics. In contrast, our disentangled representation of these modalities in the attention scores enables selective focus when appropriate , thereby providing an optimal balance between model size and effectiveness.
3 Pretraining
DocLLM is first pre-trained in a self-supervised fashion on a large number of unlabeled documents. The self-supervised pre-training objective in autoregressive language models is generally to maximize the log-likelihood of the next token prediction in a sequence based on the context provided by preceding tokens. Let denote all the parameters of the transformer model, including the projection matrices discussed above. The following cross-entropy loss is then typically minimized during the pre-training step:
Visual documents are often sparse and irregular, featuring isolated and disconnected text fragments. In such cases, it is preferable to consider coarse segments of related tokens during pre-training rather than focusing on individual tokens. A segment may represent a coherent chunk of information, similar to a text block, or it can simply be a linear sequence, similar to a text span. In Figure 2, “Name”, “John Doe” , and “Doctor” are all examples of blocks. In general, the broader context provided by multiple tokens in a block can lead to better comprehension.
Furthermore, learning to infill text, where the prediction is conditioned on both prefix and suffix tokens rather than only preceding tokens, can be beneficial. The infilling objectives enable contextually relevant completions, provide robustness to OCR noise or misaligned tokens, and can better handle relationships between various document fields. Hence we modify the standard pre-training objective to predict blocks of text given preceding and following text blocks.
Most OCR engines can provide block level information, which makes it feasible to identify coherent text blocks such as a heading or an addressNote that in order to avoid any leakage of useful information, the block information is only used for the masking objective during pre-training, and is not provided to the model as input. Concretely, masking is performed at the block level, but the model is not provided with information about the number of tokens in a given masked block. Please refer to Figure 2 for an illustrated example.. Inspired by , we follow an autoregressive block infilling objective, where text blocks are randomly masked, and the masked blocks are shuffled and reconstructed in a sequential left-to-right fashion. Block information and block infilling are solely utilized for the pre-training phase, not in instruct-tuning or downstream tasks.
4 Instruction Tuning
Following recent work in the field of VRDU and prior work in NLP , we instruction-tune DocLLM on a variety of instructions derived from DocAI datasets using various templates. Due to the high cost and time intensity of manual data collection, we leave the construction of a VRDU instruction-tuning dataset with crowdsourced instructions and preferences to future work. We employ a total of 16 datasets with their corresponding OCRs, spanning four DocAI tasks: visual question answering (VQA), natural language inference (NLI), key information extraction (KIE), and document classification (CLS).
The diversity of supervised fine tuning (SFT) instructions is critical in helping zero-shot generalization . Thus, we diversify templates per task when possible, with each template asking a different question, and in some cases, expecting different types of answers. We re-use the templates introduced in when applicable, and consider a broader selection of datasets in our instruction-tuning data mix.
We create the templates following what we believe end users would generally ask about documents (Table 1). For KIE and CLS, we hypothesize that (1) the extraction instructions can teach DocLLM to correlate names of keys in the prompts with document fields so as to retrieve values, (2) the internal classification instructions can help the model understand what intrinsically characterizes each key or document type, and (3) the multiple choice question (MCQ) instructions can teach the model to leverage its comprehension of key names included as choices in the prompt (resp. document type names) to classify extracted values (resp. entire documents). We introduce the templates in detail as follows.
Visual Question Answering. We collect DocVQA , WikiTableQuestions (WTQ) , VisualMRC , DUDE , and BizDocsBizDocs is a collection of business entity filings that is due to be released publicly., to compose the VQA instruction-tuning data mix. We use one instruction template to build our SFT inputs for VQA, as shown in table 1. An example prompt derived from DocVQA would read: "{document} What is the deadline for scientific abstract submission for ACOG - 51st annual clinical meeting?"
Natural Language Inference. We only include TabFact in our instruction-tuning data mix for NLI task, due to lack of additional DocAI NLI datasets available. The instruction template is shown in table 1. An example prompt derived from TabFact would read: "{document} \"The UN commission on Korea include 2 Australians.\", Yes or No?"
Key Information Extraction. We gather Kleister Charity (KLC) , CORD , FUNSD , DeepForm , PWC , SROIE , VRDU ad-buy (with random train-test splitting), and BizDocs to build the KIE instruction-tuning data, where we leverage three instruction templates: extraction, internal classification, and MCQ, as shown in 1. For the extraction template, we add the “None” answer if the key does not exist in the given document. To increase diversity in the SFT training data, we also derive internal classification and MCQ instructions from original KIE annotations. To stay consistent with benchmarks from previous work , we only keep the prompts derived from the extraction template in the test split of each KIE dataset. An example extraction instruction derived from KLC would read: "{document} What is the value for the \"charity number\"?"
Document Classification. We aggregate RVL-CDIP and BizDocs to build our CLS instruction-tuning data. We used two types of instruction templates for this task: internal classification and MCQ, as shown in 1. To avoid the cold start problem induced by potentially unseen types of documents in testing or even in production usage, we only keep the MCQ prompts for the test split of each CLS dataset. We also downsample RVL-CDIP in the train split to avoid hindering the other datasets. An example MCQ instruction derived from RVL-CDIP would read: "{document} What type of document is this? Possible answers: [budget, form, file folder, questionnaire]."
Experiments
We gather data for pre-training from two primary sources: (1) IIT-CDIP Test Collection 1.0 and (2) DocBank . IIT-CDIP Test Collection 1.0 encompasses a vast repository of over 5 million documents, comprising more than 16 million document pages. This dataset is derived from documents related to legal proceedings against the tobacco industry during the 1990s. DocBank consists of 500K documents, each featuring distinct layouts and a single page per document. The relevant statistics for the datasets utilized in the pre-training are detailed in Table 2. We obtain a collection of 16.7 million pages comprising a total of 3.8 billion tokens.
We have introduced the datasets used to conduct instruction tuning on Section 3.4. These datasets encompass four common DocAI tasks: VQA, NLI, KIE, and CLS. Note that when a prompt includes a list of possible answers, we create multiple copies of the prompt with one possible answer assigned to each. We only perform this “flattening” operation in the training split of the dataset. Detailed statistics for these tasks are presented in Table 3.
2 Model Setup and Training Details
Table 4 provides key settings and hyperparameters for two variants of DocLLM: DocLLM-1B, which is based on the Falcon-1B architecture , and DocLLM-7B, which is based on the Llama2-7B architecture Since Llama2 does not come with pre-trained weights at 1B parameters, we use the Falcon-1B architecture for the smaller version of DocLLM.. DocLLM-1B is composed of 24 layers, each with 16 attention heads and a hidden size of 1,536. DocLLM-7B comprises 36 layers, 32 heads, and a hidden size of 4,096. Using pre-trained weights as the backbone for the text modality, we extend the Falcon-1B and Llama2-7B models by adding the disentangled attention and block infilling objective as described in Section 3.
For DocLLM-1B, we use a pre-training learning rate of with 1,000 warmup steps, employing a cosine scheduler, and Adam optimizer with and a weight decay of 0.1. For instruction tuning we use a learning rate of with 500 warmup steps and a cosine scheduler, and the same parameters for weight decay and Adam optimizer as the pre-training phase. The Adam epsilon is set to . We pre-train for one epoch, and instruct-tune for a total of 10 epochs.
For DocLLM-7B, pre-training involves a learning rate of with 1,000 warmup steps and cosine scheduler, weight decay of 0.1, and Adam optimizer with . Instruction tuning uses a learning rate of with 500 warmup steps and a cosine scheduler, weight decay of 0.1, and Adam optimizer with . Adam epsilon is set at . We conduct one epoch of pre-training, followed by three epochs of instruct-tuning, considering available computing resources.
The maximum sequence length, or context length, is consistently set to 1,024 for both versions during the entire training process. The DocLLM-7B models are trained with 16-bit mixed precision on 8 24GB A10g GPUs using fully sharded data parallelism, implemented with the accelerate library.https://huggingface.co/docs/accelerate The DocLLM-1B model, on the other hand, is trained on a single 24GB A10g GPU.
3 Downstream Evaluation
Experimental settings. We investigate two experimental settings:
Same Datasets, Different Splits (SDDS): Following previous work in VRDU , we first evaluate DocLLM on the unseen test split (or dev split when test split is unavailable) of each of the 16 datasets composing the instruction-tuning data. The motivation behind this very typical setting is to check how DocLLM performs when tasks and domains supposedly stay the same from train to test.
Same Tasks, Different Datasets (STDD): Following , we also evaluate DocLLM on held-out datasets. More precisely, we instruction-tune the pre-trained checkpoint of DocLLM on prompts from 11 of the 16 datasets considered in SDDS, then evaluate DocLLM on the test split of the remaining three datasets. The rationale behind this evaluation setting is to assess the performance of DocLLM when tasks are unchanged but domains and layouts differ from train to test. We believe examining this setting in the DocAI field is relevant because industry use cases usually encountered in practice revolve around VQA, KIE, and CLS, while document characteristics tend to change more often in production. We specifically isolate DocVQA, KLC, and BizDocs for STDD evaluation in order to (1) exclude at least one dataset per task from SFT when possible, (2) leave enough datapoints per task in the training split of the instruction-tuning data, (3) avoid data leakage (certain datasets were obtained from the same sources), and (4) benchmark models on popular yet challenging datasets when possible. Due to the high cost of instruction-tuning, we were not able to run additional experiments with different held-out datasets.
Baselines. In SDDS and STDD, we benchmark DocLLM against comparably-sized and SOTA LLMs using Zero-Shot (ZS) prompts that contain the text extracted from each document using an OCR engine (excluding the spatial information) . In SDDS, we also report numbers from recent DocAI LLMs evaluated in a similar setting . As motivated in section 2, we do not consider DocAI models that require task-specific fine-tuning and/or dataset specific prompts , and instead focus on LLMs with out-of-the-box instruction following capability.
Metrics. Following previous work , we evaluate all VQA datasets using Average Normalized Levenshtein Similarity (ANLS) , with the exception of VisualMRC, for which we use CIDEr and WTQ, for which we use accuracyThis is done to remain consistent with the results reported by other SotA models.. Performance on all CLS and NLI datasets is measured using accuracy. We evaluate all KIE datasets with the F1 score.
Results. In the SDDS setting, as shown in the Table 5, we observe that DocLLM-7B excels in 12 out of 16 datasets, inclusively compared to ZS results of GPT4 and Llama2, and SDDS results of mPLUG-DocOwl and UReader. Among equivalent models (excluding GPT4), our model outperforms in 14 out of 16 datasets. Specifically, DocLLM demonstrates superior performance in layout-intensive tasks such as KIE and CLS. In VQA and NLI, its performance surpasses that of most multimodal language models, although it underperforms compared to GPT-4. GPT-4 outperforms DocLLM in VQA, possibly due to the higher complexity of reasoning and abstraction involved in VQA datasets compared to tasks like KIE or CLS. DocLLM-1B demonstrates performance close to that of our larger model, suggesting that the smaller model can derive significant benefits from the architecture of DocLLM.
In the STDD setting, our model demonstrates superior performance compared to Llama2 across four out of five datasets, and achieves the best score overall for two of them (KIE task again). DocLLM also outperforms mPLUG-DocOwl on DocVQA and both mPLUG-DocOwl and UReader on KLC, despite both baselines having been instruction-tuned on these datasets. However, it is important to note that classification accuracy is notably lower in our model. This discrepancy may stem from the fact that our model has been trained using only one classification dataset, limiting its ability to generalize effectively to new datasets.
Ablation Studies
We conduct ablation studies to validate the three contributions of DocLLM: (1) disentangled spatial features, (2) the block infilling pre-training objective, and (3) the masking strategy used for decoding.
For all ablations, we use Next Token Prediction (NTP) out-of-sample accuracy to compare configurations at the pre-training stage. Due to resource restrictions, each experiment uses a subset of our pre-training corpus: we randomly sample 100,000 chunks and predict on 1,000 unseen documents. A chunk is a pack of documents concatenated one by one with the total length less than maximum input length. The hyperparameters are set consistently following Table 4 across all ablation experiments.
Disentangled Spatial Attention. To measure the effect of disentangled spatial attention on cross-modal interactions, we train the models by setting the hyperparameter in Eq 6 to or . Table 7 enumerates the attention combinations, and the results suggest that keeping only the spatial-to-spatial interaction (i.e. ) yields the highest NTP accuracy. The performance differences among other configurations, such as text-to-spatial and spatial-to-text, are subtle. Notably, the vanilla text-only self-attention mechanism yields the lowest NTP accuracy, underlining the importance of incorporating spatial features for understanding documents with rich layouts. For all experiments in Section 4, we therefore set , , and . We opt for simplicity by choosing a hard mode over a soft one while acknowledging the potential advantage of flexibility for the latter.
Autoregressive Block Infilling. To evaluate the effectiveness of the proposed autoregressive block infilling objective especially comparing with the conventional left-to-right causal learning, we benchmark three configurations in our ablation study: (1) causal learning, (2) causal learning with spatial modality, and (3) block infilling with spatial modality. As highlighted in Table 8, autoregressive block infilling exhibits the best performance. Additionally, the performance gain of adding the spatial modality to the causal learning proves the advantage of the spatial modality.
Prefix Decoder and Causal Decoder. For document-conditioned generation, an intuitive choice is to employ a prefix decoder with prefix masking to make the whole document bidirectional visible in the attention, as illustrated in Figure 3(b). We investigate this assumption through experiments where we compare a prefix decoder against the conventional causal decoder. Specifically, we conduct contrast experiments on these two decoders for different settings outlined in the disentangled spatial attention to study their resulting performance.
The results in Figure 4 show marginal differences between these two decoder across the five configurations, with the causal decoder having a slight edge over the prefix. The minor difference suggests that both masking methods are comparable in modeling documents. Thus the bidirectional attention enabled by the prefix decoder may not be crucial in this context, and we consequently elect to use a causal decoder for all experiments in section 4.
Discussion and Findings
In addition to its immediate utility in visually rich document understanding tasks, we posit that DocLLM offers an opportunity to change the landscape of generative pre-training by enabling language models to go beyond next token prediction in plain text settings. By accommodating complex layout structures, DocLLM allows for e-books, e-publications, and other documents with rich layouts to be incorporated into the pre-training corpus without requiring extensive preprocessing. The spatial-aware reading approach enables the model to perceive the document as inherently structured knowledge.
Moreover, the multi-page awareness, of both page breaks and document boundaries, enhances the model’s ability to comprehend documents of various lengths. This addresses the limitations of previous smaller multi-modal models (which are mainly for single-page documents) and the existing multimodal LLMs (which are primarily designed for images). In supervised instruction tuning, we can adhere to the established practices used in other works, based on desired outputs such as text or images.
The main concept for a cohesive block is to ensure meaningful infilling during the pre-training phase, preventing disconnected predictions. However, the choice of OCR engines to obtain such cohesive blocks remains an open area for exploration. Practical comparisons with various OCR engines and/or layout parsers are left as future work, as LayoutLMs underscore the importance of accurate OCR for improved VQA results. They leverage the Microsoft Azure API, demonstrating superior performance compared to TesseractOCR, as indicated in the DocVQA leaderboard.https://rrc.cvc.uab.es/?ch=17&com=evaluation&task=1 Consequently, researchers are also encouraged to utilize more accurate OCR engines for potential enhancements, if such resources are available.
We have presented a collection of SDDS results alongside zero-shot outcomes. To mitigate prompt influence in the zero-shot results, a rigorous methodology was implemented. This involves the engagement of three independent prompt engineers, each undergoing five rounds of refinement for zero-shot settings, followed by a series of post-processing techniques to enhance result reliability. The best results are thus obtained from each of the three groups. We still acknowledge the potential for refinement and improvement.
We share some internal training experiences, acknowledging the absence of robust validation. First, we observe that a higher weight decay (e.g., 0.1 versus 0.01) generally improves performance in both pre-training and instruction-tuning. During the instruction tuning phase, a higher initial learning rate, such as 1e-4 versus 5e-5, leads to enhanced performance. Overall, we’ve observed that the cosine scheduler tends to outperform linear or constant schedulers across various settings.
Conclusions
In this paper, we introduced DocLLM, a lightweight extension to traditional large language models, tailored for generative reasoning over documents with rich layouts. Unlike existing multimodal LLMs, DocLLM strategically omits costly image encoders, instead prioritizing bounding box information to effectively capture the spatial layout structure of documents. This is achieved through a disentangled attention approach, decomposing the attention mechanism in classical transformers, and enhancing with cross-alignment between text and spatial modalities in structured documents. Notably, our model addresses the challenges posed by irregular layouts and heterogeneous content by employing a pre-training objective that focuses on learning to infill block texts. We fine-tuned the pre-trained model using a comprehensive instruction dataset. Our evaluation across various document intelligence tasks demonstrates that DocLLM surpasses equivalent models on known tasks for 14 datasets out of 16 and exhibits robust generalization to previously unseen datasets in 4 out of 5 settings, affirming its efficacy in extracting meaningful information from a wide range of visual documents. In future work, we plan to infuse vision into DocLLM in a lightweight manner.
Acknowledgments
This paper was prepared for information purposes by the Artificial Intelligence Research group of JPMorgan Chase & Co and its affiliates (“JP Morgan”), and is not a product of the Research Department of JP Morgan. J.P. Morgan makes no representation and warranty whatsoever and disclaims all liability for the completeness, accuracy or reliability of the information contained herein. This document is not intended as investment research or investment advice, or a recommendation, offer or solicitation for the purchase or sale of any security, financial instrument, financial product or service, or to be used in any way for evaluating the merits of participating in any transaction, and shall not constitute a solicitation under any jurisdiction or to any person, if such solicitation under such jurisdiction or to such person would be unlawful. © 2023 JP Morgan Chase & Co. All rights reserved.