Unifying Vision, Text, and Layout for Universal Document Processing

Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, Mohit Bansal

Introduction

Document Artificial Intelligence studies information extraction, understanding, and analysis of digital documents, e.g., business invoices, tax forms, academic papers, etc. It is a multimodal task where text is structurally embedded in documents, together with other vision information like symbols, figures, and style. Different from classic vision-language research, document data have a 2D spatial layout: text content is structurally spread around in different locations based on diverse document types and formats (e.g., invoices vs. tax forms); formatted data such as figures, tables and plots are laid out across the document. Hence, effectively and efficiently modeling and understanding the layout is vital for document information extraction and content understanding, for example, title/signature extraction, fraudulent check detection, table processing, document classification, and automatic data entry from documents.

Document AI has unique challenges that set it apart from other vision-language domains. For instance, the cross-modal interactions between text and visual modalities are much stronger here than in regular vision-language data, because the text modality is visually-situated in an image. Moreover, downstream tasks are diverse in domains and paradigms, e.g., document question answering , layout detection, classification , information extraction , etc. This gives rises to two challenges: (1) how to utilize the strong correlation between image, text and layout modalities and unify them to model the document as a whole? (2) how can the model efficiently and effectively learn diverse vision, text, and layout tasks across different domains?

There has been remarkable progress in Document AI in recent years . Most of these model paradigms are similar to traditional vision-language frameworks: one line of work inherits vision-language models that encode images with a vision network (e.g., vision transformer) and feed the encodings to the multimodal encoder along with text ; another line of work uses one joint encoder for both text and image . Some models regard documents as text-only inputs . In these works, the layout modality is represented as shallow positional embeddings, e.g., adding a 2D positional embedding to text embeddings. The strong correlation between modalities inherent in document data are not fully exploited. Also to perform different tasks, many models have to use task-specific heads, which is inefficient and requires manual design for each task.

To address these challenges, we propose Universal Document Processing (UDOP), a foundation Document AI model that unifies vision, text, and layout and different document tasks. Different from regarding image and document text as two separate inputs in previous works, in UDOP we propose to model them with the uniform layout-induced representation (Sec. 3.1): in the input stage, we add embeddings of text tokens with the features of the image patch where the tokens are located. This simple and novel layout-induced representation greatly enhances the interaction between the text and vision modalities.

Besides the layout-induced representation, to form a uniform paradigm for different vision, text, layout tasks, UDOP first builds a homogeneous vocabulary for texts and document layout that converts layout, i.e. bounding boxes, to discretized tokens. Second, we propose Vision-Text-Layout (VTL) Transformer, consisting of a modality-agnostic encoder, text-layout decoder and vision decoder. VTL Transformer allows UDOP to jointly encode and decode vision, text, and layout. UDOP unites all downstream tasks with a sequence-to-sequence generation framework.

Besides the challenges of modalities unification and task paradigms discussed above, another issue is previous works utilized self-supervised learning objectives that were originally designed for single-modality learning, e.g., masked language modeling, or classical vision-language pretraining, e.g., contrastive learning. We, on the other hand, propose novel self-supervised learning objectives designed to allow holistic document learning, including layout modeling, text and layout reconstruction, and vision recognition that account for text, vision and layout modeling together (Sec. 4). Besides sequential generation, UDOP can also generate vision documents by leveraging masked autoencoders (MAE) by reconstructing the document image from text and layout modalities. With such generation capacity, UDOP is the first document AI model to achieve high-quality customizable, joint document editing and generation.

Finally, our uniform sequence-to-sequence generation framework enables us to conveniently incorporate all major document supervised learning tasks to pretraining, i.e., document layout analysis, information extraction, document classification, document Q&A, and Table QA/NLI, despite their significant differences in task and data format. In contrast, pretraining in previous document AI works is constrained to unlabeled data only (or using one single auxiliary supervised dataset such as FUNSD ), while abundant labeled datasets with high quality supervision signals are ignored due to the lack of modeling flexibility. Overall, UDOP is pretrained on 11M public unlabeled documents, together with 11 supervised datasets of 1.8M examples. Ablation study in Table 4 shows that UDOP only pretrained with the proposed self-supervised objectives exhibits great improvements over previous models, and adding the supervised data to pretraining further improves the performance.

We evaluate UDOP on FUNSD , CORD , RVL-CDIP , DocVQA , and DUE-Benchmark . UDOP ranks the 1st place on the DUE-Benchmark leaderboard with 7 tasks, and also achieves SOTA on CORD, hence making UDOP a powerful and unified foundation Document AI model for diverse document understanding tasks,

To summarize, our major contributions include:

1. Unified representations and modeling for vision, text and layout modalities in document AI.

2. Unified all document tasks to the sequence-to-sequence generation framework.

3. Combined novel self-supervised objectives with supervised datasets in pretraining for unified document pretraining.

4. UDOP can process and generate text, vision, and layout modalities together, which to the best of our knowledge is first one in the field of document AI.

5. UDOP is a foundation model for Document AI, achieving SOTA on 8 tasks with significant margins.

Related Work

Unifying Model Architectures in Multimodal Learning. Unifying model architectures for different modalities, such as vision, language, and speech, is an emergent direction. Inspired by the immense success in natural language processing, computer vision and speech processing, model architectures in multimodal learning is converging to Transformers. One type of works concatenates text token embeddings and projected image patches as the input to a multimodal Transformer. Other models uses two-tower or three-tower architecture where each modality is encoded respectively. Projection heads or fusion networks on top of the two-tower architecture generate multimodal representations .

Unifying Tasks with the Generative Framework. Research on unifying training processes across different tasks and domains recently has made significant progress. finetunes language models with instructions on 1.8k tasks. unifies several vision-language tasks by converting training objectives to sequence generation. further combines more tasks, e.g., image generation, by converting images and bounding boxes to discrete tokens.

Document Artificial Intelligence. LayoutLM pretrains BERT models on document data with masked language modeling and document classification task, with 2D positional information and image embeddings integrated. Subsequent works also adopt VL-BERT alike architecture and includes additional pretraining tasks, e.g., masked image/region modeling proposed, and leverages the reading order in layout information . use a multimodal encoder to model region features extracted by CNN with sentence-level text representations and train with self-supervised objectives. proposes an OCR-free model to directly generate textual output from document images. trains generative language models on both unlabeled and labeled document data using generative training objectives. proposed to model documents as collections of tokens bounding boxes.

Universal Document Processing

We introduce UDOP, a novel document AI framework with unified learning objectives and model architecture for text, vision, and layout as shown in Figure 1. In this section, we will concretely discuss the proposed Vision-Text-Layout Transformer in UDOP, and will introduce the unified generative pretraining method in the next section. In document processing, given a document image v{\bm{v}}, typically optical character recognition (OCR) is used on v{\bm{v}} to extract text tokens {si}\{s_{i}\} in the document and their bounding boxes {(xi1,yi1,xi2,yi2)}\{(x^{1}_{i},y^{1}_{i},x^{2}_{i},y^{2}_{i})\}, i.e., the layout information for each token. (xi1,yi1)(x^{1}_{i},y^{1}_{i}) and (xi2,yi2)(x^{2}_{i},y^{2}_{i}) respectively represent the coordinates of the left-upper and right-bottom corner of the bounding box. Thus, suppose we have MM word tokens, the input is the triple, (v,{si}i=1M,{(xi1,yi1,xi2,yi2)}i=1M)({\bm{v}},\{s_{i}\}^{M}_{i=1},\{(x^{1}_{i},y^{1}_{i},x^{2}_{i},y^{2}_{i})\}^{M}_{i=1}). Figure 1 shows an example document (left) and downstream tasks (right).

We fuse the vision, text, and layout modalities in the input stage using one unified transformer encoder. For traditional vision-text data, the text modality is usually the high-level description of the corresponding image or task prompt (e.g., question). While in document images, text is embedded inside the image, i.e., text and image pixels have one-to-one correspondence. To leverage this correspondence, we propose a new Vision-Text-Layout (VTL) Transformer architecture to dynamically fuse and unite the image pixels and text tokens based on the layout information.

Next, we build a unified representation for vision, text, and layout as shown in Figure 2. We define the layout indicator function ϕ\phi of image patch and token embeddings as follows:

Then for each text token embedding si{\bm{s}}_{i}, the joint representation is the sum of its image patch featureSome text token like manually crafted prompts have no locations. So, we set their layout bounding boxes to be (0,0,0,0)(0,0,0,0), i.e., they fall into a pseudo image patch. and the text feature:

For image patches vj{\bm{v}}_{j} without any text tokens, i.e. ∀i, ϕ(si,vj)=0\forall i,\,\phi({\bm{s}}_{i},{\bm{v}}_{j})=0, the joint representation, vj′{\bm{v}}^{\prime}_{j} is itself:

Note we do not have a designated joint representation for image patch containing tokens, since features of these image patches are already integrated with the text embeddings. Then {si′}\{{\bm{s}}_{i}^{\prime}\} and {vj′}\{{\bm{v}}_{j}^{\prime}\} are fed into the VTL transformer encoder. These joint representations greatly enhance the interaction between vision, text and layout in the model input stage by explicitly leveraging their spatial correlations.

To further unify layout and text representation, inspired by the recent progress in generative object detection , we discretize the layout modality, i.e., continuous coordinates text bounding box, to layout tokens. Suppose we have bounding box (xi1,yi1,xi2,yi2)(x^{1}_{i},y^{1}_{i},x^{2}_{i},y^{2}_{i}) normalized in $.Theresultinglayouttokenwillbeeachcoordinatemultipliedbyvocabularysizeandthenroundedtonearestinteger.Forexample,ifwehaveboundingbox. The resulting layout token will be each coordinate multiplied by vocabulary size and then rounded to nearest integer. For example, if we have bounding box(0.1,0.2,0.5,0.6)withlayoutvocabularysizewith layout vocabulary size500$, the layout tokens will then be <50><100><250><300>. Layout tokens can be conveniently inserted into text context, and elegantly used for layout generation tasks (e.g., location detection). More details are discussed in Section 4.

Position Bias.

We follow TILT to encode 2D text token position as 2D relative attention bias, similar to the relative attention bias used in T5. However, unlike T5, TILT, or transformer models in previous Document AI works , we do not use 1D position embeddings in VTL transformer encoder, since the joint embedding and the 2D position bias already incorporate the layout structure of the input document.

2 Vision-Text-Layout Decoder

As introduced in the previous section, the VTL encoder is able to compactly and jointly encode vision, text, and their layout. To perform various document generative tasks (will be discussed in Section 4), the VTL decoder is designed to jointly generate all vision, text, and layout modalities.

The VTL decoder consists of a text-layout decoder and a vision decoder, as shown in Figure 1 (middle). The text-layout decoder is a uni-directional Transformer decoder to generate text and layout tokens in a sequence-to-sequence manner. For the vision decoder, we adopt the decoder of MAE and directly generate the image pixels with text and layout information. Details of the image decoding process will be discussed in the segment “Masked Image Reconstruction with Text and Layout ” of Section 4.1. Both text-layout decoder and vision decoder will cross-attend to the VTL encoder.

Information such as model configurations are presented in Section 5.1.

Unified Generative Pretraining

To unify across different training objectives and datasets, we create a universal generative task format with task prompt. We pretrain UDOP on large-scale documents with and without human labels. We summarize the tasks prompts and targets in Table 1 which includes all self-supervised and supervised tasks respectively in upper and lower blocks.

We propose various innovative self-supervised learning objectives for unlabeled documents. The unlabeled document contains OCR text inputs with token-level bounding boxes and the document image. In the rest of this subsection, we use the following input text as example: “Ship Date to Retail: Week of March 14, 1994”

(1) Joint Text-Layout Reconstruction requires the model to reconstruct the missing texts and locate them in the document image. Concretely, we mask a percentage of text tokens and ask the model to both the tokens and their bounding boxes (i.e. layout tokens). E.g., assume masking “Ship Date” and “of”, the input sequence and target sequence is given below:

Here and denote the text-layout sentinel tokens, <100><350><118><372> and <100><370><118><382>” represent the layout tokens of “Date to” and “of” respectively. We use masking ratio 15% similar to Masked Language Modeling (MLM) as this task can be interpreted as masked text-layout modeling.

(2) Layout Modeling asks the model to predict positions of (group of) text tokens, given the document image and context text. E.g., to predict positions of “Ship Date” and “of”, the input sequence and target sequence is given below:

Note this pretraining task has a different sentinel token, , from the previous task “Joint Text-Layout Reconstruction” because the generation content is different (layout vs. text + layout). We use large masking ratio 75% since masking with small ratio results in an easy task.

(3) Visual Text Recognition identifies text at given location in the image. E.g., to recognize the text tokens at <100><350><118><372> and <100><370><118><382>, the input and target is:

Note this pretraining task also has a different sentinel token, . We use masking ratio 50% to distinguish this task from “Joint Text-Layout Reconstruction” and set the layout (bounding box) of sentinel token, e.g., , and layout token, e.g., <0><10><2><20>, to (0,0,0,0). This objective helps model learn joint vision-text embedding by understanding vision-text correspondence.

(4) Masked Image Reconstruction with Text and Layout aims to reconstruct image with text and layout as shown in Figure 3. We adopt the MAE objective for vision self-supervised learning. Originally, MAE masks a percentage of the image patches and feed non-masked patches into a vision encoder. It then feeds encoder outputs to a vision decoder to reconstruct masked patches. MAE uses mean squared error and apply loss only on masked patches. We make the following modifications to the MAE decoding process to customize it for document image generation and our task unification framework:

(4.a) Cross-Attention with Character Embeddings. In document, the textual content mostly consists of alphabetic characters, numbers and punctuation. The character-level composition of text tokens should be helpful for the vision generation. We add cross-attention in the vision decoder that it attends to both the text token encoder features and embeddings of characters in the token (Figure 3 left upper). These characters embeddings are trainable parameters and not encoded by the encoder. This cross-attention with characters only adds linear computation complexity but considerably improves the image generation quality.

(4.b) Image Decoding. Next, we describe the MAE decoding process. For UDOP, we cannot directly feed the unified encoder output to the vision decoder, since the joint vision-text embedding only contains non-masked image patches to the unified encoder (Section 3.1), and image patches are fused with text tokens. Therefore, we propose that the vision decoder takes in a sequence of trainable placeholder embeddings. The length and order of the placeholder sequence is same as the patches of target image. We use two types of placeholder embeddings to indicate whether the image patch is masked in the input document image. The vision decoder attends to encoder vision-text output AND character embeddings via cross-attention. The above process is illustrated in Figure 3. We show the high quality generation visualization in Section 6.1.

2 Supervised Pretraining Tasks

Self-supervised tasks leverage large-scale unlabeled data to learn robust representations. On the other hand, supervised tasks use labeled data for fine-grained model supervision. We include the following supervised tasks in pretraining: document classification, layout analysis, information extraction, question answering, and document natural language inference. Details of the following supervised dataset are in Appendix D. Note that we do not conduct self-supervised tasks on the supervised datasets since we already have large-scale and diverse unlabeled data. Note that the validation or test set of downstream tasks is not used in supervised pretraining.

Classification. The task is to predict the document type. The task prompt is “Document Classification on (Dataset Name)” like “Document Classification on RVLCDIP”, then followed by text tokens. The target is the document class. We use RVL-CDIP with 16 document categories.

Layout Analysis. This task is to predict locations of an entity in the document like title, paragraph, etc. The task prompt is “Layout Analysis on (Dataset Name)”, then followed by the entity name. The target are all bounding boxes that cover the given entity. We use PubLayNet .

Information Extraction. This task predict the entity type and location of a text query (e.g., the abstract paragraph). The task prompt is “Information Extraction on (Dataset Name) (Text Query)”. The target is the entity label and the bounding box of each token of the query. We use DocBank , Kleister Charity (KLC) , PWC , and DeepForm .

Question Answering. The task is to answer a given question associated with the document image. The task prompt is “Question Answering on (Dataset Name)”, then followed by the question and all document tokens. The target is the answer. We use WebSRC , VisualMRC , DocVQA , InfographicsVQA , and WTQ (WikiTableQuestions) .

Document NLI. Document Natural Language Inference predicts the entailment relationship between two sentences in a document. The prompt is “Document Natural Language Inference on (Dataset Name)”, then followed by the sentence pair. The target is the “Entailment” or ”Not Entailment”. We use TabFact for this task.

Experimental Setup

Model Configuration. In UDOP, the unified encoder and text-layout decoder follows the encoder-decoder architecture of T5-large . The vision decoder is MAE-large decoder . Overall UDOP has 794M trainable parameters. For tokenizer, we use T5 tokenizer and embedding from Hugging Face Transformers . We also extend the vocabulary to accommodate special tokens (e.g., new sentinel and layout tokens).

Data. For self-supervised learning, we use IIT-CDIP Test Collection 1.0 , a large-scale document collections commonly-used in previous works . It contain 11 million scanned document with contains text and token-level bounding boxes extracted by OCR. Supervised datasets are as introduced in Section 4.2.

Curriculum Learning. We use large image resolution, 10241024, in our final settings since low resolution makes document text unidentifiable for both detection and generation. It will result in (1024/16)2=4096(1024/16)^{2}=4096 image patch sequence length which takes longer training time than small image resolution, e.g., 224224. Therefore, we use curriculum learning to start from a relatively small resolution and gradually scale up to 1024 resolution. In practice, we use scale with 3 resolutions during the pretraining 224→512→1024224\rightarrow 512\rightarrow 1024. We show the performance of the 3 stages in Appendix E.

Training. We use Adam optimizer with learning rate 5e-5, 1000 warmup steps, batch size 512, weight decay of 1e-2, β1=0.9\beta_{1}=0.9, and β2=0.98\beta_{2}=0.98. For each curriculum learning stage, we train for 1 epoch.

2 Downstream Evaluations

We report the results on FUNSD , CORD , RVL-CDIP , and DocVQA in Table 3 and describe their respective settings in below. We also report the results on 7 datasets of DUE-Benchmark in Table 2. Finetuning training details are available in Section D.6 and performance variance is available in Table 9 and Table 10. Note that for all downstream tasks, we use the original OCR annotations provided in the datasets.

FUNSD (Form Understanding in Noisy Scanned Documents ) has 149 and 50 samples for train and test. We evaluate on the entity recognition task: predicting the entity, "question", "answer", "header", or "other", for the text token. The task format is, suppose we have the title, "The Title", and its entity "[I-Header]", then the encoder input is "The Title" and the generation target is "The Title [I-Header]". The metric is F1 scores.

CORD (Consolidated Receipt Dataset for Post-OCR Parsing) is a key information extraction dataset with 30 labels under 4 categories such as "total" or "subtotal". It has 1,000 receipt samples. The train, validation, and test splits contain 800, 100, and 100 samples respectively. The metric is F1 and the task format is the same as FUNSD.

RVL-CDIP is the document classification dataset that we have discussed previously. It has 320k/40k/40k images for training/validation/test. The metric is classification accuracy.

DUE-Benchmark contains 7 datasets and 3 domains, including document question answering (DocVQA , InfographicsVQA), key information extraction (KLC, PWC, DeepForm), and Table QA/NLI (WTQ, TabFact). Task prompt formats can be found in Section 4.2 and details of datasets can be found in the appendix.

Results. Pretrained models are finetuned on each evaluation dataset. As shown in Table 2, our models UDOP achieve SOTA performance on all 7 tasks of DUE-Benchmark, ranking the 1st place on the leaderboard as of November 11, 2022. It also sets SOTA on CORD and (Table 3). It is worth noting that UDOP is an open-vocabulary generative model and uses one single model for all tasks. In comparison, most baselines leverage task-specific network for each dataset and are classification-based models. Nonetheless, UDOP still exhibits better results than those models.

Curriculum learning on image resolution (appendix Table 8) shows that with larger resolution, UDOP steadily gains stronger performance. E.g., UDOP average performance on DUE-Benchmark with 224, 512 and 1024 resolution is 63.9, 64.3 and 65.1 respectively. Note our model with 224 resolution already outperform previous best models (e.g., average 62.9 on DUE-Benchmark). We then train UDOP only with self-supervised objectives (224 resolution). Its performance (Table 4) also surpasses baselines, which shows the effectiveness of the unified representations, TVL transformer and the proposed self-supervised objectives.

Analysis

Masked Image Reconstruction. Figure 6 presents masked image reconstruction. Even with high masking ratio, the model can reconstruct the document image from text and layout signals with high quality: reconstructed contents are clear, consistent, and almost identical with the original image (all demonstrations are conducted on unseen documents.).

Document Generation & Editing. For the first time in Document AI, UDOP achieves controllable high-quality document generation and editing. As shown in Fig. 4), one can edit and add to the document image content with customized contents. The generated content is of high resolution and is consistent with the context in font, size, style and orientation (e.g., vertical numbers in Fig. 4). More generation examples are available in Appendix B. This is done by masking the regions to edit in the document image, and specifying the customized content in the text input, and their positions through layout embeddings. This novel functionality can generate augmentation document data for future research.

Layout Customization. UDOP can perform controllable high-quality document layout edits. We show examples in Figure 5, where our model can edit the layout of the document by regenerating the document from scratch. This is done by keeping only a few image patch as prompt, change the bounding boxes of the content, and then regenerate the document image with the new layout.

2 Ablation Analysis

Table 4 presents the ablation study of pretraining objectives on DocVQA and RVL-CDIP validation sets. We first develop a MLM (Masked Language Modeling) baseline that is a UDOP model pre-trained only on the BERT’s MLM that masks 15% of the input tokens. UDOP models (224 image resolution) pretrained with layout/text self-supervised objectives (“Layout Modeling”, “Visual Text Dataition”, and “Joint Text-Layout Reconstruction”) outperforms the one trained with masked language modeling (MLM), confirming their effectiveness. Table 4 also shows relative effectiveness of each pretraining task. Layout modeling improves upon Joint Text-Layout Modeling; Masked Image Reconstruction improves on text-based pretraining tasks. Adding vision self-supervised learning (masked image reconstruction) and supervised learning further improves the performance.

Modality-Specific Model Variant.

In the field of multimodal learning, a common model architecture is the two-tower model, where vision and text are encoded by two modality-specific encoders respectively . Therefore, we explore an variant of UDOP such that instead of having one unified encoder, we separately use a text encoder (to encode both text and layout tokens) and a vision encoder. Position bias are used in both encoders to represent layout information following previous works. We name this variant UDOP-Dual. For UDOP-Dual, the text-layout encoder-decoder follows T5-large, and the vision encoder-decoder has the same configuration as MAE-large. It has in total 1098M trainable parameters. As shown in Table 5 and Table 11, using one unified encoder is better than having separated encoders in most datasets. The exceptions are WTQ and RVL-CDIP on which UDOP-Dual achieves SOTA.

Additional Supervised Training Stage

TILT performs additional training on a wide range of QA datasets, such as reading comprehension dataset SQuAD , before the finetuning on DocVQA. This results in considerable performance improvement of the TILT model on DocVQA and InfographicsVQA. To have a fair comparison, we also finetune UDOP on the same set of datasets before testing on DocVQA or InfographicsVQA. As shown in Table 6, UDOP is further improved with this auxiliary training and outperforms TILT.

3 Effectiveness of the Vision Modality

In the field of Document AI, the effectiveness of the vision modality, i.e., document images, is unclear. We explore this by removing the visual embedding from the model input, with results shown in Table 7. It shows that the vision modality is more prominent on visually-rich tasks, e.g., InfographicsVQA, compared with text-dominant data such as DocVQA.

Conclusion

In this work, we propose UDOP, a foundation model for document AI. UDOP unifies the vision, text and layout modalities of documents by utilizing their strong spatial correlations through layout-induced vision-text representations and Vision-Text-Layout transformer. It also unites all self-supervised and supervised document tasks with a generative framework. UDOP achieves SOTA on 8 tasks and currently ranks the 1st place on the Document Understanding Benchmark Leaderboard. For the first time in document AI, UDOP achieves customizable realistic document generation and editing. We discuss the limitations and societal impact of our work in the appendix.

References

Appendix A Appendix Overview

Vision demonstrations of UDOP localizing answers in documents, the effectiveness of the cross attention with character embeddings in vision generation, and more neural editing examples Appendix B.

More details for pretraining and evaluation datasets, and finetuning experiment set up in Appendix D.

Experiment results of curriculum learning in Appendix E.

Performance variance of UDOP in Appendix F.

Discussion of limitations and societal impacts in Appendix G.

Appendix B Visualization Analysis

Creative Image Generation. UDOP achieves controllable high-quality document generation and editing as described in Section 6.1. We show additional examples here in Fig. 7. Our model can edit and add to the document image content with customized contents. Note that even if the document content is vertical (the first subfigure of Fig. 7), UDOP can still achieve high generation quality.

Answer Localization for Document QA. UDOP can perform question answering while predicting the location of the answer. We show examples on VisualMRC in Figure 8 and our model can answer the questions regarding the document correctly while locating the area of interest.

Appendix C UDOP-Dual Performance

We list the performance of UDOP-Dual on FUNSD, CORD, and RVL-CDIP in Table 11.

Appendix D Supervised Pretraining Tasks

In this section, we list more details about the supervised datasets in pretraining and evaluations.

RVL-CDIP contains 16 document categories, such as “invoice”, “scientific publication” and “form”. The dataset has 320k training, 40k validation and 40k test images.

D.2 Layout Analysis

PubLayNet is a layout analysis dataset created from medical publications. It contains over 360k document images and labeled with typical document layout elements such as titles, paragraphs, etc.

D.3 Information Extraction

DocBank is a richly-annotated large-scale IE dataset. It consists of 500K document pages, where 400K for training, 50K for validation and 50K for testing. It has 12 semantic structure labels like abstract, title, and author. Each token has corresponding bounding box and semantic structure label.

Kleister Charity is an IE dataset with complex invoice page layout and has 21.6k entities and 2.7k document images from UK Charity Commission. Its entities for extraction include invoice date, invoice number, net amount, vendor name, etc.

PWC is an IE dataset which has 2,291 leaderboards, where the data is collected from the Papers with Code labelling interface. It asks information like task, dataset, metric, etc. Different from original implementation, DUE-Benchmark provides complete papers as input instead of tables.

DeepForm is an IE dataset collected from political television ads in US elections and has 20k receipts and over 100k document images. This task is to extract entities like advertiser name, contract number, amount paid, etc.

D.4 Question Answering

WebSRC stands for Web-based Structural Reading Comprehension. It consists of 0.44M questions collected from 6.5K web pages with corresponding HTML, screenshots and metadata. The answer is either the text span of context or yes/no.

VisualMRC stands for visual machine reading comprehension. It consists of 10,197 images 30,562 abstractive questions-answers.

DocVQA is a QA dataset for excerpts from industry documents and has 50k questions on 12k document images. It asks questions on topics like text content, non-textual elements like marks or diagrams, layout, style, etc.

InfographicsVQA is a QA dataset with a focus on infographic images and has 30K questions on 5.3k document images. It requires reasoning on text content, images, data visualizations, layout, etc.

WTQ is a table-based QA dataset on HTML tables collected from Wikipedia. It has 2.1k tables and 22k questions hand crafted by humans and cover a wide range of topics like table lookup, superlatives, arithmetic operations, etc.

D.5 Document NLI

TabFact is an open-domain table-based NLI task and has 16k Wikipedia tables for 118k statements by human annotations.

D.6 Finetuning Experiment Setting

For all DUE-Benchmark finetuning experiments, we use Adam optimizer with learning rate 5e-5, 1000 warmup steps, batch size 16, weight decay of 1e-2, β1=0.9\beta_{1}=0.9, and β2=0.98\beta_{2}=0.98. For FUNSD and CORD, we use learning rate 3e-4 and for RVL-CDIP, we use learning rate 1e-3 both with 1000 warmup steps, batch size 16, weight decay of 1e-2, β1=0.9\beta_{1}=0.9, and β2=0.98\beta_{2}=0.98.

Appendix E Curriculum Learning

In this section, we present the results of curriculum learning of input image resolution (224, 512, 1024) on the validations sets of evaluation benchmarks. As shown in Table 8, while the model already performs competitively well on 224 resolution, its performance further increases on 512 and 1024.

Appendix F Performance Variance

For results in Table 2 and Table 3, we report their standard deviations as shown in Table 9 and Table 10. The deviations are computed from 5 runs with different seeds for parameter initialization.

Appendix G Limitations and Societal Impact

UDOP can assist users with document analysis, understanding and information extraction. This automatic processing technology will make the document processing workflow more efficient and potential more accurate. It is also worth noting that, similar to all AI generation technology, the document generation capacity of UDOP can be potentially abused for malicious document counterfeit, e.g., signature forgery, tampering monetary amount in checks, fake medical/financial records generation, etc. To avoid abuse, for model release we plan to open source the vision generation model only with limited access, e.g., through an API. Documents submitted by users that are classified as sensitive (the classifier can be a finetuned UDOP model), such as checks and personal ID, will be denied.

Applying UDOP on non-English data, especially those with non-Latin writing systems, may require further modifications to the model. For example, in Sec. 4.1, the vision decoder cross-attends with character embeddings. Then for non-English data, we need to include more character embeddings to attend with.