TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Yangfan He, Kuan Lu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, Wei Shi, Yuliang Liu, Hao Liu, Yuan Xie, Xiang Bai, Can Huang

Introduction

Recent research on multimodal large language models (MLLMs) has achieved significant advancements in the text-centric visual question-answering(VQA) domain , with several closed-source state-of-the-art (SOTA) models leading the way. Two representative examples are GPT4V and Gemini , which have demonstrated remarkable performance and have even surpassed human-level capabilities in certain aspects. Nevertheless, as illustrated in Figure 1, the performance of open-source models still lags significantly behind that of pioneering closed-source models. This phenomenon can be attributed to various factors, including model architecture, the scale of model parameters, image resolution, the volume of pretraining and instruction tuning data, and training strategies, among others.

Many pioneering studies have recently conducted data-centric research into the challenges of insufficient instruction tuning data. For instance, Monkey initially employed expert models to generate descriptions of different aspects of images, which were then summarized by GPT-4 to produce high-quality and detailed image caption data. For better text-based knowledge injection, For better text-based knowledge injection, LLaVAR and TG-Doc used GPT-4 to generate conversations for text-rich images by integrating OCR results into the instructions. In order to improve the image caption ability for MLLMs, ShareGPT4V constructs a high-quality image caption dataset through GPT4V. While these efforts have achieved remarkable success, they also left some challenges unresolved. Image caption data and VQA data belong to different domains, with inconsistencies in the granularity and scope of image content presentation. Furthermore, the scale of synthetic data remains relatively small, preventing MLLMs from fully realizing their potential. The exploration of methods that leverage large-scale text-centric VQA data for instruction tuning of existing open-source models remains limited.

To bridge the gap, this paper proposes a strategy termed Square for obtaining massive, high-quality text-centric VQA data from sophisticated and versatile closed-source MLLMs, resulting in the construction of a dataset (Square-10M) comprising tens of millions of instances for instruction tuning. Specifically, the method consists of four steps: Self-Questioning, Answering, Reasoning, and Evaluation. The self-questioning step involves utilizing the MLLM’s capabilities in text-image analysis and understanding to generate questions related to the textual content of images. The answering step involves answering these generated questions, leveraging various prompting techniques such as Chain-of-Thought and few-shot prompting. The reasoning step entails probing the model for the reasoning behind its answers, leveraging the powerful reasoning abilities of MLLMs. The evaluation step involves evaluating the question-answer pairs, assessing the validity of the questions and their relevance to the textual content of the images, as well as the correctness of the answers, thereby improving data quality and mitigating hallucinations. Overall, Square comprehensively leverages the capabilities of MLLMs in various aspects, significantly enhancing the data quality.

Besides, enriching the diversity of images is also crucial. We collect a diverse set of text-rich images from various public sources, including natural scenes, charts, tables, receipts, books, slides, PDFs, documents, products, and web images. Subsequently, deduplication is performed on this collection. By applying the Square method to these images, Square-10M is constructed.

Based on Square-10M, we achieve several remarkable results with extensive and rigorous experiments. First, as shown in Figure 1, our model (TextSquare) achieves comparable or superior performance to advanced closed-source models and substantially outperforms recent state-of-the-art open-source models on various benchmarks. It is notable that the image resolution of TextSquare is 700700 and the parameters are 8.68.6B. Second, our experiments validate the beneficial impact of reasoning data on VQA tasks, demonstrating its ability to enhance model performance while mitigating hallucinations. With reasoning data for instruction tuning, TextSquare has a strong reasoning capability to provide elaborate explanations for VQA scenarios. Last but not least, by leveraging the dataset’s massive scale, we unveil the relationships between instruction tuning data scale, training convergence loss, and model performance. Whereas a few instruction tuning data can motivate MLLM well, it is not sufficient. Large amounts of high-quality data can still significantly reduce convergence loss and improve performance. The performance of TextSquare grows and the loss of convergence decreases while continuously scaling up the instruction tuning data, which also demonstrates the effectiveness of our dataset.

In summary, the main contributions of this paper can be categorized into four points:

A high-quality dataset (Square-10M) comprising tens of millions of instances for text-centric VQA instruction tuning is constructed by comprehensively collecting text-rich images from various scenarios and employing the Square (Self-Questioning, Answering, Reasoning, and Evaluation) strategy on closed-source MLLMs.

Leveraging Square-10M, TextSquare achieves a significant outperformance of existing open-source models and even comparable or superior performance to SOTA closed-source models on various benchmarks, e.g., +0.9% on ChartQA, +2.1% on WTQ, +4.3% on SROIE. Notably, TextSquare outperforms GPT4V in overall rankings across ten text-centric benchmarks (ranking 2.2 v.s. 2.4).

Reasoning data is demonstrated to be beneficial in improving model performance and mitigating hallucinations in text-centric VQA scenarios, as it can deliver rich question-specific contextual information.

Through extensive experiments, we reveal the relationships between data scale, convergence loss, and model performance for text-centric VQA instruction tuning, which demonstrates the effectiveness and necessity of Square-10M.

Related Work

Recent work has increasingly focused on introducing visual knowledge into LLMs . General attempts connect a visual encoder and an LLM with intermediate modules like Projector , Q-Former , Perceiver Resampler , etc, and go through pre-training alignment and instruction fine-tuning for vision-language understanding.

Recently, several researches propose to enhance MLLMs’ capabilities in understanding textual elements (OCR, text-centric VQA, etc). Among them, mPLUG-DocOwl creates novel instruction-following datasets to enhance the tuning process. TextMonkey adopts shifted window attention and filters out significant tokens. DocPedia and HRVDA enlarges input resolution to bridge the gap between MLLMs and visual document understanding.

Despite the extraordinary progress of existing open-source MLLMs, they still suffer from the huge gap against SOTA closed-source models like GPT4V and Gemini Pro . In this paper, we propose to mitigate this gap by training with large-scale and high-quality instruction-following data.

2 Text-Centric Visual Question Answering

Text-Centric Visual Question Answering aims to understand the interactions between the textual and the visual elements in the image. Donut first proposes an end-to-end training method based on a Transformer without OCR. Pix2Struct introduces a variable-resolution input representation to adapt to document images. DoCo enhances the visual representation of the image encoder in LVLMs by aligning the document object of multi-modal inputs. BLIVA enlarges the input token space by concatenating learned query embeddings and encoded patch embeddings. Several studies have performed data-centric attempts in this regard. UniDoc construct 600k document-oriented image-text pairs from PowerPoint presentations. LLaVAR and TG-Doc prompt text-only GPT-4 to generate conversations for text-rich images by integrating OCR results into the instructions. These researches are restricted to small-scale annotations or generation based on uni-modal inputs.

3 Generating Instruction-Tuning Data via LLMs

The success of LLMs has inspired recent work to employ them as training data generators . In this regard, we anchor on generating instruction-following data. Self-Instruct took the initial step towards synthesizing instructions via language models and improving the instruction-following capabilities. Llama-GPT4 uses GPT-4 to generate instruction-following data for LLM fine-tuning. Synthetic Prompting leverages a few handcrafted examples to prompt LLMs to generate more examples. Bonito converts unannotated text into task-specific training datasets for instruction tuning. Recently, ALLAVA employs GPT4V to generate reasoning instructions and detailed answers from unlabeled images. All of the above attempts suffer from the low quality of the generated data and are typically performed on a small scale. In contrast, we collect massive text-centric images (i.e., tens of millions) and devise comprehensive generating methods and filtering rules to ensure the quantity and quality of the instruction tuning dataset.

Square-10M: A Massive and High-quality Text-Centric VQA Instruction Tuning Dataset

Square-10M is synthesized by our proposed Square pipeline, i.e., Self-Questioning, Answering, Reasoning, and Evaluation.

Figure 3 presents an overview of our proposed Square. Square generally consists of three stages for synthesizing high-quality instruction tuning data for text-centric VQA: (1) Data Collection for collecting large-scale images with textual elements of diverse properties. (2) Data Generation involves self-questioning, answering, and reasoning of the collected data. In this phase, the MLLM is prompted to generate VQA pairs based on the given image, as well as the reasoning behind its answers. (3) Data Filtering for self-evaluation of the generated content, aiming to discard meaningless questions and erroneous answers by employing the evaluation capabilities of MLLMs.

The above procedures result in our Square-10M dataset, standing out with its massive and high-quality text-centric VQA pairs and reasoning context. To be more specific, a total of 3.8 million images with rich textual elements are collected from diverse sources. After that 20 million question-answer pairs are obtained from Data Generation. Finally, 9.1 million QA pairs as well as the reasoning context are distilled with our Square strategy. A more precise analysis of Square-10M is depicted in Figure 2.

2 Data Collection

The data collection strategy is driven by the primary objective of encompassing a broad range of real-world text-rich scenarios. To this end, we collect 3.8 million unlabeled text-rich images (Figure 2). These images exhibit diverse properties. For instance, Chart and Table focus on textual elements with intense statistical information; Slide, Screenshot, and WebImage are designed for the interaction between text and prominent visual messages; Document/PDF, Receipt, and e-commerce contain images with fine and dense text; Street-View is derived from natural scenes. The collected images form a mapping of the textual elements in the real world and constitute the foundation of our research on text-centric VQA.

3 Data Generation: Self-Questioning, Answering, and Reasoning

We build our Square-10M dataset by employing the multi-modal understanding capabilities of Gemini Pro, one of the most advanced LLMs. For each image selected from a specific data source, Gemini Pro is instructed to generate VQA pairs and reasoning context through the subsequent three stages:

Stage 1: Self-Questioning. In this stage, Gemini Pro is prompted to generate profound, meaningful, and non-trivial questions about the given image. We ask Gemini Pro to first comprehensively analyze the image and then raise questions based on its understanding, as shown in Figure 3. Considering that advanced MLLMs typically have weaker understanding capabilities of the textual elements than visual elements, we also prepend the extracted text to the prompt by employing expert OCR models.

Stage 2: Answering. Gemini Pro is then instructed to give appropriate answers to the generated questions. We leverage various prompting techniques to enrich the contextual information and improve the reliability of the generated answers, such as Chain-of-Thought and few-shot prompting. Figure 3 shows an example prompt for generating answers to a given question.

Stage 3: Reasoning. We require Gemini Pro to elaborate on the detailed reasons behind its answers. Such an effort enforces Gemini Pro to think more about the connections between the questions and the visual elements, thus reducing hallucinations and providing accurate answers. Moreover, the generated reasons could serve as extra contextual information specific to individual questions, favoring possible research on the mechanism behind in-context learning. We present an example prompt for self-reasoning in Figure 3.

4 Data Filtering: Self-Evaluation and Answering Consistency

Despite the effectiveness of Self-Questioning, Answering, and Reasoning, the generated image-text pairs could face hallucinatory content, meaningless questions, and erroneous answers. We thus devise filtering rules based on the Evaluation capabilities of LLMs to select high-quality VQA pairs. The whole filtering system is established upon three aspects.

Self-Evaluation of MLLMs. We prompt Gemini Pro as well as other advanced MLLMs to judge whether the generated questions are meaningful and whether the answers are good enough to correctly address the questions.

Figure 3 depicts an example prompt for self-evaluation.

Multi-Prompt Consistency. Besides direct evaluation of the generated content, we manually augment the prompt and context space in Data Generation. A correct and meaningful VQA pair should be semantically consistent when provided with different prompts. Specifically, in the stage of Answering we provide Gemini Pro with different but semantically similar prompts to answer the given question. Then we discard the VQA pairs if the generated answers are not stable in semantics. An example is given in Figure 3.

Multi-Context Consistency. Similar to Multi-Prompt Consistency, we further validate the VQA pairs by prepending the question with varied context information. Given the generated question, three types of answers are produced by Gemini Pro with different contexts: (1) Answering with reasoning. Gemini Pro answers the question with a detailed explanation prepended (i.e., content generated in the stage of Reasoning). (2) In-Context answering. Gemini Pro answers the question with chain-of-thought or few-shot prompts prepended. (3) Naive answering. Gemini Pro answers the question with no extra context. We then discard the VQA pairs if the generated answers are not semantically consistent.

TextSquare: A Text-Centric Multimodal Large Language Model

The model architecture of TextSquare follows the paradigm established by InternLM-Xcomposer2 , including three integral components: (1) A Vision Encoder modified from OpenAI CLIP ViT-L-14-336 , where the resolution is increased to 700 for improved performance. (2) A LLM based on InternLM-2 , utilizing InternLM2-7B-ChatSFT as the practical variant. (3) A Projector, which semantically aligns the vision token and the text token.

2 Supervised Fine-Tuning with Square-10M

TextSquare is achieved by performing Supervised Fine-Tuning (SFT) with Square-10M. The SFT process comprises three stages: In the first stage, we unfreeze all the three components (i.e., the Vision Encoder, the LLM, and the Projector) and train the model in a resolution of 490. In the second stage, the input resolution is increased to 700 and only the Vision Encoder is trained to adapt to the resolution change. In the third stage, we further perform full-parameter fine-tuning in the resolution of 700. TextSquare demonstrates that with our Square-10M dataset, a model with 8B parameters and normal-size image resolution can achieve extraordinary performance on text-centric VQA, surpassing most available MLLMs and even the closed-source SOTA models.

Experiment

The training data contains Square-10M and in-domain datasets (consistent with Monkey’s SFT data). The training process is divided into three phases, using the same data and the AdamW optimizer with 64 A100-80G GPUs. In the first phase, we fine-tune InternLM-Xcomposer2 with full parameters, and the learning rate decreases from 1e-5 to 1e-6, taking about 9520 GPU hours; In the second phase we scale up the image resolution to 700, and train only VIT, with the learning rate decreasing from 1e-4 to 1e-5, taking about 7280 GPU hours; In the third stage, we perform full-parameter fine-tuning at 700 image resolution, and the learning rate drops from 1e-5 to 1e-6, spending about 12350 GPU hours.

2 Benchmark Evaluation

We report the results on Scene Text-centric VQA, Document-oriented VQA, Table VQA, Text-centric KIE, OCRBench, and General VQA for a comprehensive comparison of the performance of our model with existing models. The metrics of each benchmark are listed in Table 6 in the Supplementary Material.

Document-Oriented Benchmark. While the documents have a clean background, dense text and complex typography pose distinct challenges. To effectively evaluate our model, we select representative benchmarks including DocVQA , ChartQA , and InfographicVQA . The results, detailed in Table 1, show that TextSquare outperforms all the open-source models in these three document-oriented VQA tasks with an average improvement of 3.53.5%, specifically, DocVQA 84.384.3% vs. 81.681.6% (Cogagent and mPLUG-DocOwl 1.5), ChartQA 79.479.4% vs. 72.772.7% (Intern-Xcomposer2), InfographicVQA 51.551.5% vs. 50.450.4% (mPLUG-DocOwl 1.5). On the ChartQA dataset, TextSquare outperforms GPT4V and Gemini Pro by a slight margin. Note that TextSquare employs an image resolution of 700, which is smaller than most document-oriented MLLMs. Our model relies on comprehensively high-quality VQA information specific to the text in the document, improving its ability to recognize and understand various document elements such as text, diagrams, infographics, and so on. If the image resolution is further increased, it is believed that the model performance will be further improved, as demonstrated by Monkey et al.

Scene Text-centric Benchmark. The ability to answer text-based questions in images becomes an important aspect of the answering task as textual information is usually present in real-world scenes. In the evaluation, we utilize two datasets: TextVQA and AI2D . As shown in Table 1, in this scenario, although TextSquare achieves SOTA performance on the AI2D dataset, there is no major improvement over our baseline Intern-Xcomposer2, which may be due to the fact that Intern-Xcomposer2 has been adequately optimized with high-quality in-domain data.

Table VQA Benchmark. Due to the complex structure of tables and the dense text, the understanding of the content of tables remains a challenging issue. In order to evaluate the performance of the comprehension of table content and structure, we choose two widely utilized datasets, Wiki Table Questions (WTQ) and Table Fact (TabFact) , as shown in Table 1. On the Table VQA benchmarks, TextSquare achieves optimal performance among the leading models with an average 3.03.0% improvement. This demonstrates that our model has reached a new level of table understanding, where high-quality generated table VQA and reasoning data play a key role.

Text-centric KIE Benchmark. Text-centric key information extraction tasks are frequently encountered in the information processing of various types of products, certificates, and receipts. We select a receipt information extraction dataset (SROIE) and a product information extraction dataset (POIE) , and the KIE task is converted to the VQA task. TextSquare achieves optimal performance in both datasets, with a major average lift of 14.814.8% (shown in Table 1). It is worth noting that there is no training set of POIE added to the training set and there is not much data in the domain of product scenarios. This illustrates the extensive textual comprehension capabilities of our model.

OCRBench. OCRBench is a comprehensive benchmark consisting of 29 OCR-related assessments, with text recognition, formula recognition, text-centric VQA, KIE, etc. TextSquare achieves optimal performance in OCRBench except for the closed-source models and becomes the first MLLM that exceeds 600600 points with about 1010B parameters. It indicates that the model performs well in both text-centric perception and comprehension tasks, especially in text recognition, where little in-domain data is included in the training set.

General VQA and Hallucination Evaluation Benchmark. General VQA requires the ability to learn both visual and textual information and a deep understanding of their inter-relationships. For general VQA, we validate on four benchmarks: VizWiz , VQAv2 , GQA , and POPE . The VizWiz and POPE benchmarks are also relevant for hallucination evaluation. The results are shown in Table 2. On VQAv2 and GQA, TextSquare does not have a significant degradation compared to InternLM-Xcomposer2 and still maintains comparable performance. TextSquare exhibits superior capabilities in VizWiz and POPE, outperforming the closest competing method by an average of 3.63.6%. These results highlight the effectiveness of our approach, which is also able to mitigate model hallucinations in particular with large-scale instruction tuning. We observe that it is partly attributed to the high-quality reasoning data that provides detailed explanations for VQA.

3 Qualitative Analysis

As illustrated in Figure 4, TextSquare has a formidable capability to provide plausible explanations of the answers to questions in a variety of text-centric VQA scenarios. Figure 4(a) shows that TextSquare has simple arithmetic capabilities. Figure 4(b) shows the ability to understand textual content and provide approximate location in dense text. Figure 4(c) shows the comprehension of table structure and the ability to extract contextual information relevant to the question.

4 Ablation Study

The Effect of Incorporating Square-10M for Instruction Tuning.

In order to verify the effectiveness of Square-10M, we fine-tune the baseline model InternLM-Xcomposer2 on the public text-centric VQA instruction tuning dataset (consistent with Monkey’s training data). As shown in Table, TextSquare substantially outperforms Xcomposer2∗ (fine-tuned) on various text-centric VQA benchmarks by 7.77.7%, which corroborates that Square-10M can fully exploit MLLM’s ability in text-centric VQA scenarios and that a large amount of high-quality instruction tuning data has a major improvement in performance.

The Effect of Evaluation Step of the Square Strategy. As shown in Table 5, there is a distinct improvement in model performance after incorporating the evaluation of the generated VQA data, which verifies that the evaluation step of the Square strategy improves the quality of VQA instruction tuning data.

The Effect of VQA Reasoning Data on Model Performance and Hallucination Evaluation. From Table 5, we can find that VQA Reasoning data is helpful in both improving VQA performance and mitigating hallucinations. Specifically, in terms of enhancing VQA performance, there is a 1.4% and 1.3% gain on DocVQA and ChartQA. In terms of mitigating hallucinations, there is a 2.72.7% and 3.23.2% gain on POPE and WizViz.

5 Relationships between Instruction Tuning Data Scale, Convergence Loss, and Model Performance

To explore the relationship between instruction tuning data scale, convergence loss, and model performance based on the merged large-scale Square-10M and the in-domain instruction tuning dataset, we conduct 10 sets of experiments for different data scales. The average performance of the models is evaluated on DocVQA, ChartQA, InfoVQA, WTQ, and SROIE. As shown in Figure 5(a)(b), the convergence loss of the model continues to decrease as the data scale grows, whereas the rate of decrease becomes progressively slower. The relationship between the convergence loss and the instruction tuning data scale approximately conforms to a logarithmic function. Similarly, from Figure 5(c)(d), it can be seen that as the instruction tuning data grows, the model performs better and better, but the rate of growth continues to slow down. Their relationship is also approximately in accordance with a logarithmic function. Holistically, there is a corresponding scaling law in the instruction tuning phase in text-centric VQA scenarios, where model performance is proportional to the logarithm of the scale of data. It can guide the construction of potentially larger datasets and predict model performance.

Limitation

Although our approach achieves remarkable results in various scenarios, there are some limitations. Firstly, large-scale data requires plenty of GPUs for long-time training, which greatly increases the training consumption. Second, while the Square strategy improves the quality of synthetic data, it still cannot reach the human level.

Conclusion

In this paper, we present the Square strategy for constructing a high-quality text-centric instruction tuning dataset(Square-10M). Leveraging this dataset, TextSquare significantly surpasses recent open-source models and even achieves performance comparable to GPT4V across various benchmarks. Furthermore, we derive the relationship between instruction tuning dataset scale, convergence loss, and model performance in order to pave the way for constructing even much larger datasets. Our approach provides a data-centric perspective that revisits the role of instruction-tuning data in text-centric VQA, confirming that both the quantity and quality of data are crucial to model performance. We believe that there is a promising direction on how to further improve the data quantity and quality for closing the gap between open-source models and the leading ones.

References

Supplementary Material

We summarize the evaluation benchmarks used in this paper in Table 6.