MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Jingqun Tang, Qi Liu, Yongjie Ye, Jinghui Lu, Shu Wei, Chunhui Lin, Wanqing Li, Mohamad Fitri Faiz Bin Mahmood, Hao Feng, Zhen Zhao, Yangfan He, Kuan Lu, Yanjie Wang, Yuliang Liu, Hao Liu, Xiang Bai, Can Huang
Introduction
In the era of burgeoning AI, especially in LLMs/MLLMs , Text-Centric Visual Question Answering (TEC-VQA) has served as a de facto gold proxy to evaluate AI models in the domain of text-centric scene understanding. Compared with general VQA , TEC-VQA places greater emphasis on answering questions that require understanding textual information within images. It provides a streamlined avenue for individuals without specialized expertise to articulate their requirements and access applications in text-centric visual environments. However, the majority of advancements in TEC-VQA have predominantly concentrated on high-resource languages, e.g., English , Chinese , Japanese and etc., thus restricting the applicability of AI models to the global community, particularly populations speaking low-resource languages.
To tackle the problem of language diversity, several seminal studies in the general VQA field, leverage off-the-shelf translation engines to expand existing question-answer pairs from high-resource languages to their multilingual counterparts including low-resource ones. However, when applied to TEC-VQA, this translation-based approach may fall prey to the “Visual-textual misalignment” problem as only the text in question-answer pairs can be processed, while the visual text present in the images is overlooked. Not to mention issues such as nuanced meaning, contextual distortion, language bias, and question type diversity further render the transferability of the translation protocol infeasible for TEC-VQA. The status quo begs for a question: “How can we address the visual-textual misalignment problem for multilingual TEC-VQA and what we stand in the MLLM era?”
In this work, to answer the question above, we establish MTVQA, a novel and high-quality multilingual TEC-VQA benchmark, where all images are collected from real-world and meticulously annotated by human experts in nine languages: Arabic (AR), Korean (KO), Japanese (JA), Thai (TH), Vietnamese (VI), Russian (RU), French (FR), German (DE), and Italian (IT). More concretely, to ensure the visual-textual alignment at most, the annotation process follows the raise-then-correct paradigm, where a group of human annotators raises several distinct questions, ranging from simple content extraction to text-related reasoning, and subsequently provides answers. These QA pairs are then double-checked by another group to ensure accuracy and consistency. Consequently, as illustrated in Fig. 5, 6,678 training images and 21,829 question-answer pairs, as well as 2,116 test images and 6,778 question-answer pairs are obtained, covering several fine-grained scenarios, such as menus, logos, maps, bills, PPTs, research papers, and etc. To our best knowledge, MTVQA is the first TEC-VQA dataset to provide native human annotations for multilingual text-rich scenarios, especially for low-source languages. Furthermore, we investigate recent representative MLLMs, including GPT-4V, Gemini, QwenVL etc., by juxtaposing experimental results regarding their performance on our newly proposed MTVQA. Both for general MLLMs and document-focused ones, the results unequivocally demonstrate that opportunities for improvement persist within these MLLMs when applied in multilingual text-rich scenarios.
In summary, the main contributions of this paper can be categorized into three points:
We introduce the MTVQA dataset, to the best of our knowledge, which is the first multilingual TEC-VQA benchmark to provide human expert annotations for text-centric scenarios.
We benchmark the state-of-the-art MLLMs on our new dataset and show there is still room for performance improvement for these models under multilingual text-rich scenarios.
We propose a set of baselines for multilingual TEC-VQA tasks.
Related Work
Recent advancements in LLMs/MLLMs have revolutionized VQA tasks, as demonstrated by the remarkable zero-shot performance of these models. Notably, the high generalizability of LLMs/MLLMs, when explicitly trained on visual text understanding datasets and fine-tuned with instructions, has significantly enhanced their application in text-centric VQA scenarios . For example, LLaVAR , UniDoc , which extend LLaVA into the realm of document understanding, pioneering the text-centric VQA of MLLMs by training them to predict texts and coordinates from document images. Furthermore, DocPedia operates visual input in the frequency domain rather than in space, which enables higher input resolution without increasing the input sequence. Lately, mPLUG-DocOwl , Qwen-VL , and TextMonkey leverage publicly available document-related VQA datasets to further enhance the text-centric VQA capabilities. Despite the promising results achieved by existing LLMs/MLLMs in text-centric VQA tasks, their focus on high-resource languages such as English or Chinese has posed challenges in achieving reasonable performance for low-resource languages. This is primarily due to the lack of data or benchmarks for these low-resource languages.
2 Multilingual text-centric VQA Benchmarks
VQA has garnered significant attention in recent years, with numerous studies, datasets, and benchmarks being proposed to advance the field . Many datasets have been created that encompass scene text of various domains, including natural images , scanned documents , book and movie covers . One notable limitation of these datasets is their predominant focus on English or other high-resource languages such as Chinese and Japanese , which restricts the applicability of VQA systems for low-resource languages such as Thai and Vietnamese.
There is a recent effort toward extending VQA tasks to a wider range of languages by providing a multilingual VQA datasets. For example, Gao et al. created a free-form bilingual VQA dataset (FM-IQA) contains over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations. Raj Khan et al. developed a large-scale multilingual and code-mixed VQA dataset (MuCo-VQA) supporting five languages. Of more relevance are the works xGQA (8 languages) and MaXM (7 languages) , which apply translation-based protocols to expand VQA data beyond English. However, the translation-based multilingual VQA datasets inherently face issues, such as the “Visual-textual misalignment” problem, where only the text in question-answer pairs is processed, while the visual text in images is overlooked. Additionally, the nuanced meaning and context are often distorted; language bias introduced by machine translation models, and the coverage of certain question types is limited, as highlighted by Changpinyo et al. . Moreover, none of the previous multilingual datasets focus on text-centric scenarios where multilingual text frequently occurs.
Our benchmark distinguishes itself by focusing on multilingual text-centric VQA scenarios using human expert annotations. To the best of our knowledge, the MTVQA benchmark is the first dataset to provide native human annotations for such scenarios. It covers 9 languages, thereby facilitating the training and evaluation of multilingual models in diverse linguistic contexts. Additionally, our dataset can gauge the VQA system’s ability for not only high-resource languages but also those that are typically underrepresented in current datasets .
The MTVQA benchmark addresses a significant gap in existing datasets by catering to the crucial needs of low-resource languages through annotations from native speakers across multiple languages. Our pioneering efforts distinctly position the MTVQA benchmark as a unique multilingual VQA resource, advancing the frontier of machine learning research.
MTVQA Benchmark
The MTVQA Benchmark covers 9 languages: Arabic (AR), Korean (KO), Japanese (JA), Thai (TH), Vietnamese (VI), Russian (RU), French (FR), German (DE), and Italian (IT). In this section, we describe in detail how we establish the MTVQA benchmark, including the collection of raw image data and two-round human expert annotations, which are independent of each other.
Our purpose is to develop a multilingual VQA benchmark capable of evaluating the QA performance of MLLMs in multilingual text-centric scenarios, thus the raw data collection process is mainly oriented towards text-centric images from natural scenarios and document scenarios. To ensure the diversity and quality of data, we collect not only the raw image data from publicly available datasets, including the multilingual scene text recognition images from MLT2019 and PowerPoint slides (PPTs) sourced from the internet, but also the data from countries of each language. Furthermore, the collected data includes multiple fine-grained scenarios (Fig. 1), such as menus, logos, maps, bills, PPTs, research papers, and etc. As a result, we gather a total of 1,220 images from document scenarios and 876 images from natural scenarios in the test set of the MTVQA benchmark. To ensure the visual-textual alignment, for text-rich images lacking text and language annotations, we subject them to a standardized data cleaning process, which includes text recognition and language classification. Afterward, we organize all the text-rich images we have obtained into language-specific groups, preparing them for the subsequent stage of data annotation.
2 Human Expert Annotation
In order to obtain informative and accurate text-related QA pairs on the language-specific grouped images, we recruit a group of annotators with expertise from local regions of each language. It is worth noting that all these annotators are native speakers of their respective languages, ensuring their deep understanding and proficiency in the linguistic nuances and cultural context necessary for precise annotations. Considering the subjective nature of the text-image understanding task, we have implemented a further division within the annotation team. This division involves separating the team into two independent groups, with one group dedicated to generating and responding to questions based on the provided images, while the other group focuses on evaluating and correcting the QA pair results. This raise-then-correct paradigm ensures a comprehensive and reliable assessment of the text-image understanding process. Additionally, each language’s annotation results undergo a 10% sampling inspection by a quality inspector. If the QA pairs fail to meet the criteria, they are sent back for re-annotation. Prior to commencing the formal human expert annotation task, all annotators undergo unified training and receive annotation examples. The brief diagram of the two-round annotation process is shown in Figure 3 and we elaborate on it in the following subsections.
First Round Questioning and Answering. For the first round of annotation tasks, we assigned 3 annotators for each language to manually generate original QA results. Given a text-centric image from our collection, annotators are first required to read the texts in the image and analyze other contents in the image in a comprehensive and detailed manner. They must then raise 4 meaningful and distinct questions based on the content in the image and give the answers. All annotators adhere to the following criteria: (1) the first three questions should satisfy that answering these questions requires direct reading of the textual information in the image, (2) the fourth question requires reasoning about the text in the image to answer (3) the questions and answers must be reasonably correct and consistent with the content of the image, and (4) the answer should be as concise as possible and free of nonsense (e.g., when the question is “When is the volunteer recruitment period”, the answer should be “9:00-16:00” rather than “The volunteer recruitment period is 9:00-16:00”). It’s worth mentioning that our requirement for concise answers is to make the evaluation process more friendly and more reliable, cause we try to keep the evaluation metrics unaffected by extraneous content in the answer sentence.
Second round Evaluation and Correction. To reduce the effect of human subjective cognitive bias on our MTVQA benchmark and get high-quality question-answer pairs, we assigned 2 annotators for each language for the annotation evaluation and correction process. Based on the provided images and the first-round annotation results, the annotators must follow these rules of judgment and steps for the annotation: (1) Whether the question is related to the text in the image. If not, discard the current question-answer pair, (2) Whether the answer is correct. If not, modify the answer, and (3) Whether the answer repeats the content from the question. If so, remove the repeated content to ensure a concise answer.
3 Data Statistics
We instruct the annotators to complete the above human expert annotation work towards the text-centric VQA tasks and construct the MTVQA benchmark consisting of 8,794 images and 28,607 question-answer pairs that cover the 9 languages. The MTVQA benchmark is divided into a training set containing 6,678 images and 21,829 question-answer pairs, and a test set containing 2,116 images and 6,778 question-answer pairs. The detailed data distribution can be seen in Figure 1. To visualize the vocabulary richness of our benchmark, we calculate the word frequencies for each language and present them in the form of word clouds as shown in Figure 4. In Figure 5 we demonstrate the statistics of the question and answer lengths using GPT-4o tokenizer.
Experiments
For the MTVQA benchmark, we evaluate the following instruction-tuned general MLLMs, (1) Open-source MLLMs: InternVL-V1.5 , InternLM-Xcomposer2-4KHD , Mini-Gemini-HD-34B , Llava-Next-34B , DeepSeek-VL , YI-VL-34B , TextSquare , TextMonkey and mPLUG-DocOwl 1.5 ; (2) Closed-source MLLMs: GPT-4V, Gemini Ultra, QwenVL Max, QwenVL Plus, Claude3 Opus, Claude3 Sonnet and GLM4V. For the closed-source MLLMs, we use the chat version through the official APIs, while for the open-source MLLMs, we utilize the instruct versions. It is noted that all the model weights of the open-source MLLMs evaluated in our experiments could be downloaded from the HuggingFace Model Hub. For the open-source MLLMs, the model size varies from 7b to 34b.
2 Implementation Details
We conduct the evaluation experiments over the baseline MLLMs with their default settings, ignoring the effect of generation configuration on the results. To make the output of MLLMs more evaluation-friendly, we design the following prompt format to limit the output length: “Answer the question using a word or phrase in the language of the question. +
3 Evaluation Results
Zero-shot testing To demonstrate the quantitative comparison results in the above MLLMs, we follow TextMonkey with accuracy as the evaluation metric. That is, the model output is only counted as correct if it contains the ground truth. The complete evaluation results are shown in Table 2, where Claude3 Opus achieves the highest average accuracy of 25.7 on the 9 languages. It indicates that the multilingual text-centric VQA tasks remain a big challenge, even for the state-of-the-art open-source and closed-source MLLMs. From the metrics across languages, both open-source and closed-source models performed significantly better on Indo-European languages using the Latin alphabet, including DE, FR, and IT in our benchmark, compared to other languages, which results from the distribution of realistically available training data and the genetic relationship of different languages. In addition, all closed-source models except GLM4V outperform the open-source model overall across the nine languages, which may be due to the contribution of pre-training on multilingual data. We also found that the document-focused MLLMs, like TextSquare and TextMonkey , do not significantly outperform other open-source models on the metrics of these 9 languages.
Instruction tuning As shown in Table 2, the instruction tuning experiment on MTVQA benchmark brings a 8.5 improvement in average accuracy. With respect to specific languages, French sees the largest improvement of 14.2 in accuracy, while Russian has the smallest improvement of 1.7 in accuracy. The results demonstrate that MLLMs vary in their ability to understand and learn from text-centric data in different languages, leaving great potential for future research of multilingual text-centric MLLMs pre-training.
Limitation
The current iteration of MTVQA exhibits certain constraints that warrant attention. Primarily, the linguistic diversity incorporated is not exhaustive; several lesser-spoken languages remain unrepresented. Future enhancements will aim to broaden the multilingual scope of the dataset. Additionally, the dataset currently offers a singular canonical response for each question. Recognizing the multifaceted nature of the inquiry, subsequent versions will endeavor to include a spectrum of plausible answers to reflect the varied perspectives inherent to each question.
Conclusion
In this paper, we introduce MTVQA, a multilingual TEC-VQA benchmark featuring high-quality human expert annotations in 9 diverse languages. We believe that MTVQA is the first benchmark of its kind to provide fully manual annotations specifically tailored to text-centric scenarios. The results obtained from both closed- and open-source MLLMs on our MTVQA dataset indicate that there is still room for improving their performance in multilingual text-centric scenarios. Although the current version of MTVQA has constraints regarding linguistic diversity and singular responses per question, we are confident that this dataset can still inspire researchers within the TEC-VQA community with new perspectives and ideas.