TAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation
Jun Wang, Mingfei Gao, Yuqian Hu, Ramprasaath R. Selvaraju, Chetan Ramaiah, Ran Xu, Joseph F. JaJa, Larry S. Davis
Introduction
Visual question answering (VQA) task [Antol et al.(2015)Antol, Agrawal, Lu, Mitchell, Batra, Zitnick, and Parikh] aims at inferring the answer to a question based on a holistic understanding of an image. It facilitates many AI applications such as robot interactions [Anderson et al.(2018b)Anderson, Wu, Teney, Bruce, Johnson, Sünderhauf, Reid, Gould, and Van Den Hengel], document analysis [Mishra et al.(2019)Mishra, Shekhar, Singh, and Chakraborty] and assistance for visually impaired people [Bigham et al.(2010)Bigham, Jayant, Ji, Little, Miller, Miller, Miller, Tatarowicz, White, White, et al.]. Text-VQA specifically addresses question answering requests where reasoning text in an image is essential to answer a question. It is a more challenging task in a sense that it requires not only understanding the question and the visual context, but also the embedded text in an image [Singh et al.(2019)Singh, Natarajan, Shah, Jiang, Chen, Batra, Parikh, and Rohrbach]. To achieve this goal, Text-VQA methods [Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach, Kant et al.(2020)Kant, Batra, Anderson, Schwing, Parikh, Lu, and Agrawal, Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] aim at studying the interactions among question words, visual objects, and scene text in an image. Recent approaches have focused on either improving transformer-based architectures [Vaswani et al.(2017)Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, and Polosukhin] in a multi-modal manner [Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach, Kant et al.(2020)Kant, Batra, Anderson, Schwing, Parikh, Lu, and Agrawal, Gao et al.(2021)Gao, Zhu, Wang, Li, Liu, Van den Hengel, and Wu, Liu et al.(2020)Liu, Xu, Wu, Du, Jia, and Tan, Han et al.(2020)Han, Huang, and Han, Zhu et al.(2021)Zhu, Gao, Wang, and Wu, Lu et al.(2021)Lu, Fan, Wang, Oh, and Rosé], or adopting pre-training using additional large-scale data [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] to further boost their model performance. All of these methods heavily rely on annotations of question-answer (QA) pairs for model training. Intuitively, the more annotated pairs are leveraged, the better performance a model can achieve. Thanks to the development of text-related VQA datasets [Gurari et al.(2018)Gurari, Li, Stangl, Guo, Lin, Grauman, Luo, and Bigham, Wang et al.(2021a)Wang, Xiao, Lu, Jin, and He, Biten et al.(2019)Biten, Tito, Mafla, Gomez, Rusinol, Valveny, Jawahar, and Karatzas, Mishra et al.(2019)Mishra, Shekhar, Singh, and Chakraborty], Text-VQA has achieved rapid progress.
However, the amount of Text-VQA annotations available is still limited due to the sparse labeling of QA pairs in recent datasets. Consider for example the TextVQA dataset [Singh et al.(2019)Singh, Natarajan, Shah, Jiang, Chen, Batra, Parikh, and Rohrbach] whose statistics are illustrated in Figure 1. It shows that only one or at most two QA pairs are annotated in the training images. Meanwhile, we also compute the number of text words presented in each imageWe compute the average of different OCR tokens acquired by the Microsoft-OCR system [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] for the entire training set. and observe that most of the images contain at least 5 text words. This observation indicates that scene text is not fully utilized in the annotations, and hence not fully leveraged by recent methods. A natural question would be – can we fully take advantage of text words in images without incurring extra annotation costs?
As illustrated in Figure 2, we propose to tackle the problem by learning to generate large-scale and diverse text-related QA pairs from existing Text-VQA datasets, using the generated QA pairs to expand the training set and ultimately improving Text-VQA models. Towards this end, we introduce TAG, a text-aware QA generation model, that generates novel text-related QA pairs at scale. It takes text words (the answer) as one of the inputs and aims at generating a question corresponding to this answer by leveraging the rich visual and scene textual cues. TAG is trained using the originally annotated QA pairs and adapts to generate new QA pairs containing scene text words in images that are not utilized in original annotations. No extra human annotation is required in our framework, so the size and diversity of the training data could be easily and largely increased. Since our generation process is disentangled with the training of Text-VQA models, our generated QA pairs can be used by most of the recent methods.
In summary, we introduce a simple yet efficient text-aware generation approach, which automatically and efficiently generates new QA pairs to improve the performance of the current Text-VQA methods. The main contributions of our work are three-fold:
We identify and analyze possible deficiencies of current Text-VQA datasets– sparse annotations of QA pairs - and propose to better utilize unused scene text information within each image to improve the model performance.
To the best of our knowledge, TAG is the first method that explores scene text-related QA pairs generation for improving Text-VQA tasks without additional labeled data.
We consistently demonstrate the effectiveness of our method with two recent Text-VQA models on two Text-VQA datasets. The experimental results suggest that the existing Text-VQA algorithms can benefit from training with the high-quality and diverse QA pairs generated by our method.
Related Work
To study and evaluate the Text-VQA task, several scene text-based datasets are introduced, including VizWiz [Gurari et al.(2018)Gurari, Li, Stangl, Guo, Lin, Grauman, Luo, and Bigham], OCR-VQA[Mishra et al.(2019)Mishra, Shekhar, Singh, and Chakraborty], TextVQA [Singh et al.(2019)Singh, Natarajan, Shah, Jiang, Chen, Batra, Parikh, and Rohrbach], and ST-VQA [Biten et al.(2019)Biten, Tito, Mafla, Gomez, Rusinol, Valveny, Jawahar, and Karatzas]. With the help of these datasets, numerous approaches have been proposed in recent years which increasingly improve Text-VQA performance [Jiang et al.(2018)Jiang, Natarajan, Chen, Rohrbach, Batra, and Parikh, Anderson et al.(2018a)Anderson, He, Buehler, Teney, Johnson, Gould, and Zhang, Singh et al.(2019)Singh, Natarajan, Shah, Jiang, Chen, Batra, Parikh, and Rohrbach, Liu et al.(2020)Liu, Xu, Wu, Du, Jia, and Tan, Gao et al.(2020)Gao, Li, Wang, Shan, and Chen, Han et al.(2020)Han, Huang, and Han, Zhu et al.(2021)Zhu, Gao, Wang, and Wu, Gao et al.(2021)Gao, Zhu, Wang, Li, Liu, Van den Hengel, and Wu, Lu et al.(2021)Lu, Fan, Wang, Oh, and Rosé, Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach, Kant et al.(2020)Kant, Batra, Anderson, Schwing, Parikh, Lu, and Agrawal, Zhang and Yang(2021), Zeng et al.(2021)Zeng, Zhang, Zhou, and Yang, Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo, Biten et al.(2022)Biten, Litman, Xie, Appalaraju, and Manmatha]. LoRRA [Singh et al.(2019)Singh, Natarajan, Shah, Jiang, Chen, Batra, Parikh, and Rohrbach] is an early work that extends the original VQA models [Jiang et al.(2018)Jiang, Natarajan, Chen, Rohrbach, Batra, and Parikh, Anderson et al.(2018a)Anderson, He, Buehler, Teney, Johnson, Gould, and Zhang] with an extra OCR attention branch to select the answer from either a fixed word vocabulary or OCR tokens. Recent studies [Devlin et al.(2018)Devlin, Chang, Lee, and Toutanova, Chen et al.(2020)Chen, Li, Yu, El Kholy, Ahmed, Gan, Cheng, and Liu, Dosovitskiy et al.(2020)Dosovitskiy, Beyer, Kolesnikov, Weissenborn, Zhai, Unterthiner, Dehghani, Minderer, Heigold, Gelly, et al., Carion et al.(2020)Carion, Massa, Synnaeve, Usunier, Kirillov, and Zagoruyko, Liu et al.(2021)Liu, Lin, Cao, Hu, Wei, Zhang, Lin, and Guo, Zhao et al.(2021)Zhao, Jiang, Jia, Torr, and Koltun, Guan et al.(2022)Guan, Wang, Lan, Chandra, Wu, Davis, and Manocha, Zhou et al.(2020)Zhou, Palangi, Zhang, Hu, Corso, and Gao, Baevski et al.(2020)Baevski, Zhou, Mohamed, and Auli, Wang(2022), Wang et al.(2022)Wang, Chen, Wu, Luo, Zhou, Zhao, Xie, Liu, Jiang, and Yuan] show the benefits of transformer for different vision, language and speech tasks. M4C [Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach] develops a transformer-based architecture to fuse different input modalities and iteratively predicts answers through a multi-step answer decoder. Inspired by M4C, more transformer-based models have been proposed with varied structure modifications. Among them, CRN [Liu et al.(2020)Liu, Xu, Wu, Du, Jia, and Tan] constructs a graph network to model the interactions between text and visual objects. LaAP-Net [Han et al.(2020)Han, Huang, and Han] predicts a bounding box to explain the generated answer. SSBaseline [Zhu et al.(2021)Zhu, Gao, Wang, and Wu] proposes to split the OCR token features into separate visual and linguistic attention branches. SMA [Gao et al.(2021)Gao, Zhu, Wang, Li, Liu, Van den Hengel, and Wu] reasons over structural text-object graphs and produces answers in a generative way. LOGOS [Lu et al.(2021)Lu, Fan, Wang, Oh, and Rosé] introduces a question-visual grounding pre-training task to connect question text and image regions. SA-M4C [Kant et al.(2020)Kant, Batra, Anderson, Schwing, Parikh, Lu, and Agrawal] builds a spatial graph to explicitly model relative spatial relations between visual objects and OCR tokens. TAP [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] presents three text-aware pre-training tasks to align representations among scene text, text words, and visual objects.
However, most of the existing works focus on designing sophisticated architectures that leverage the annotated text in an image and overlook the rich text information that is underused by the annotated QA activities. We fully explore the embedded scene text in images and explicitly generate novel QA pairs that can be used to boost the performance of downstream Text-VQA models.
2 Data Augmentation for VQA
Data augmentation has been demonstrated to be an effective approach to improve the performance of the VQA task [Kafle et al.(2017)Kafle, Yousefhussien, and Kanan, Shah et al.(2019)Shah, Chen, Rohrbach, and Parikh, Ray et al.(2019)Ray, Sikka, Divakaran, Lee, and Burachas, Agarwal et al.(2020)Agarwal, Shetty, and Fritz, Tang et al.(2020)Tang, Ma, Zhang, Wu, and Yang, Wang et al.(2021b)Wang, Miao, and Specia, Kant et al.(2021)Kant, Moudgil, Batra, Parikh, and Agrawal]. Kafle et al. [Kafle et al.(2017)Kafle, Yousefhussien, and Kanan] propose to generate new questions using the existing semantic segmentation annotations and templates. Shah et al. [Shah et al.(2019)Shah, Chen, Rohrbach, and Parikh] introduce a cycle-consistent scheme generating question rephrasings to make VQA models more robust to linguistic variations. Ray et al. [Ray et al.(2019)Ray, Sikka, Divakaran, Lee, and Burachas] propose a consistency-improving data augmentation module to make VQA models answer consistently. Agarwal et al. [Agarwal et al.(2020)Agarwal, Shetty, and Fritz] explore data augmentation to improve the VQA model’s robustness to semantic visual variations. Tang et al. [Tang et al.(2020)Tang, Ma, Zhang, Wu, and Yang] use data augmentation to inject proper inductive biases into the VQA model. Wang et al. [Wang et al.(2021b)Wang, Miao, and Specia] introduce a generative model for cross-modal data augmentation on VQA. Kant et al. [Kant et al.(2021)Kant, Moudgil, Batra, Parikh, and Agrawal] adopt the contrastive loss to make the VQA model robust to linguistic variations in generated questions. However, these approaches are designed for the traditional VQA systems that do not emphasize the importance of scene text in their QA tasks. Our method is tailored for the problem of Text-VQA. It takes advantage of the underexploited scene text in images and enlarges the training samples by generating novel text-related QA pairs without the extra labeling cost.
Our Approach
The proposed framework is illustrated in Figure 2, which consists of a transformer-based text-aware visual QA generation module named TAG, followed by a downstream Text-VQA model. Our core module, TAG, carries out text-aware data augmentation tailored for the Text-VQA task and generates novel QA pairs by leveraging underused scene text in an image. After the TAG module generates a large amount of new QA pairs, we directly augment the training data by combining the generated set and the originally labeled set. The augmented set is used by the downstream Text-VQA models to boost the model performance.
The workflow of our method is as follows. Given an image, an OCR system and an object detector are used to detect scene text and visual objects, respectively. As illustrated in Figure 3, our TAG takes the scene text words of interest (the answer words), the visual objects and all the detected OCR tokens in the image as inputs and generates a question explicitly corresponding to the answer. Specifically, the answer words, visual objects, and all the OCR tokens are first represented by high-dimensional features (Section 3.1). Then, the multi-modality information is fully aggregated through a transformer architecture with the attention mechanism (Section 3.2). Finally, the enriched features are used to predict a question to the answer through iterative decoding in an auto-regressive manner (Section 3.3). More details can be found in the supplementary.
We describe the feature embedding strategy of our work. The answer words, detected visual objects, and all the detected OCR tokens are embedded as high-dimensional features and then projected into a common d-dimensional embedding space.
Embedding of extended answer words. We follow [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] to use an extended representation to embed answer words. Given an answer input , we extend the words with labels of objects (detected from the object detector) and scene text OCR words (generated from the OCR system) as a set of text words. A trainable BERT-style model [Devlin et al.(2018)Devlin, Chang, Lee, and Toutanova] is adopted to extract the embedding of those text words, , where , and is the d-dimensional feature vectors for text word. The embeddings of the set of words are used jointly as the feature of the answer.
Embedding of detected objects. Following M4C [Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach], we run a pre-trained 2D object detector, Faster R-CNN [Ren et al.(2015)Ren, He, Girshick, and Sun] to localize visual objects for each image. Two visual object features, including appearance and location features are extracted and then combined together to encode each detected object, , where and is the projected d-dimensional feature vectors for object. Specifically, the feature vector output of the object detector (from the fc7 layer) is used to encode the appearance feature and the relative bounding box coordinates are employed as the location feature.
Embedding of OCR tokens. For the OCR tokens extracted by an OCR system, we construct the embedding for each token containing both its visual and text feature. The visual feature extraction follows the strategy of the above visual object embedding. Additionally, FastText [Bojanowski et al.(2017)Bojanowski, Grave, Joulin, and Mikolov] and PHOC features [Almazán et al.(2014)Almazán, Gordo, Fornés, and Valveny] are extracted for each OCR token to represent its textual cues. A rich OCR representation is thus obtained, , where and is the projected d-dimensional feature vectors for OCR token.
2 Multi-modality Fusion
Once the feature embedding representation from individual modality, , and are generated, they are able to dynamically attend to each other from a stack of transformer layers [Vaswani et al.(2017)Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, and Polosukhin] as shown in Figure 3. The input sequence to the multi-modal transformer is . The multi-modal transformer leverages feature embeddings from different modalities and accordingly models interaction among them through the multi-head attention mechanism. From the output of the multi-modal transformer, we extract a sequence of d-dimensional feature vectors for each modality, which is an enriched feature from a joint semantic embedding space.
3 Text-aware Visual Question Prediction
With the enriched embedding from the multi-modal transformer, the multi-step decoding module predicts a question to the input answer and iteratively generates the question word by word. At each iterative decoding step, we feed in an embedding of previously predicted words, and then the next output word could be either selected from the fixed frequent word vocabulary or from the extracted OCR tokens. Similar to [Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach, Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo], two special tokens and are appended to the word vocabulary, where is used as the input to the first decoding step and indicates the end of the decoding process. Alternatively, the decoding process ends when the maximum number of steps is reached.
During training, our TAG is supervised with the binary cross-entropy loss applied using the originally annotated QA pairs and adapts to generate novel QA pairs during generation. During the QA pairs generation process, we pass an input answer, each of which is selected from the extracted OCR tokens, into the TAG module and generate the corresponding question accordingly. In this way, the generated QA pairs cover a diverse set of scene text which was not directly exploited in the original annotation set. For answer selection, we perform a simple yet efficient strategy that is feeding the OCR token with the largest bounding box as the answer candidate to the proposed TAG. The intuition behind this design is that the scene text with the largest bounding box region is likely to encode semantically meaningful information for scene text-based understanding and reasoning. Also, scene text with a larger font size has a higher chance to be detected correctly without recognition error in general. As we illustrate in our experiments, our simple design facilitates a better understanding of the visual content and provides promising Text-VQA performance. Note that, more high-quality QA pairs could be continuously augmented with a more sophisticated answer-candidate selection strategy. We leave this direction as future work.
Experiments
We evaluate TAG both qualitatively and quantitatively on the TextVQA [Singh et al.(2019)Singh, Natarajan, Shah, Jiang, Chen, Batra, Parikh, and Rohrbach] and the ST-VQA [Biten et al.(2019)Biten, Tito, Mafla, Gomez, Rusinol, Valveny, Jawahar, and Karatzas] datasets. We first present a brief overview of the datasets and implementation details. Then, we empirically validate the effectiveness of our proposed method by comparing it with the existing Text-VQA approaches. Our framework clearly outperforms previous work by a significant margin on both datasets.
TextVQA dataset [Singh et al.(2019)Singh, Natarajan, Shah, Jiang, Chen, Batra, Parikh, and Rohrbach] is a widely used benchmark for the Text-VQA task. It consists of 28,408 images sourced from the Open Images dataset [Kuznetsova et al.(2020)Kuznetsova, Rom, Alldrin, Uijlings, Krasin, Pont-Tuset, Kamali, Popov, Malloci, Kolesnikov, et al.], with human-annotated questions that require reasoning over text in the images. We follow the standard split on the training, validation and test sets [Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach, Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo]. For each question, the answer prediction is evaluated based on the soft-voting accuracy of 10 human-annotated answers [Goyal et al.(2017)Goyal, Khot, Summers-Stay, Batra, and Parikh, Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach, Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo].
ST-VQA dataset [Biten et al.(2019)Biten, Tito, Mafla, Gomez, Rusinol, Valveny, Jawahar, and Karatzas] is another popular dataset for the Text-VQA task. It contains 23,038 images from multiple sources including ICDAR 2013 [Karatzas et al.(2013)Karatzas, Shafait, Uchida, Iwamura, i Bigorda, Mestre, Mas, Mota, Almazan, and De Las Heras], ICDAR 2015 [Karatzas et al.(2015)Karatzas, Gomez-Bigorda, Nicolaou, Ghosh, Bagdanov, Iwamura, Matas, Neumann, Chandrasekhar, Lu, et al.], ImageNet [Deng et al.(2009)Deng, Dong, Socher, Li, Li, and Fei-Fei], VizWiz [Gurari et al.(2018)Gurari, Li, Stangl, Guo, Lin, Grauman, Luo, and Bigham], IIIT STR [Mishra et al.(2013)Mishra, Alahari, and Jawahar], Visual Genome [Krishna et al.(2017)Krishna, Zhu, Groth, Johnson, Hata, Kravitz, Chen, Kalantidis, Li, Shamma, et al.], and COCO-Text [Veit et al.(2016)Veit, Matera, Neumann, Matas, and Belongie]. The standard evaluation protocol on the ST-VQA dataset consists of accuracy and Average Normalized Levenshtein Similarity (ANLS) [Biten et al.(2019)Biten, Tito, Mafla, Gomez, Rusinol, Valveny, Jawahar, and Karatzas].
2 Implementation Details
We use PyTorch to implement our TAGOur implementation is built upon the codebase: https://github.com/microsoft/TAP. that is used to augment the initially labeled data. The augmented dataset is used to improve two recent Text-VQA models, M4C [Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach] and TAP [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo]. M4C† is a variant version [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] of M4C, where the detected object labels and scene text tokens are also included in the text encoder, which further improves the performance.
TAG. We project the multi-modality feature embedding to be = 768 channels. We extract the embedding of extended answer words using the same trainable structure as [Devlin et al.(2018)Devlin, Chang, Lee, and Toutanova]. Specifically, we initialize the weights of the model from the first three layers of and eliminate the separate text transformer. In terms of the object embedding, a Faster R-CNN object detector [Ren et al.(2015)Ren, He, Girshick, and Sun] pre-trained on the Visual Genome dataset [Krishna et al.(2017)Krishna, Zhu, Groth, Johnson, Hata, Kravitz, Chen, Kalantidis, Li, Shamma, et al.] is adopted to extract top-scoring objects on each image and represents each object with its appearance and location features. The Microsoft-OCR system [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] is used to extract OCR tokens per image with each token represented with its appearance, location, FastText [Bojanowski et al.(2017)Bojanowski, Grave, Joulin, and Mikolov] and PHOC features [Almazán et al.(2014)Almazán, Gordo, Fornés, and Valveny]. The multi-modality fusion module is a four-layer transformer with 12 attention heads, which has the same hyper-parameters as . We use decoding steps to predict the output question word by word in an auto-regressive manner.
Training parameters. Experiments are conducted on 4 Nvidia P6000 GPUs. We train TAG for 24K iterations with a batch size of 128. We adopt the Adam optimizer [Kingma and Ba(2014)] with a learning rate of 1e-4 and a staircase learning rate schedule, where we multiply the learning rate by 0.1 at 14K and at 19K iterations. We keep the original parameter settings of downstream Text-VQA models except that we increase their maximum iteration in proportion to the increased size of the augmented data to accommodate the enlarged number of training samples.
3 Main Results
TextVQA dataset. To perform a fair comparison with prior work, we conduct experiments in both the constrained setting (top part of Table 1) and the unconstrained setting (bottom part of Table 1) on the TextVQA dataset [Singh et al.(2019)Singh, Natarajan, Shah, Jiang, Chen, Batra, Parikh, and Rohrbach].The constrained setting means training without extra data and the unconstrained one indicates otherwise. The number of our augmented training QA examples for Text-VQA is 69.2K compared with 34.6K for the original one. In the constrained setting (top), our TAG improves the corresponding M4C and TAP baselines by 1.18% and 2.63% on the validation set, respectively. We note that although LOGOS [Lu et al.(2021)Lu, Fan, Wang, Oh, and Rosé] in Table 1 uses an extra grounding dataset with 1.1 million images for pre-training and yet our method performs better. In the unconstrained setting (bottom), TAG further boosts M4C and TAP baselines by 1.11% and 3.06% on the validation set, respectively. On the TextVQA test set, TAG also obtains significant performance gains over existing methods. This validates the effectiveness of TAG.
We also visualize the generated QA pairs of our TAG in Figure 4. It shows that our TAG generates meaningful QA pairs that are novel compared to the originally annotated ones.
ST-VQA dataset. We also compare our approach with the state-of-the-art (SOTA) methods under both the constrained setting and the unconstrained setting on the ST-VQA dataset [Biten et al.(2019)Biten, Tito, Mafla, Gomez, Rusinol, Valveny, Jawahar, and Karatzas]. We compute the accuracy and ANLS score as the evaluation metrics. The number of the newly built training QA examples for the ST-VQA task after augmentation is 46.8K compared with 23.4K for the original one. Table 2 suggests that TAG achieves SOTA performance and significantly outperforms the baselines. In particular, TAP [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] achieves 50.83%, and 0.598 in terms of the accuracy and ANLS score on the validation set with additional TextVQA and 1.4 million large-scale pre-training data, while TAG improves these results by a significant 2.70% and 0.022 with only additional TextVQA data. In addition, we submit the prediction results of test set on the ST-VQA test server. The results show that TAG with TAP achieves the SOTA performance with ANLS score of 0.602 on the test set. Without bells and whistles, our approach greatly outperforms the baselines, M4C [Hu et al.(2020)Hu, Singh, Darrell, and Rohrbach] and TAP [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo].
4 Ablation Studies
We conduct extensive ablation studies to demonstrate the effectiveness of TAG using TAP [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo] under the constrained setting on the TextVQA validation set.
Contribution of each modality in TAG. To understand the contribution of different input modalities to the success of TAG, Table. 3 summarizes the performance of our framework when a certain modality is removed. It suggests that when both the visual objects and OCR tokens modalities are removed, the performance of our TAG decreases by 3.78%. On the other side, when removing the visual objects modality and OCR tokens modality separately, the performance drops by 3.59% and 3.41%, respectively.
Impact of the answer selection strategy. To better explore the performance of our TAG, and understand how different answer selection strategies would affect the model performance, we design several experiments over the choice of input answer selection strategy. Our method adopts the largest OCR word as the answer candidate for TAG. We compare this strategy with other possibilities in Table. 4. The table shows that, if we use a random OCR token as the input answer, the performance drops by 3.28%. On the other hand, if we increase the number of answer candidates by including the top-3 largest OCR tokens to augment the labeled data by 3, the performance boosts additional 0.19% as compared to the largest strategy while it introduces 3 training time. To achieve a better balance between training efficiency and accuracy, we consider the OCR token with the largest bounding box as our final setting for the input answer to TAG. As we have mentioned previously, more high-quality QA pairs could be continuously augmented with a more sophisticated answer-candidate selection strategy. We leave this direction for future work.
Conclusion
We propose a novel architecture TAG, a text-aware visual question-answer (QA) generation method to deal with the sparse annotation of existing Text-VQA datasets. Our approach leverages the rich yet underexplored visual and scene text information and directly enlarges the existing training set by generating high-quality and rich QA pairs without extra labeling cost. Without bells and whistles, experimental results show that our generated QA pairs boost the performance of recent Text-VQA models by a large margin on both TextVQA and ST-VQA datasets.
References
Appendix A Additional Details of Our Approach
Text-aware Visual Question Prediction. We leverage the powerful capability of the attention mechanism in transformers [Vaswani et al.(2017)Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, and Polosukhin] to capture the interactions among the extended answer words, visual objects, and OCR tokens. Our decoding module is based on a dynamic pointer network [Vinyals et al.(2015)Vinyals, Fortunato, and Jaitly], which allows both copying words via pointing, and generating words from a fixed vocabulary obtained from the training set.Our decoding module is implemented following the implementation of the decoding module in TAP [Yang et al.(2021)Yang, Lu, Wang, Yin, Florencio, Wang, Zhang, Zhang, and Luo]. We keep their default hyper-parameters except otherwise noted.
Hyper-parameters. Table. A1 overviews the hyper-parameter settings of TAG. We use the original parameter settings of downstream Text-VQA models except that we increase the maximum number of iteration in proportion to the increased size of the augmented data to accommodate the enlarged number of training samples.
Appendix B Additional Qualitatively Visualization of TAG
We present additional visualization results of generated QA pairs by TAG on the TextVQA training set in Figure. A1.