WebQA: Multihop and Multimodal QA
Yingshan Chang, Mridu Narang, Hisami Suzuki, Guihong Cao, Jianfeng Gao, Yonatan Bisk
Introduction
Web search is a multimodal experience: Will I find my answer on the image search tab or within text snippets? In contrast, most deployed Question Answering (QA) systems treat the web as a text-only landscape of facts to be extracted, ignoring the knowledge present in images. This has two fundamental limitations: 1. The text-based web is impoverished , and 2. This form of information extraction is inefficient. For example, when searching to see if a park has picnic tables, surfacing an image of the picnic area answers the question immediately, rather than wading through pages of reviews hoping someone happened to mention this fact. QA engines need to move to treating the Internet as a multimodal trove of information, but this requires multihop reasoning on either images or text.
To this end, datasets are rapidly emerging. But they either use pre-defined templates for the curation of multihop multimodal QA pairs , or encourage a “question decomposition + rerouting to uni-modal model” approach to superficially solve the problem. However, when humans absorb knowledge, there is no need to distinguish whether the knowledge was learned from books versus images, or whether a piece of knowledge is a composite of multiple scattered fragments versus being carried by a single one. We argue that genuine progress in reasoning over linguistic notions of meanings and visually grounded meanings under the same representation framework depends on the development of a unified system that indiscriminately treats snippets and images as knowledge carriers. On top of that, the goal includes better extraction, integration and summarization abilities in a heterogeneous information landscape.
To facilitate this research intersection, in this work we propose a novel benchmark, WebQA, for multi-hop, multi-modal, open-domain question-answering where all questions are knowledge-seeking and resemble real-world use cases. Success on WebQA requires a system to a) incorporate both text and images, b) retrieve relevant knowledge in either modality, c) aggregate information from multiple sources via logical or numerical reasoning, and d) generate answers in natural language. We experiment with state-of-the-art multimodal reasoning and text generation models, whose failures indicate promising future directions.
Related Work
Many datasets and tasks can be broadly considered “question answering.” For example, VQA is one of the widely studied tasks at the intersection of language and vision. Nevertheless, it is unclear how VQA models should be adapted to open-domain scenarios. This is largely due to the simplification of VQA tasks into classification over a fixed vocabulary of frequent answers. Recent work on video has also adopted a multiple-choice format. In contrast, OK-VQA broadens the task to knowledge-seeking questions. OK-VQA and our task differ in the role of images. Images in OK-VQA are regarded as part of the query rather than as part of the knowledge source, and can only be processed after retrieval.
Within the natural language community, QA datasets are experiencing a similar transition from multiple-choice and span prediction to the harder free-form answer generation paradigm. Multi-hop question answering has recently taken the spotlight as it aligns with the multi-hop nature of how humans perform reasoning during knowledge acquisition, leading to a proliferation of benchmarks .
There have been several recent benchmarks for reasoning over input and contexts in multiple modalities . MultiModalQA made the first foray into complex questions that require reasoning over snippets, tables and images. It focuses on cross-modal heterogeneous knowledge extraction. However, questions are generated from templates. Once a template is detected the task reduces to filling in blanks with modality-specific answering mechanisms.
ManyModalQA also deals with snippets, images and tables. However, the primary challenge their design addresses is the choice of answer modality — rather than knowledge aggregation or extraction. Our focus is more about representing world knowledge in a unified space, than about distinguishing the answer modality, since mastering the former may naturally eliminate the need to classify questions according to the answer modality.
Finally, MIMOQA introduces a new concept of “Multimodal Input Multimodal Output” which highlights accompanying a textual answer with an image in order to enhance cognitive understanding. MIMOQA requires selecting a text span and an image from the context as an output pair. Their approach is nicely complementary to ours. Where we differ, is that our task also requires aggregation and summarization before producing the final natural language answer, whereas the outputs required by MIMOQA are not completely digested by the model. Here, we refer “digesting” to the ability to produce a reasonable output which cannot be directly copied from the input.
Tables 1, 2 and Appendix E provide comparisons between WebQA and related datasets. No existing multimodal or knowledge-seeking benchmark requires the answers to be complete, free-form natural language sentences, as opposed to extractive spans, or elements from a finite set. Additionally, previous work has not supported both natural language generation (NLG) evaluation and accuracy-style evaluation as we do. To this end, we highlight that a) in WebQA more importance is attached to digesting, aggregating and summarizing information as answers cannot be simply copied from an existing text span or image patch, b) WebQA requires the source retrieval stage in addition to VQA, which better simulates the full reasoning pipeline during a web search, and c) answers in the form of a natural language sentence better transit to downstream applications such as conversational agents and voice assistants.
Task Formulation
As in Fig 1, examples consist of a question , a set of positive sources (in green), a set of distractor sources (in red) and an answer . Each source can be either a snippet or an (image, description) pair. Each image is accompanied by a description to resolve names or geographic information not present in the image itself, but serve as critical links to references in the question. We include both a restricted () and full () setting.
We decompose the task into two stages. First, given and , the model identifies the sources from which to derive the answer. The second stage is question answering where the model takes and the chosen sources as context , to generate an answer . Ideally, a single-stage system would jointly process to produce , but we are unaware of any modeling approaches that can consume sufficiently large multimodal contexts to achieve this, so this is left to future work.
WebQA
Following the paradigm popularized by search engines, we structure our data as having answers that can be found either via image search or general web (text) search. Note, WebQA does not contain questions that need an image and an (independent) snippet as knowledge sources. However, all image-based questions already require processing both images and text as image descriptions provide necessary information. Below we outline how both types of questions are collected, structured, and filtered for quality.
We collect both multi-image questions that require stitching two images to answer and complex single-image questions. Rich multi-image questions do not naturally exist at scale in user search logs, likely because users do not issue queries they believe search engines cannot handle, thus we turn to crowdsourcing.
We presented annotators with a set of six related images and asked them to produce three QA-pairs by selecting one or two images per pair that are necessary to answer the question. We require that at least one of the three pairs utilizes two distinct images. Additionally, we instructed annotators to avoid questions that: a) are simple facts (e.g. “How many wheels does a car have”); b) are easily answered by a text-only search; c) are bound to a specific image; d) ensure every question is meaningful without paired context. This elucidates one of the key differences between the well-known VQA task and ours. In most VQA style tasks, every question is about a paired image, whereas in our task images serve as knowledge sources over which to reason, and do not serve the role of augmenting the question. To assist annotators, each image is accompanied by a description extracted from Wikipedia. This description is only to be used to confirm the name or location of the objects depicted. The answer has to be derived from visual clues.
Images were crawled from Wikimedia Commons via the Bing Visual Search API. Wikimedia’s topic list cannot be used directly as most categories are (visually) uninteresting. We seeded with natural scenes and iteratively refined the image pool by removing categories flagged as (visually) uninteresting. This resulted in categories like animals, plants, attractions, and architecture (Fig 3).
We produce a set of both text- and image-based hard negatives for models to sift through for every question. Text sources are extracted from relevant passages on Wikipedia based on noun chunks in the question, while limiting overlap to avoid false negatives. For images, we leverage Bing APIs to find similar images with respect to both the description (via Bing Image Search) and the visual content (via Bing Image Insight). In total, we collect 25K image-based questions, each requiring an average of 1.4 visual sources, and paired with 15.3 text and 15.9 visual distractors. Question prefixes are visualised in Fig 2.
Categorization
We categorize questions into open and closed classes. Closed class questions include: color, shape, number (i.e.“how many”), yes/no (Y/N), and “multi-choice” (MC). The rest are open class questions.
Adversarial splits
We construct our test set to be out-of-distribution when possible to reward models with better generalization and reasoning. For color, shape, and number questions, we partition the answer set and ensure that the majority class during training does not carry over to testing. For the “Y/N” and “MC” classes, we trained models on 10 random train-test splits and consistently difficult samples across splits were placed in the test set. Finally, we randomly split questions from the open-class “other”.
2 Answers From Text
We collected multi-hop QA pairs that involve combining knowledge from 2 snippets. To generate diverse, yet consistent, topics for mining difficult multi-hop reasoning questions, we construct clusters of similar entities, but where text snippets had low overall n-gram overlap or semantic similarity (yielding 8K clusters). We provide annotators with four snippets to prevent and allow them to contribute facts they researched to help answer the question.
For text distractors we mine passages from Wikipedia that contain noun phrases from the question and choose those with the highest lexical overlap but lacking reference to the answer. For image distractors, we use the images and descriptions present on the aforementioned Wikipedia pages, again filtering for those with high lexical overlap. In total, we collected 24K text-based questions, each requiring 2.0 text sources, and paired with 14.6 text and 11.6 visual distractors. Lacking clear criteria for question categorization, we do not construct an adversarial test split, but instead simply sample randomly.
3 Quality Control
We ensure the data quality via crowdworkers training and expert-feedback-in-the-loop, which are found to be effective ingredients in crowdsourcing . The initial pool of annotators were trained with a tutorial and selected via a qualification task. Additionally, we released the annotation task in batches to spot check quality after every batch, followed by sending constructive feedback to correct any deviation from our expectations. Workers who failed multiple times were de-qualified. Crowdsourcing data is challenging in that crowdworders are usually income-driven and will stick to a fixed answer generation pattern once they find it lucrative. To better align the crowdworkers’ incentives with our goal, we generously bonus out-of-the-box thinking. All data was then also run through additional validation HITs to ensure agreement. Annotator pay averaged $13/hr overall (lower on the initial qualification and higher on the annotation/validation). Appendix A contains rubics and interfaces.
4 Dataset Statistics
In total, WebQA has over 34K training QA pairs, with an additional 5K and 7.5K held out for development and testing. Overall Statistics are summarized in Table 4 and language distributions are presented in Table 3.
44% of image-based queries and 99% of text-based queries require two or more knowledge sources. This is verified by crowdworkers during validation to ensure that multiple knowledge sources provide non-overlapping information and cannot be replaced by each other. Additionally, as image sources also require understanding the caption, even single-image queries require multi-source reasoning.
Topics
Fig 3 provides a qualitative sense of the wide range of topics covered in WebQA. In contrast to MultiModalQA, the images in WebQA concentrate on the natural-world, events, and locations rather than digital artifacts (e.g. posters/logos). Snippets also exhibit a wide range of topics from contemporary science to ancient mythology. When comparing the topic clouds, it is clear that image-based queries more often relate to physical entities while text-based queries tend to be more abstract.
Metrics
WebQA requires a model to answer open-domain questions and cite its sources. Therefore, we evaluate model performance with respect to both relevant fact prediction and question answering. While fact retrieval is easily evaluated via F1, language fluency and accuracy metrics are nuanced.
Our task expects fluent and complete sentences as answers, which we believe are appropriate for applications such as voice assistants or conversation agents. Therefore, the quality is measured as both fluency and accuracy. On each testing sample we collected five full-sentence answers written by humans. In addition, we collected one keyword answer by asking human annotators to rephrase the full-sentence answer into a succinct minimal semantic form.
We measure fluency via BARTScore , a newly proposed NLG evaluation metric based on accurate measurement of paraphrase quality. BARTScore(, ) measures the probability of generating from . In our setting, this is computed as BARTScore(r, c), which can be interpreted as the probability of generating a candidate given a reference. Since BARTScore is based on the generation likelihoods, it does not distribute neatly across $$. So we normalize BARTScore(r, c) by the identity score BARTScore(r, r). On top of that, we make the normalized score bounded by 1. Finally, we choose the best score for a candidate across all references, as illustrated in Eq. 1.
This formulation a) prioritizes semantic agreement and is robust to functional words misplacement, b) does not heavily punish short sentences (i.e. 4 words) as BLEU4 does, c) penalizes word reordering / disfluencies d) and unlike BERTScore , which indiscriminately treats all colors or all shapes as nearly identical, BARTScore better captures small but critical differences. However, no language based embedding metrics accurately evaluate visual phenomena, so we also introduce an accuracy metric.
Accuracy
To ensure answer accuracy we use the collected keywords. Note, our paradigm differs from both open-domain text QA which focuses on lexical F1 and visual QA which uses a multiple choice evaluation. F1 rewards copying the question even if the key information is missing (e.g. the wrong color or count is chosen). Conversely, multiple-choice paradigms are not applicable to evaluate generated sentences. The goals of measuring accuracy on WebQA are: 1. Detect the presence of key entities. 2. Penalize the use of any incorrect entities. 3. Avoid penalizing semantically relevant but superfluous words. We are unaware of any solution to all of these criteria in the naturally mixed setting of our data (open-domain entities with a nearly closed-domain set of properties), so we propose an appropriate metric to tackle the different styles of answers.
Given the aforementioned question categorization for visual queries, questions having closed answer domains should be evaluated via F1 that tests for precision (avoiding a model producing both Yes and No to game the metric). We define the answer domains of those question categories () in Table 5. For the remaining visual queries and all textual queries, they have diverse and unrestricted answer domains. So, there are good reasons to believe that the probability of cheating by guessing a long list of keywords is small and would be penalized by BARTScore, so we evaluate accuracy via recall (RE). With as a candidate output, for correct answer keywords, and for question category, Equation 2 sketches our score.
Finally, we report the average combined fluency and accuracy score * across all test samples as a single evaluation result for a system.”. Unlike the categories in Table 5 where it is wrong to output incorrect elements, including more elements in additional to the correct element in an answer may be correct if asked to compare the elements. We leave this to future NLG evaluation research as outside the scope of this work.
Baseline Models
We test existing models on WebQA in both fune-tuned and few-shot settings. The former fine-tunes a pre-trained vision-and-language transformer on our source retrieval and QA tasks, while the latter (PICa ) prompts GPT-3 with engineered prefixes. Note, since the answer space in WebQA is inappropriate for the classification approach (3K answers) considered by most VQA models, these models , cannot be applied in our generative task. At present, VLP and Oscar are the top generative multimodal transformers. Oscar is built on VLP so we chose VLP as more canonical but include the state-of-the-art visual features of VinVL implemented in Oscar+ . Other recent models may also have complementary strengths. To test the largest possible language model, we also run PICa which leverages VinVL based captioning to augment GPT-3 with oracle source knowledge. Finally, to simulate the full retrieval setting, we ran zero-shot sparse and dense retrieval models over the entire collection of sources.
We train two separate models for source retrieval and question answering on from released VLP weights.
Text segments, including the questions, answers, textual sources and image captions, are tokenized by the Bert-base-cased tokenizer. Each image is represented by 100 regions predicted by an object detection model, which is a variant of Faster RCNN with an ResNeXt-101 FPN backbone, pretrained on Visual Genome . We take the output of fc1 layer from the object detection network an 2048-dim feature and finetune the fc2 layer. We also experiment with the latest state-of-the-art representations from VinVL. Comparing to ResNeXt-101 FPN, the major advances of VinVL include a larger backbone (ResNeXt-152), replacement of FPN by C4 and better pretraining enriched by attribute information.
Source Retrieval
Candidate sources are fed to the model one by one. Each pass takes the concatenation of [CLS], , [SEP], , [SEP] and estimates probability of a particular source being selected. Let and denote the set of gold sources and distractors for a sample. The loss function is as follows.
Question Answering
We feed [CLS], , [SEP], , , [SEP] to the Transformer, where attention masks are applied to tokens in to satisfy the auto-regressive property. We use standard Masked-Language-Modeling loss during fine-tuning. We decode by iteratively appending a [MASK] to the end of the input, replacing it with a predicted token and appending a new [MASK] for the next timestep. Generation stops upon seeing [SEP], [PAD], or reaching a maximum length. We use beam search () and choose the most confident output for evaluation.
Model Variants
In addition to the standard VLP trained on full data, we also include two modality-specific variants VLPI and VLPT, which are trained on image- or text-based queries only as opposed to the full data, in order to reveal gains and losses resulted from the complexity of presenting models with data from both modalities,
2 Zero-shot Full-scale Retrieval Approach
For end-to-end performance in an open-domain setting, we consider the entire collection of sources as our retrieval space (390k images and 540k text sources). Since running VLP-based retrieval of the test set over the entire source collection is prohibitively expensive (3 years), we consider both sparse retrieval (BM25 ) and dense retrieval for a coarse filtering. Dense retrieval was achieved via CLIP encoding all image and text sources, as well as all questions. Next, using the modality knowledge, we rank all image/text sources based on the question-source similarity.
3 Few-shot Question Answering Approach
PICa is the strongest model on OK-VQA , where GPT-3 is prompted to generate answers given a few training samples as prefix. We adapt PICa to our QA task using oracle sources to provide an upper-bound for the best possible performance of the strongest known models. PICa (and GPT-3) exhibit unstable behaviors on source prediction when presented with 4 choices (as it is most familiar with 4-way multiple choice tasks). Due to the inability to fine-tune, we cannot construct a truly fair comparison of PICa with our other baselines on our full pipeline.
We construct an input prefix by concatenating the pre-selected training examples, the context and the question of a testing sample. Since PICa’s transformer backbone does not accept visual input, each image is described by three text segments, namely 1) a wikipedia description, b) a caption generated by Oscar+ and c) a list of tags predicted by Oscar+. Limited by the maximum input length, we experimented with an 8-shot setting. If the input length exceeds the maximum length, we decrease the number of shots until it fits in the length budget.
Training examples to be included in the prefix for each testing sample are selected according to both question and source similarities. We use CLIP to extract text or image encodings for questions, oracle snippets and oracle images. When multiple sources exist, we take the average of pairwise similarities between sources in one sample and sources in the other.
Prompt Design
We use XML-style brackets to denote different text segments. See Fig 4 for what constitutes a prompt for a text- or image-based query.
Results & Analysis
Below we present results and analysis of our baselines’ performance on WebQA. We include question-only baselines for both VLP and PICa to investigate how effectively models use the sources. VLP scores 22.6 on the proposed metric when evaluated end-to-end (Table 6). Modest improvement can be achieved by knowing the gold sources, showing room for growth on retrieval correctness. We observe that the latest best-performing visual encoder, VinVL, does not lead to significant gains. This may support the argument that the missing aspects from the status quo are more reflected in cross-modal information sharing than in the imperfection of uni-modal representations. PICa achieves a large gain over VLP. Promising as it is, we later show that, while pursuing the benefits of scaling up is one thing, there is still a lot remaining to be done to combat the diminishing returns involved with scale . We show that humans can perform our task with ease (i.e. achieving 94 and 55 ) computed via cross-evaluation on multiple (3-6) references provided by different annotators to prove robustness and consensus. While models’ scores are high, reaching human-level accuracy is not within sight.
Crucial to a complete system design is multimodal source retrieval. We investigate the effect of retrieval scale (Table 7) and dense versus sparse retrieval approaches. For the VLP-based model, sources are selected if its binary classification confidence is above a specified threshold. While the optimal thresholds for different models may vary, for fair comparisons we use 0.2, which is optimal for VLP on the development set.
VLP achieves 68% F1 given a restricted set of candidates. Indicating that it can model semantic relevance, despite its lack of scalibility. In comparison, we use simpler and less expensive approaches when scaling up to the full collection which causes our overall performance to degrade substantially (likely due both to the ambiguity and the weaker underlying document representations). The dense retrieval method suffers from a greater performance drop compared to sparse retrieval. Having VLP rerank the top 20 sources predicted by CLIP doubles F1, which holds promise for a future of large-scale coarse-to-fine retrieval that strikes a better accuracy-efficiency balance. See Appendix D for additional retrieval results.
2 Question Answering
Table 8 provides an accuracy breakdown with respect to question categories. A noticeable pattern is that models are more capable of solving text-based queries than image-based queries. Both VLP and PICa greatly surpasses the question-only baseline and VLP performs favorably against VLPT, demonstrating reasonable use of sources and the effectiveness of combined training.
On the other hand, image-based queries pose a much harder challenge. VLP and VLPI are no better than the question-only baseline on image-based queries. While this may be an issue of the sources being ignored, we also attribute this to the fact that the image-based testing samples are intentionally constructed to prevent the success of any superficial correlations that can be drawn from the training set (e.g. the majority answers in each category). We observe a similar issue with PICa. Although PICa consistently outperforms VLP, it does not demonstrate an appropriate utilization of the provided sources, which is especially true on “Y/N”, “MC”, “Shape”, “Number” and “Other” question categories. PICa has a surprising amount of knowledge embedded in its parameters, but unlike with text, on images it shows very little improvement from the inclusion of visual sources, as such it is still lacking the ability to explicitly and effectively use the retrieved sources, which might be crucial for further progress towards human accuracy.
We argue that performance is bottlenecked by the lossy textual representation of images consumed by PICa, thereby calling for concerted effort from both language and vision sides to build a unified representation rather than simply relying on one modality being translated to the other. For future research, we expect to explore whether symbolic or compositional representations in a structured problem space could equip a generative model with skills to perform aggregation beyond simple extraction.
3 Qualitative Analysis
Finally, we perform a qualitative analysis of the model’s failures for both image- and text-based questions. Table 9 includes two image-based and two text-based examples with commentary (additional analysis in Appendix F). Both image questions are clean examples of producing logically consistent and fluent sentences which are incorrect. The first matches the negation but the answer should have been yes, while in the second, the model runs away with a very logical hallucination (heads wear helmets).
In the text examples, we see a different pattern. Here the model is more easily able to copy facts from the source texts, but still demonstrates a lack of understanding or reasoning. In the first example, the model appears to know it is looking for a number, but choosing one via direct copying rather than performing the arithmetic necessary to combine both facts. In the second case, the model finds a relevant span selection (as is commonly the only thing necessary for text QA tasks), but does not understand that the question is asking about a method of diagnoses versus a symptom.
None of the questions presented here require complex problem-solving skills. They follow rather simple implication, addition, or visual extraction patterns which are out of reach for current models (uni- or multi-modal).
Conclusion
WebQA is a new multi-hop, multi-modal question answering challenge for our community. Designed to simulate the heterogeneous information landscape one might expect during a web search, WebQA covers a series of open-domain general visual queries while also forcing models to still reason about text. Our task requires a system to determine relevant sources, perform aggregation and reasoning. We also propose a novel general recipe for evaluation on WebQA which measures both fluency and accuracy.
Neither the versatile V&L transformer nor the large-scale text generator present a nearly-there solution. We provide both a restricted and full retrieval setup, to bridge multimodal QA and IR research. This dataset not only mirrors our everyday experience on the web, but provides a playground for the community to explore important sub-challenges, targeting the creation of a single model for multimodal reasoning, knowledge aggregation, and open-domain visual understanding.
WebQA aims to facilitate research into constructing a single model which can 1) retrieve relevant documents, and 2) integrate information across a large context window including multiple paragraphs and images, in order to 3) generate fluent natural language answers.
References
Appendix A Data Annotation Details
For quality control, we included a qualification task with 15 hard coded QA pair annotations, some of which obviously violate the annotation guidelines. Annotators had to point out the problematic pairs and explain in what ways they did not follow the instructions. We restricted to crowdworkers located in the US or Canada, with a general requirement of over 1,000 previously approved HITs with at least 95% approval rate. Additionally, one has to score 80% or higher on our qualification task before getting access to our main task. We gave workers who achieved 60% - 80% at their first attempt a second chance because we believe that workers who had the patience to complete their first attempt were more coachable than others.
Image Filter HIT
We designed a Filter HIT as a pre-step to obtain groups of related images as prompts for the QA-pair creation task. We present 10 images at a time, which are returned by an Image Search API call using the same search term. Annotators were told to a) select 3 out of the 10 that are distinct but related in some ways, and b) give a label that best summarizes the commonality. After having all these image triples, we paired up triples to form groups 6 according to the cosine similarity between their topic labels. We tuned similarity thresholds to make sure that within each group all images fall under the same topic but still have enough dissimilarity to facilitate both connection-based and comparison-based QA-pair construction.
QA Pair Creation HIT
The main annotation task (QA-pair creation task) was released batchwise. We spot checked data quality after every batch and sent targeted feedback when we noticed any deviation from our expectations. Workers who constantly failed to follow the guidelines were de-qualified. Crowdsourcing data is challenging in that crowdworders are usually income-driven and will stick to a fixed answer generation pattern once they find it lucrative. To better align the crowdworkers’ incentives with our goal, we gave generous bonuses to the annotations that demonstrate out-of-the-box thinking.
QA Pair Validation HIT
Multiple Human References Generation HIT
Appendix B Visualization of Image Question Prefixes
Appendix C Classification Based Coverage
The figure below shows the test set coverage of Top-K training keywords (image-based). All keywords (5k) provides only 70% coverage. The full sentence answers are almost entirely unique, suggesting that classification-based approaches are at a significant disadvantage on WebQA.
Appendix D Additional Results on Full-scale Retrieval
Assuming known answer modality, CLIP achieves 91% and 64% recall rate for image- and text-based queries when 2,000 candidates are retrieved. Without the modality knowledge, the recall rate for image-based queries is zero because the question-image similarities are systematically lower than question-text similarities. Future work may fine-tune dense multimodal retrieval models to close the gap between question-image and question-text similarities.
Appendix E Comparing WebQA and recent benchmarks
We succinctly contrast WebQA against existing knowledge-aware and multimodal datasets in the main paper. We provide here a more complete clarification of the new contributions of WebQA over relevant datasets in prior work in terms of data size, modalities and reasoning levels.
WebQA differs from QAngaroo, HotpotQA, ComplexWebQuestions, HybridQA and NaturalQuestions either in the knowledge-awareness or the involvement of both text and image modalities. OK-VQA, MultiModalQA, ManyModalQA and MIMOQA qualify as both knowledge-seeking and multimodal. Thus we explain them in detail.
OK-VQA OK-VQA and our task differ in the role of images. Images in OK-VQA are regarded as part of the query rather than the knowledge source, so source retrieval is not required. However, images in WebQA serve as the knowledge rather than part of the query and can only be processed after retrieval. OK-VQA Topics:
MultiModalQA MultiModalQA and WebQA differ in the way qa-pairs were constructed and the answer schema. First, MultiModalQA questions are generated from templates. While this facilitates the data generation process, it does not mirror the way real users construct queries. Once the question template is detected, the task reduces to filling in blanks with modality-specific answering mechanisms. This problem-solving manner might not generalize to queries issued by real users where an underlying template is less obvious. In contrast, queries in WebQA are written by annotators, and more structurally diverse. Second, MultiModalQA requires different answer schemas for TextQA, ImageQA and TableQA. TextQA expects a span, “yes” or “no” as an answer. ImageQA expects selection from a fixed answer vocabulary determined by the training set. TableQA expects “yes”, “no”, a table cell, or a summary of more than one table cells via a predicted aggregation operation (i.e. SUM / MEAN / COUNT). We unify the answer schema to be a complete natural language sentence and use an open answer set, so neither span prediction nor classification over a fixed vocabulary suffice. MultiModalQA Topics:
ManyModalQA The primary challenge ManyModalQA addresses is the choice of answer modality – rather than knowledge aggregation or extraction. Our focus is less about distinguishing the answer modality, than about representing world knowledge in a unified space, since mastering the latter may naturally eliminate the need to classify questions into different buckets according to their answer modality. Also, to avoid ambiguity and for easy evaluation, ManyModalQA restricts all answers to be a single word. Therefore, the following question answering is a multiple choice task from [all words in the given context + a pre-defined answer vocabulary]. We argue that multiple choice is an unnatural simplification, because the finite and static answer space imposes a hard limit on the capacity of an answering system, especially when we consider unfamiliar domains, constant shift of world states, and unlimited coverage of the Web. This leads to us formulating WebQA as a free-form generation task, which, although it introduces new challenges for evaluation, better resembles real-world use cases and suits the needs of downstream applications such as voice assistants or conversational agents. Last but not least, ManyModelQA is much smaller than WebQA in size. ManyModalQA Topics:
MIMOQA requires selecting a text span from a given context and an image from a set of related images as a multimodal output pair. However, this task formulation does not support queries whose answers should be a digested and summarized version of the given sources instead of a span. WebQA requires further information aggregation and summarization through either numerical or logical reasoning, highlighting the major advantage over MIMOQA in reasoning levels. Plus, WebQA tests natural language generation ability while MIMOQA only requires span prediction and retrieval, both under the classification banner.
Appendix F Additional Qualitative Analysis
Appendix G Datasheet for WebQA
WebQA was created to drive the research progress in multihop, multimodal question answering, which would bridge the gap between the natural language and vision community.
The initial version of WebQA was created by Yingshan Chang and Yonatan Bisk on behalf of Language Technology Institute, Carnegie Mellon University, and Mridu Narang at Microsoft Bing.
Who funded the creation of the dataset?
Microsoft Research and Bing provided the funds for crowdsourcing and web crawling.
G.2 Composition
Each instance is a tuple of (Knowledge Sources, Question, Answer), where a knowledge source can be either an image assisted by a caption, or a snippet. Questions and Answers are in textual form.
How many instances are there in total (of each type, if appropriate)?
WebQA is structured as having answers that can be found either via image search or general web (text) search. So there are two folds of data, containing 22,423 image-based queries and 24,343 text-based queries, respectively. There are 600K images crawled from Wikipedia and 750K snippets crawled from the general Web (mostly from Wikipedia) serving as potential knowledge sources.
WebQA is a sample of instances. It is presumably intended to be a random sample of instances representing what one might encounter during a real web search experience. Manual efforts were put in to ensure reasonable coverage and diversity. Only qualitative tests were run to show the inclusiveness.
What data does each instance consist of?
Each data instance consists of text and images.
Is there a label or target associated with each instance?
The answer component is regarded as the target. Each instance is associated with one human-written answer in the format of a complete natural language sentence. Additionally, each instance in the testing set has multiple (3-6) full sentence answers as well as a keyword answer annotated by humans, which is supposed to be a succinct rephrasing of the corresponding long-form answer.
Is any information missing from individual instances?
Are relationships between individual instances made explicit (e.g., users’ movie ratings, social network links)?
There are no relationships between instances except for the fact that multiple instances may share knowledge sources.
Are there recommended data splits (e.g., training, development/validation, testing)?
The dataset comes with specified train/dev/test splits. The split on the text-based fold was determined randomly while the test split on the image-based fold was adversarialy selected to prevent spurious shortcut learning from inflating the metrics.
Are there any errors, sources of noise, or redundancies in the dataset?
Erroneous instances were pruned during the validation process after the initial collection, where we had human annotators report mistakes and inconsistency. The released version is clean.
No. All the information crawled from the Web was downloaded and fixed when the dataset was constructed.
No. All data was derived from crowdsourcing and publicly available content on the web.
No, data was specifically pulled from known vetted resources (e.g. Wikipedia / Wikimedia).
Does the dataset relate to people?
G.3 Collection Process
The questions and answers were curated by crowdsourcing. The knowledge sources were mined from the web that were directly observable.
Crowdsourcing relied on the Amazon Mechanical Turk platform. Web crawling was assisted by Bing Visual Search and Wikipedia APIs.
All question-answer pairs were human-curated. Knowledge sources for each sample are determined by their relevance to the question-answer pair.
Crowdworkers are paid with an average hourly wage above $13.
Over what timeframe was the data collected?
WebQA was collected and validated from Oct 2020 to Aug 2021.
Were any ethical review processes conducted (e.g., by an institutional review board)?
Does the dataset relate to people?
G.4 Preprocessing/Cleaning/Labeling
After the initial collection, each sample was validated by 2 or 3 crowdworkers. Problematic samples were discarded. Testing samples with low human agreements were discarded. Besides, each sample in the image-based fold was assigned a question category label produced by a text analysis algorithm.
The raw unprocessed data (consisting of crowdsourcing output, history versions of unpruned dataset) is saved.
Is the software used to preprocess/clean/label the instances available?
While a script running a sequence of commands is not available, all codes used to process the data is open source on Github.
G.5 Uses
The dataset was introduced in the paper WebQA: Multihop and Multimodal QA.
Is there a repository that links to any or all papers or systems that use the dataset?
Papers using this dataset will be listed in https://webqna.github.io/ or linked from the EvalAI leaderboard.
What (other) tasks could the dataset be used for?
WebQA can be used for modelling works in the areas of knowledge retrieval, multimodal reasoning and open-domain question answering.
No. There is minimal known risks for harm.
Are there tasks for which the dataset should not be used?
G.6 Distribution
Yes. WebQA will be made publicly available.
How will the dataset will be distributed (e.g., tarball on website, API, GitHub)?
See https://webqna.github.io/ for downloading instructions.
When will the dataset be distributed?
WebQA will be released to the public in Sep 2021.
The crawled data copyright belongs to the websites that the data originally appeared in (e.g. Wikimedia Foundation). WebQA will be distributed under freely to academic researchers upon request.
Have any third parties imposed IP-based or other restrictions on the data associated with the instances?
Do any export controls or other regulatory restrictions apply to the dataset or to individual instances?
G.7 Maintenance
WebQA is supported and maintained by Language Technologies Institute @CMU and Microsoft Research, and the leaderboard is hosted on EvalAI.
How can the owner/curator/manager of the dataset be contacted (e.g., email address)?
Is there an erratum?
All changes to the dataset will be announced on https://webqna.github.io/
Will the dataset be updated (e.g., to correct labeling errors, add new instances, delete instances)?
All updates (if necessary) will be posted on https://webqna.github.io/
Will older versions of the dataset continue to be supported/hosted/maintained?
All changes to the dataset will be announced on https://webqna.github.io/. Outdated versions will be kept around for consistency.
If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so?
Any extension/augmentation by an external party should be made after contacting the original authors.