AmbigQA: Answering Ambiguous Open-domain Questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, Luke Zettlemoyer
Introduction
In the open-domain setting, it can be difficult to formulate clear and unambiguous questions. For example, Figure 1 shows a Google search query Kwiatkowski et al. (2019) that, perhaps surprisingly, has two possible interpretations given the evidence in Wikipedia. Although open-domain question answering (QA) systems aim to answer any factoid question (Voorhees et al., 1999), existing methods assume questions have a single well-defined answer. Nonetheless, ambiguity arises frequently in open-domain QA, where questions are written during information gathering (e.g., search queries) without knowledge of the answer. As we will see in Section 4, over 50% of the questions we sampled from a set of Google search queries are ambiguous. Furthermore, identifying ambiguities is difficult both for humans and machines. As shown in Figure 1, ambiguity is a function of both the question and the evidence provided by a large text corpus.
To study this challenge, we introduce \taskname (Answering Ambiguous Open-domain Questions), a new task which involves disambiguating and answering potentially ambiguous questions. Specifically, the model must (1) find a set of distinct, equally plausible answers to the question, and (2) provide minimal yet unambiguous rewrites of the question that clarify the interpretation which leads to each answer. Figure 1 shows two such disambiguated questions and their answers.
To support the study of this task, we construct a dataset called AmbigNQ using 14,042 questions from an open-domain version of Natural Questions (Kwiatkowski et al., 2019), denoted NQ-open. For each question, annotators search for, navigate, and read multiple Wikipedia pages to find as many answers as possible. The high prevalence of ambiguity makes the task difficult even for human experts; it is inherently difficult to know if you have found every possible interpretation of a question. Nonetheless, we are able to collect high quality data covering high levels of ambiguity (2.1 distinct answers per question on average) with high estimated agreement (89.0 F1) on valid answers. The types of ambiguity are diverse and sometimes subtle (Table 1), including ambiguous entity or event references, or ambiguity over the answer type; many are only apparent after examining one or more Wikipedia pages.
To establish initial performance levels on this data, we present a set of strong baseline methods. We extend a state-of-the-art QA model (Karpukhin et al., 2020) with three new components: (1) set-based question answering with a sequence-to-sequence model, (2) a question disambiguation model, and (3) a modification to democratic co-training (Zhou and Goldman, 2004) which leverages the partial supervision available in the full NQ-open dataset. We also do an ablation study and qualitative analysis, which suggest there is significant room for future work on this task.
To summarize, our contributions are threefold.
We introduce \taskname, a new task which requires identifying all plausible answers to an open-domain question, along with disambiguated questions to differentiate them.
We construct AmbigNQ, a dataset with 14,042 annotations on NQ-open questions containing diverse types of ambiguity.
We introduce the first baseline models that produce multiple answers to open-domain questions, with experiments showing their effectiveness in learning from our data while highlighting avenues for future work.
Related Work
Open-domain Question Answering requires a system to answer any factoid question based on evidence provided by a large corpus such as Wikipedia (Voorhees et al., 1999; Chen et al., 2017). Existing benchmarks use questions of various types, from open-ended information-seeking (Berant et al., 2013; Kwiatkowski et al., 2019; Clark et al., 2019) to more specialized trivia/quiz (Joshi et al., 2017; Dunn et al., 2017). To the best of our knowledge, all existing formulations assume each question has a single clear answer.
Our work is built upon an open-domain version of Natural Questions (Kwiatkowski et al., 2019), denoted NQ-open, composed of questions posed by real users of Google search, each with an answer drawn from Wikipedia. NQ-open has promoted several recent advances in open-domain question answering (Lee et al., 2019; Asai et al., 2020; Min et al., 2019a, b; Guu et al., 2020; Karpukhin et al., 2020). Nonetheless, Kwiatkowski et al. (2019) report that the answers to such questions are often debatable, and the average agreement rate on NQ-open test data is 49.2%,111The NQ-open test data has 5-way annotations; we compute their pairwise agreement based on string match. in large part due to ambiguous questions. In this work, we embrace this ambiguity as inherent to information seeking open-domain QA, and present the first methods for returning sets of answers paired with different interpretations of the question.
Clarification Questions have been used to study question ambiguity in other settings. Research on community Q&A (Braslavski et al., 2017; Rao and Daumé III, 2018, 2019) studies finding underspecification in the question, but it does not find the answer to the original question. In recent work, Xu et al. (2019) study clarification of questions that are intentionally annotated with pre-specified entity reference ambiguities. Aliannejadi et al. (2019) and Zamani et al. (2020) use clarification questions to refine intents of simple query logs without immediately apparent information needs (e.g., single keywords like dinosaur222The average query length in Zamani et al. (2020) is 2.6.).
In contrast, we study open-domain factoid questions asked by real users: these present clear information needs, but carry diverse naturally occurring ambiguities (see Table 1). Furthermore, instead of prolonging the user’s information-seeking session with clarification questions, our task formulation provides a complete and immediate solution with unambiguous rewrites of the original question.
Question Rewriting is a novel, well-defined task which we propose for differentiating distinct answers. To the best of our knowledge, it has not been studied for resolving ambiguity; we are only aware of Elgohary et al. (2019) which use question rewriting to convert conversational questions into self-contained questions.
Task: \taskname
Figure 1 depicts the \taskname task. The input is a prompt question , and the output is a list of question-answer pairs , where each is an equally plausible answer to , and each is a minimally edited modification of whose answer is unambiguously . We consider two subtasks.
Given a question , output a set of semantically distinct and equally plausible answers , where is unknown.
Question Disambiguation.
Given and a set of answers , generate disambiguated questions , where each is a minimal edit of which makes it unambiguous so that is a correct answer and all for all are incorrect. When , this task is trivial, as .
We choose to represent ambiguity with a set of disambiguated questions because it is well-defined, immediately human-interpretable, and allows for straightforward annotation of a wide range of ambiguities without complex guidelines.
2 Evaluation Metrics
To evaluate model performance, we present several ways to compare a model prediction with question-answer pairs with a gold reference set with pairs . Since there may be more than one way to refer to a single answer (e.g., Michael Jordan and Michael Jeffrey Jordan) each gold answer is a set of acceptable answer strings, where all are disjoint.
We assign each predicted question-answer pair a correctness score based on a string similarity function valued in $$.
Intuitively, considers (1) the correctness of the answer and (2) the similarity between the predicted and reference question. We calculate F1 treating the as measures of correctness:
Data: AmbigNQ
We construct AmbigNQ using prompt questions from NQ-open and English Wikipedia as the evidence corpus. We use Amazon Mechanical Turk for crowdsourcing.
The crucial annotation challenge is maximizing recall: finding all possible distinct answers to a question. This is difficult, as ambiguities are often only apparent after carefully searching the evidence for multiple possible answers. However, we can collect high quality data with high levels of ambiguity using careful worker selection and a two stage pipeline: generation and validation.
Workers in the first stage are given a prompt question and a search box that uses the Google Search API restricted to English Wikipedia. Allowing annotators to find Wikipedia pages on their own closely approximates the real process people use to answer open-ended questions—an approach with no existing large-scale dataset.444 For instance, answers in NQ-open are annotated over pre-specified Wikipedia pages from the Google search engine.
Workers find all plausible answers to the question; when there are multiple, each answer is paired with a minimal edit of the prompt question which differentiates it from the other answers, in line with our task requirements. A distinct answer may be annotated as multiple possible spans (e.g., Michael Jordan and Michael Jeffrey Jordan).
As a special case, some questions contain temporal deixis which depends on the time of writing, e.g., “When does the new family guy season come out?”. To avoid unmanageably many answers, we instruct workers to remove the time-dependence by rewriting the prompt question for up to three most recent events before Jan 1, 2018, e.g., “When does family guy season 16 come out?” (see Table 1).
Validation.
Workers in the validation stage review the annotations provided by multiple generators. Validators mark each generator’s annotations as correct or incorrect, or provide a new set of question-answer pairs by combining the valid ones from each generator. They search Wikipedia as generators do, and are additionally given Wikipedia pages that generators viewed to speed up the process. Validation is skipped when annotated answers from all generators exactly match (37% of cases).
Quality control.
We recruit highly qualified workers through a qualification test (details in Appendix A). Although the task was difficult for most workers, we found that our highly qualified full-time workers, given quick and detailed feedback on their work, produced high accuracy and recall. For development and test data, we use two generators and one validator per prompt question. For training data, we skip validation and only use one generator per question.
Inter-annotator agreement.
2 Data Analysis
The final dataset contains 14,042 annotated examples, split consistently with NQ-open. As shown in Table 2, over 50% of development and test examples contain multiple question-answer pairs. This indicates a high rate of ambiguity in NQ-open, even though previous work has studied it with the assumption that each question has a single answer. We also find a discrepancy between development and test; this is likely due to the way in which NQ-open is constructed, which over-samples difficult questions in the test set (see Appendix B for details). The training set contains relatively fewer ambiguous examples (47%), presumably because using only one worker per training example yielded slightly lower recall.
Table 1 shows a breakdown of the types of ambiguity in AmbigNQ. They are diverse, including ambiguity in entity references, event references, properties, and answer types, with a relatively uniform distribution between them. In comparison to Xu et al. (2019), who intentionally elicit questions with ambiguous entity references, our analysis shows that unintended ambiguity comes from diverse sources. In many cases, ambiguity is not apparent from the prompt question alone, but only after researching the question on Wikipedia, as evidenced by differences in model performance (Section 6.2).
Annotator behavior.
Figures 2(a) and 2(b) show the number of unique Wikipedia pages and the number of search queries used by workers during annotation. More often than not, workers used multiple queries and navigated multiple Wikipedia pages, showing how our setup captures ambiguity in the retrieval step of open-domain question answering, which is missed in approaches that assume a pre-specified evidence document.
Distribution of edits.
Figure 2(c) shows unigram edits made to questions in the development data, where we remove stopwords except wh-words and group numeric values by the number of digits. Adding numerals such as years is common, as they can easily disambiguate entity or event references or remove time dependence. Wh-word changes are also common, especially for specifying the answer type (e.g., from who to which group; see Table 1). The distribution of edits is fairly long-tailed, with the 100 most frequent edits covering 36% of the total, and the top 1,000 covering 69%.
Model
To set initial performance levels on AmbigNQ, we present a baseline \taskname model combining ideas from recent advances in open-domain QA Karpukhin et al. (2020) and generation Lewis et al. (2020). Given a prompt question , our model predicts answers , and generates corresponding questions conditioning on , the answers , and the evidence passages. A novel co-training step also allows the model to leverage the partial supervision available in NQ-open.
Here we describe SpanSeqGen, our model for multiple answer prediction. Following Karpukhin et al. (2020), a state-of-the-art model on NQ-open, SpanSeqGen first retrieves 100 passages with a BERT-based (Devlin et al., 2019) dual encoder, and reranks them using a BERT-based cross encoder. Then, instead of predicting an answer span from the top 1 passage as Karpukhin et al. (2020) does, SpanSeqGen uses another sequence-to-sequence model based on BART (Lewis et al., 2020). Specifically, it conditions on the concatenation of and the top passages in order up to 1024 tokens, and sequentially generates distinct answers token-by-token, separated by [SEP]. We pretrain SpanSeqGen on NQ-open and finetune it on AmbigNQ.
We develop SpanSeqGen primarily because Karpukhin et al. (2020) is designed for generating a single answer, but SpanSeqGen also boosts the performance on NQ-open (41.542.2 on the test data). We include ablations on different approaches and models in Section 6.2.
Question Disambiguation.
We design a question disambiguation (QD) model based on BART. The model generates each question () conditioning on the concatenation of , the target answer , other answers , and the top passages as used by SpanSeqGen. We pretrain on NQ-open to generate questions given an answer and passage, and then finetune it on the full task data in AmbigNQ. We include ablations on different variants of the model in Section 6.2.
Co-training with weak supervision.
Given the prevalence of unlabelled ambiguity in NQ-open, we introduce a method that treats the NQ-open annotations as weak supervision and learns to discover potential ambiguity in the data. We modify a democratic co-training algorithm (Zhou and Goldman, 2004) as described in Algorithm 1. We iteratively grow the training set from AmbigNQ () with silver data from NQ-open () predicted by a majority of a set of SpanSeqGen models trained on . The key step is injecting the known answer from NQ-open as a prefix to SpanSeqGen’s output during prediction. In each step, if a majority of predict an additional answer, we assume we have found a false negative and add the result to the training set . If all models predict no additional answer, we add the example to with as a single answer.
Experiments
This baseline disambiguates the prompt question without any context from plausible answers or reference passages. Specifically, it implements the following pipeline: (1) Feed the prompt question into a BERT-based binary classifier to determine whether it is ambiguous. (2) If is ambiguous, pass it into a BART-based model which generates a sequence of disambiguated questions (), separated by [SEP]; otherwise, consider only . (3) Feed each into a state-of-the-art model on NQ-open (Karpukhin et al., 2020) to produce its answer .
Thresholding + QD.
We also include a model based on Karpukhin et al. (2020), with thresholding for multiple answer prediction and our question disambiguation (QD) model. Karpukhin et al. (2020) outputs a likelihood score for each span; we obtain by taking valid spans with likelihood larger than a hyperparameter . The model is trained to maximize the marginal likelihood of any span in the gold answer set . As with SpanSeqGen, we pretrain on NQ-open and finetune on AmbigNQ. We then produce disambiguated questions using our BART-based QD model (Section 5).
2 Results
Table 3 reports the performance of our baselines; example model outputs are provided in Table 5.
We first find that Disambig-first is significantly worse than other models. In particular, classification accuracy on whether the prompt question is ambiguous is 67%, close to the majority baseline (60%). When the model does identify an ambiguous question, its rewrites often look reasonable on the surface, but do not match the facts. For instance, in example 1 of Table 5, it asks about filming in 2017 and during season 1 for Snow White and the Huntsman, which was actually a film released in 2012. This shows that reading evidence documents is crucial for identifying and characterizing ambiguities.
There is a substantial difference in performance between development and test overall, likely due to distributional differences in the original questions in NQ-open; detailed discussion is in Appendix B.
Effect of co-training.
The last two rows of Table 3 reports the effect of our co-training method. As co-training requires multiple trained models, we compare with a naive ensemble. While we see gains from ensembling alone, an ensemble trained with the co-training method achieves the best performance on all metrics. This result demonstrates the potential of jointly using AmbigNQ and partial supervision from NQ-open.
Ablations on question disambiguation.
Performance is low overall, even given the gold answers, highlighting the challenge of the task. We think there are two major reasons. First, maximizing the likelihood of the output sequence can miss the importance of edits to the prompt question, leading the QD model to miss the information that is most important to differentiate one answer from the others. Second, there is a lack of annotated data, especially for question disambiguation which does not benefit from weakly supervised learning with NQ-open; future work can explore how to maximize the use of supervision from other available data. It is also worth noting that the metric may miss edits that are semantically correct, but phrased differently (see Table 5, example 2).
3 Zero-shot results
4 Error Analysis
Table 6 reports an analysis of predictions by SpanSeqGen with co-training, based on 50 random samples from the development data; examples can be found in the Appendix (Table 10). When there are multiple reference answers, the model rarely gets all correct answers, although often generates a subset of them. In 15 out of 20 partially correct cases, the model produces only one answer, consistent with the under-generation we found in Section 6.2. In four out of those 15 cases, the model prediction is arguably the most likely answer,888For example, a question “Who did
Conclusion & Future Work
We introduced \taskname, a new task that involves providing multiple possible answers to a potentially ambiguous open-domain question, and providing a disambiguated question corresponding to each answer. We constructed AmbigNQ, a dataset with 14,042 annotations on NQ-open questions. Our analysis shows the dataset contains diverse types of ambiguity, often not visible from the prompt question alone. We also introduced a first baseline model for producing multiple answers to open-domain questions, with experiments showing its effectiveness in learning from our data while highlighting possible areas for improvement.
Future research developing on AmbigQA models may include explicitly modeling ambiguity over events and entities or in the retrieval step, as well as improving performance on the difficult problems of answer recall and question disambiguation. Furthermore, future work may build on the AmbigQA task with more open-ended approaches such as (1) applying the approach to QA over structured data (such as ambiguous questions that require returning tables), (2) handling questions with no answer or ill-formed questions that require inferring and satisfying more complex ambiguous information needs, and (3) more carefully evaluating usefulness to end users.
Acknowledgments
This research was supported by ONR N00014-18-1-2826, DARPA N66001-19-2-403, the NSF (IIS-1252835, IIS-1562364), an Allen Distinguished Investigator Award, and the Sloan Fellowship. We thank Mandar Joshi, H2Lab members and the anonymous reviewers for their helpful comments and suggestions.
References
Appendix A Data Collection Details
We use Amazon Mechanical Turk999www.mturk.com and Spacro (Michael et al., 2018)101010github.com/julianmichael/spacro for crowdsourcing. All data was collected in February and March of 2020. We use the Google Search API111111developers.google.com/custom-search/ restricted to English Wikipedia for the search tool.
Figure 3 shows the interface used for generation and validation. We use an iframe to render Wikipedia pages in a mobile view, in order to provide the document format that they are familiar with, rather than the plain text with no formatting. When workers write the questions and the answers in the generation stage, we show appropriate error messages (e.g. when the written question is the same as the prompt question) or warning messages (e.g., when the answer is composed of more than 20 words) in order to give tight feedback. Workers produce free text answers which we instruct them to copy and paste from Wikipedia.
We pay 0.75 and 0.15 USD per prompt question for generation and validation, respectively. Generators may skip the prompt question if the answer is not found in Wikipedia, or the question is ill-formed, too subjective or too ambiguous, e.g., “When did the new tax cuts go into effect?”
Quality control.
We only recruit full-time workers that are dedicated to our task. We were able to recruit full-time workers by requiring the minimum number of HITs that can be achieved by working 40 hours a week. We also host a public website for them to monitor the validated statuses, ask questions on examples that they do not understand the validated result, or claim on the validation which is incorrect in their opinion. We found it very useful to communicate with workers, give feedback, and fix the incorrect annotations.
Inter-annotator agreement.
Appendix B Discrepancy between development and test in NQ-open
In our experiments on AmbigNQ, we found a significant discrepancy between the development and test sets. Upon further investigation, we identified that this is at least in part due to a distributional difference between the development and test sets of NQ-open, upon which we built the data. As this may be important for other researchers working on NQ-open, we detail our findings here.
Following Lee et al. (2019), NQ-open is constructed by filtering Natural Questions to questions where at least one annotator provided a non-null short answer to the question.121212 Natural Questions annotators answered each question with a set of short answers, which could be empty if there was no reasonable short answer. We refer to the empty cases as null answers. See Kwiatkowski et al. (2019) for details. While the training and development sets of NQ-open were all drawn from the training set of Natural Questions, in which one annotator answered each question, the test set of NQ-open is taken from its development set, which had five annotators per question.
This difference in number of annotators introduces a sampling bias: questions for which an annotator is less likely to find an answer are overrepresented in the NQ-open test set, in comparison to training and development. Suppose, for example, that a randomly sampled annotator has a 50% chance of producing a short answer for some question . Then has a 50% chance of making it into NQ-open’s development set, but a () 97% chance of making it into test. Concretely, when each annotator is considered independently, 34.6% of the short answer annotations in the test set of NQ-open are null answers, and the majority of annotations are null for 33.9% of questions.
As a consequence, there is a significant gap in model performance between development and test when they are evaluated under the same conditions. The official evaluation protocol for NQ-open counts a prediction as correct if it matches any of the gold reference answers. Under these conditions, the gap between development and test appears marginal (Table 8, first two columns). However, as the NQ-open test set was more comprehensively annotated than development, it has a more generous evaluation; the number of unique reference answers is 1.2 and 1.8 on development and test, respectively. In order to make the evaluation more consistent, we try evaluating models against the first reference answer only, and find a significant gap between development and test (5–8%) across all models (Table 8, last two columns).131313 It is unlikely that this discrepancy is due to overfitting on development, because the effect is consistent across models and not present on the other datasets that they are evaluated on.
Despite this discrepancy, AmbigNQ follows the setup and data split from NQ-open providing consistency with prior work. Since the AmbigNQ development and test sets were annotated under the same conditions, this discrepancy now shows up in the metrics. We leave the distribution shift of questions on the test data as one of challenges on AmbigNQ.
Appendix C Data Analysis Details
29.4% of AmbigNQ development examples do not include the NQ-open answer. We analyze a random sample of 50 such questions, and present a breakdown in Table 9. We find that our answers are correct in 92% of cases, among which 44% of disagreements are due to mismatched spans, 22% are due to the NQ-open answer being incorrect, and 14% are due to time-dependence in the question. Of the 8% of cases where our answer is incorrect, the NQ-open answers are also incorrect over half the time, indicating that these may be difficult questions.
Appendix D Baseline Implementation Details
We use English Wikipedia dump from 2018-12-20 and 2020-01-20 for NQ-open and AmbigNQ, respectively. Following Karpukhin et al. (2020), we take the plain text and split passages to be up to 100 words each.
Model implementation.
Details in ensemble and co-training.
Appendix E Error Analysis of SpanSeqGen
Table 10 reports an analysis of predictions by SpanSeqGen, on 50 random samples from the development set. We refer to Section 6.4 for the discussions.