MultiReQA: A Cross-Domain Evaluation for Retrieval Question Answering Models

Mandy Guo, Yinfei Yang, Daniel Cer, Qinlan Shen, Noah Constant

Introduction

Retrieval-based question answering (QA) investigates the problem of finding answers to questions from an open corpus Surdeanu et al. (2008); Yang et al. (2015); Chen et al. (2017); Lee et al. (2019); Ahmad et al. (2019); Chang et al. (2020); Ma et al. (2020). There is a growing interest in building scalable end-to-end question answering systems for large scale retrieval. Retrieval question answering (ReQA) Ahmad et al. (2019), illustrated in Table 1, defines the task as directly retrieving an answer sentence from a corpus.This can be contrasted to a two stage approach that first retrieves supporting text and then identifies the correct answer span Chen et al. (2017); Lee et al. (2019) Motivated by real applications such as Google’s Talk to Books https://books.google.com/talktobooks/, where sentence-level answers from books are retrieved to answer users’ queries, ReQA is different from traditional machine reading for question answering or “reading comprehension” which aims to extract a short answer span from a given passage. Rather than just identifying answers within a short preselected passage that is provided to the model effectively by an oracle, retrieving sentence-level answers from a large pool of candidates directly addresses the real-world problem of searching for answers within a corpus. Sentences retrieved as answers in this manner can be used directly to answer questions. Alternatively, retrieved sentences, as well as possibly the passages that contains them, can be provided to a traditional Open Domain QA model Chen et al. (2017); Karpukhin et al. (2020).

We introduce a new common evaluation suite and strong baselines for ReQA across eight publicly available QA tasks. Five in-domain tasks include training and test data, while three out-of-domain tasks contain only test data. Our experiments investigate using two competitive neural models, based on BERT Devlin et al. (2019) and USE-QA Yang et al. (2019), respectively, and BM25, a strong information retrieval baseline. BM25 performs surprisingly well on many retrieval question answering tasks, achieving the best performance on two of five in-domain tasks and all three out-of-domain tasks. Neural models achieve the highest performance on three of five in-domain tasks, outperforming BM25 by a wide margin on tasks with less token overlap between question and answer. Comparing general models trained on a mixture of QA training sets to specialized in-domain models trained on a single QA task reveals that models trained jointly on multiple datasets rarely outperform those trained on only in-domain data.

Retrieval QA (ReQA)

ReQA formalizes the retrieval-based QA task as the identification of a sentence in-context that answers a provided question Ahmad et al. (2019). Retrieval QA models are evaluated using Precision at 1 (P@1) and Mean Reciprocal Rank (MRR). The P@1 score tests whether the true answer sentence appears as the top-ranked candidateRetrieval models are often measured by P@N (N=1,3,5,10). However, as our main concern is whether the question is correctly answered, we focus on P@1.. MRR, introduced for the evaluation of retrieval based QA systems Voorhees (2001); Radev et al. (2002), is calculated as MRR=1N∑i=1N1ranki\textrm{MRR}=\frac{1}{N}\sum_{i=1}^{N}\frac{1}{\textit{rank}_{i}}, where NN is the total number of questions, and ranki is the rank of the first correct answer for the iith question.

Multi-domain ReQA (MultiReQA)

The multi-domain ReQA (MultiReQA) test suite is composed of select datasets drawn from the MRQA shared task Fisch et al. (2019a).We exclude NewsQA, RACE, DROP, and DuoRC, as the majority of their questions are underspecified when taken out of their original context, making them inappropriate for a large-scale retrieval evaluations. We follow the training, in-domain test, out-of-domain test splits defined in MRQA. The individual datasets are described below:

Jeopardy question-answer pairs augmented with text snippets retrieved by Google Dunn et al. (2017).

TriviaQA

Trivia enthusiasts authored question-answer pairs. Answers are drawn from Wikipedia and Bing web search results, excluding trivia websites Joshi et al. (2017b).

HotpotQA

Wikipedia question-answer pairs. This dataset differs from the others in that the questions require reasoning over multiple supporting documents Yang et al. (2018).

SQuAD 1.1

Wikipedia question-answer pairs Rajpurkar et al. (2016a).

NaturalQuestions (NQ)

Questions are real queries issued by multiple users to Google search that retrieve a Wikipedia page in the top five search results. Answer text is drawn from the search results Kwiatkowski et al. (2019).

BioASQ

Bio-medical question-answer pairs with answers annotated by domain experts and drawn from research articles Tsatsaronis et al. (2015).

RelationExtraction (R.E.)

Entity relation question-answer pairs, created by slot filling using the WikiReading dataset Ahmad et al. (2019).

TextbookQA

Multi-modal question-answer pairs taken from middle school science curricula Kembhavi et al. (2017).

Table 2 provides example question-answer sentence pairs. Datasets are converted from a span identification task to sentence-level retrieval. The questions from the original data are used without modification. Supporting documents are split into sentences using NLTK. All resulting sentences become retrieval candidates. Answer spans are used to identify the sentences containing the correct answers. Retrieval candidates other than the sentence identified by the answer span could also provide the correct answer to a question. We investigate the prevalence of such false negatives in our subsequent analysis (6.4). As the datasets SearchQA, TriviaQA and HotpotQA contain special tags [DOC], [PAR], [SEP], and [TLE], we perform dataset-specific pre-processing to handle context splitting and tag removal. TriviaQA has [DOC] [TLE] [PAR] tags, but with no clear divisions to mark where the span of each kind of tags ends. We remove all the tags, and tokenize the article as if it does not have special tags. SearchQA uses [DOC] to separate the supporting snippets, [TLE] to mark the start of title, and [PAR] to mark start of the snippet content. We treat contents between two [DOC] tags as individual context. We then use NLTK to split the sentences within each context. The contents between [TLE] and [PAR] are used as a title feature. If the answer appears in the title feature, we do not add it as a positive answer. There are about 500 examples where the answer span is only in the title span, and we remove the corresponding questions. We follow the same procedure for HotpotQA, which uses [PAR] to separate supporting documents, and [SEP] to separate title and document content. Spans covering multiple sentence are excluded.This is typically due to sentence splitting errors by NLTK. Tables 3 and 4 provide dataset statistics.

Models

Two neural models, based on BERT Devlin et al. (2019) and USE-QA Yang et al. (2019), respectively, are evaluated on the MultiReQA test suite. Performance is contrasted with a strong term-based information retrieval baseline, BM25.

Given the strong performance of BERT Devlin et al. (2019) on many language understanding tasks, we explore adapting BERT into a dual encoder as our first neural baseline. Figure 1 illustrates our BERT dual encoder architecture. The question and answer are encoded separately. On the left side, the question is fed into a BERT transformer network, and we take the embedding output of the CLS token as the question encoding. On the right side, the answer text and context are concatenated as a long sequence, using segment IDs to separate them. The concatenated input is fed into the same BERT transformer network. As with the question encoder, we take the CLS embedding as the answer encoding. To distinguish questions and answers, we add an additional input type embedding to each input token. Note that we switch the final activation layer of the BERT CLS token from tanh to gelu. The final embeddings are l2l_{2} normalized.

We employ the BERTBASE model,The BERTBASE model uses 12 transformer layers with 12 attention heads, a hidden size of 768 and a filter size of 3072. The final embedding size is 768. due to memory constraints during training.We use in-batch negative sampling in the dual encoder training, which requires relatively large batch size. For more details of dual encoder training with negative sampling, see Gillick et al. (2018) and Guo et al. (2018).

2 Universal Sentence Encoder QA

Following Ahmad et al. (2019), we also employ Universal Sentence Encoder QA (USE-QA) Yang et al. (2019)https://tfhub.dev/google/universal-sentence-encoder-multilingual-qa/1 as a neural baseline. It is a multilingual QA retrieval model pre-trained on billions of examples from web-crawled question answering corpora. In USE-QA, the question and answer are encoded separately using a dual encoder architecture. On the left side, the question is encoded using a transformer (Vaswani et al., 2017) network with final average pooling. The pooled output is then fed into a fully-connected network. On the right side, the answer text and answer context are encoded using a transformer network and a deep averaging network (DAN) Iyyer et al. (2015) respectively. The answer text transformer encoder is shared with question transformer encoder, and employs average pooling. Then the answer encoding and context are concatenated as a single vector and fed into another fully-connected network. Both question and answer embeddings are l2l_{2} normalized before being fed into the dot product operation. The model architecture is illustrated in Figure 2. USE-QA uses a 6 layer transformer with 8 attention heads, a hidden size of 512 and a filter size of 2048. The context DAN encoder uses hidden sizes with residual connections. The feed-forward networks for question and answer both use hidden sizes , so the final dimension of the encodings is 512.

3 BM25

Term frequency inverse document frequency (TF-IDF) based methods remain the dominant method for document retrieval, with the “Best Matching 25” (BM25) family of ranking functions providing a strong baseline (Robertson and Zaragoza, 2009). In previous work on open domain question answering, BM25 has been used to retrieve evidence text, and has been shown to be a particularly strong baseline on tasks where the question is written with advance knowledge of the answer (Lee et al., 2019).

The BM25 score of document DD given query QQ which contains words q1,...,qnq_{1},...,q_{n} is given by:

where f(qi,D)f(q_{i},D) is qiq_{i}’s term frequency in the document, ∣D∣|D| is the length of the document in words, and avgdlavgdl is the average document length across all documents. Scalars k1k_{1} and bb are free parameters.

We concatenate the answer sentence and context as the document when applying BM25 for answer retrieval. In this setup, the answer sentence is duplicated twice in its context. Thus, the score for each sentence in context remains unique.

Experiments

We use the BM25 implementation in the Gensim library Řehůřek and Sojka (2010) with default k1k_{1} and bb settings. Inverse document frequency is calculated for each constructed dataset independently. We deploy two different tokenization methods for BM25: NLTK Bird et al. (2009) and a WordPiece model (wpm) Wu et al. (2016) following the BERT implementationThe wpm vocab is from BERTBASE.. Note that NLTK does not normalize the text, while the WordPiece model does. We also experimented on SQuAD with removing normalization from wpm, and found that wpm still outperforms NLTK. Our results in Table 5 for BM25word use NLTK without normalization, while BM25wpm uses wpm with normalization.

The USE-QA model was pre-trained specifically for retrieval question answering tasks. So we first evaluate the default model without any dataset specific fine-tuning. We further fine-tune the USE-QA model using the same discriminitive objective for retrieval used for the original USE-QA training Yang et al. (2019):

Where xx is the question, yy is the correct answer, Y\mathcal{Y} is all answers in the same batch that are used as sampled negatives, and ϕ(x,y)\phi(x,y) is the dot product of question and answer representations.

We fine-tune each USE-QA model on the in-domain training set using batch size 64, and SGD optimizer with learning rate decaying exponentially from 0.01 to 0.001. All model are trained 10 epochs.

BERT was pre-trained for masked language modeling and next-sentence prediction, rather than for retrieval. To adapt BERT for retrieval, we fine-tune our BERT dual encoder, with the same discriminative objective used to fine-tune the USE-QA models. We use in-batch random negative sampling with batch size 128, and the default AdamW optimizer with learning rate 0.0001. Each BERT based model is trained 10 epochs. Note that neural model hyper-parameters are tuned on a validation set (10%) split out from the training data.

2 Results

Table 5 shows baseline model performance of precision at 1 (P@1) and Mean Reciprocal Rank (MRR) on the constructed retrieval QA datasets. The highest score for each task is bolded. For P@1, the first two rows shows the results for BM25word and BM25wpm. Notably, BM25wpm performs better on 7 of 8 tasks, indicating that a careful selection of tokenization and normalization can improve the term-based model considerably. The advantage of BM25wpm is particularly noticeable on datasets where the question is constructed without seeing the answer: SearchQA, TriviaQA, NQ, BioASQ and Relation Extraction. BM25wpm also achieves the highest P@1 on 2 of 5 in-domain datasets and on all out-of-domain dataests.

The remaining rows show the results of the neural models: the off-the-shelf USE-QA model, as well as fine-tuned versions of USE-QA and the BERT dual encoder model. We finetune on each in-domain dataset separately, and the performance on the out-of-domain datasets is the average across all five fine-tuned models. The default USE-QA model is overall not competitive with BM25wpm. However when fine-tuned on in-domain data, USE-QA outperforms BM25wpm on 3 of 5 in-domain datasets. For most datasets, fine-tuned BERT (pre-trained over generic text) performed nearly as well fine-tuned USE-QA model (pre-trained over question-answer pairs). This indicates that it is not critical to pre-train on question answering data specifically. However, large-scale pre-training is still critical, as we will see in section 6.2.

We observe that the best neural models outperform BM25wpm on Hotpot and NQ by large margins: +11.68 and +12.68 on P@1 respectively. This result aligns with the statistics from Table 3, where token overlap between question and answer/context is low for these sets. For datasets with high overlap between question and answer/context, BM25wpm performs better than neural models.

The same conclusion for P@1 can be drawn for MRR, with the exception that BERTfinetune outperforms the other models on BioASQ and TextbookQA. We observe that the vocabulary of BioASQ and TextbookQA are different from the other datasets, including more specialized technical terms. Comparing with other models, the good MRR performance of BERTfinetune may be due to the better token embedding from the masked language model pre-training.

3 Transfer Learning across Domains

The previous section shows that neural models are competitive when training on in-domain data, with USE-QA slightly outperforming the BERT dual encoder. In order to better understand how fine-tuning data helps the neural models, in this section we experiment with training on different datasets, focusing on the USE-QA fine-tuned model. Table 6 shows the performance of models trained on each individual dataset, as well as a model trained jointly on all available in-domain datasets.

Each column compares the performance of different models on a specific test set. The highest numbers of each test set are bolded. Rows 1 through 5 show the results of the models trained on each in-domain dataset. In general, models trained on an individual dataset achieve the best (or near-best) performance on their own eval split, with the exception of TriviaQA. It is interesting to see that the model fine-tuned on TriviaQA performs poorly on nearly all datasets. This suggests the sentence-level training data quality from TriviaQA might be lower than other datasets. TriviaQA requires reasoning across multiple sources of evidence Joshi et al. (2017a), so sentences with annotated answer spans may not directly answer the posed question.

Rows 6 and 7 are models trained on the combined datasets. In addition to the model trained jointly on all the datasets, we also train a model without TriviaQA, given the poor performance of the model trained individually on this set. The model trained over all available data is competitive, but the performance on some datasets, e.g. NQ and SQuAD, is significantly lower than the individually-trained models. By removing TriviaQA, the combined model gets close to the individual model performance on NQ and SQuAD, and achieves the best P@1 performance on TriviaQA and TextbookQA.

Analysis

Candidate answers may be not fully interpretable when taken out of their surrounding context Ahmad et al. (2019). In this section we investigate how model performance changes when removing context. We experiment with one BM25 model and one neural model, by picking the best performing models from previous experiments: BM25wpm and USE-QAfinetune. Recall, USE-QAfinetune models are fine-tuned on each individual dataset.

Figure 3 illustrates the change in performance when models are restricted to only use the candidate answer sentence.We report P@1 here, but observed similar trends in MRR. Even without the surrounding context, both the BM25 model and USE-QA model are still able to retrieve many of the correct answers. For the USE-QA model, the performance drop is less than 5% on all datasets. The drop in BM25 performance is larger, supporting the hypothesis that BM25’s token overlap heuristic is most effective over large spans of text, while the neural model is able to provide a “deeper” semantic understanding and squeeze more signal out of a single sentence.

2 Pre-training

We perform a simple ablation by training the USE-QA model architecture directly on the fine-tuning data, starting from randomly initialized parameters. The results are shown in Table 7. The “No Pre-training” model is trained jointly on all available in-domain data, as we found that training from scratch on individual datasets performed even worse. The model performs worse than the out-of-the-box pre-trained model on seven of the eight datasets, and is dramatically worse across all datasets compare to the fine-tuned pre-trained model. This indicates that large-scale pre-training is critical for getting good QA retrieval performance from neural models. However, recalling the strong performance of the BERT dual encoder in Table 5, it is not critical to pre-train over question answering data specifically.

3 Error Analysis

In this section we examine some typical failure cases of the BM25wpm and USE-QAfinetune models. As a first observation, the two models retrieve very different answers. For example, we find that on Natural Questions, the two models’ top-ranked answers disagree on 64.75% questionsNote that even if the models retrieve different answers, both answers could still be correct.. The other datasets have similar levels of disagreement. This suggests that the models have different strengths, and that a combination of these modeling techniques could leads to a significant improvement.

Table 8 shows examples where the models retrieve different answers, and both are incorrect. In the first example, the BM25wpm retrieves the correct context by matching the keyword “Salton Sea”. But it fails to retrieve the correct sentence, as none of the keywords in the question appear in the target answer. On the other hand, the USE-QAfinetune model understands the question is asking about some sort of animal living in the sea, but fails to connect to the Salton Sea specifically. Similarly, in the second example, both models retrieve sentences that match some keywords from the question. The BM25wpm matches keywords “Spencer” and “Maine”, but misses that the question is looking for an invention. The USE-QAfinetune matches “Spencer”, and is able to connect “invent” with “discover”, but surfaces the wrong discovery.

Overall, we observe that the term based model is able to retrieve the correct context in most cases, but often fails to select the correct answer sentence, as that sentence may not have the highest token matching score with the question. On the other hand, the neural model seems to “understand” the question a little better, but sometimes fails to recognize important keywords.

Related Work

Open domain QA answers questions by querying a large collection of documents (Voorhees and Tice, 2000). Existing open domain QA datasets usually measure if a system’s output matches the ground-truth answer of the given question, often a word or a short phrase. For example, the DrQA (Chen et al., 2017) task treats Wikipedia as a knowledge base to answer factoid questions from SQuAD (Rajpurkar et al., 2016b), CuratedTREC (Baudiš and Šedivý, 2015), and other datasets. The task measures how well a system can successfully extract a string containing the answer to a question. Instead, our work follows ReQA task, and differs from this type of task by retrieving a complete sentence-level answer.

Similar to ReQA task, Seo et al. (2018) constructs a phrase-indexed QA challenge benchmark retrieving phrases, allowing for a direct F1F_{1} and exact-match evaluation on SQuAD. An extended work Seo et al. (2019) demonstrates the phrase-indexed QA system can be built using a combination of dense (neural) and sparse (term-frequency based) indices. Roy et al. (2020) investigates the retrieval of sentence-level answers from a language agnostic candidate pool. Chang et al. (2020) investigates the pre-training tasks for retrieving answers from a large scale candidate pool.

Finally, Surdeanu et al. (2008) provides a dataset consisting of 142,627 question-answer pairs from Yahoo! Answers “how to” questions, with the goal of retrieving the correct answer to a given question from the set of all answers. WikiQA (Yang et al., 2015) is another sentence-level answer selection dataset consisting of 3,047 questions and 29,258 candidate answers, split into train, dev, and test. These datasets, however, are either limited to a specific type of question, or limited to a small set of candidates.

We propose a more comprehensive eval covering multiple domains and include tasks at a much larger scale. Additionally, folding the various MRQA in-domain and out-of-domain datasets into a single eval allows us to directly investigate cross-domain generalization.

Conclusion

In this paper, we convert eight existing QA tasks from the MRQA shared task Fisch et al. (2019b) into retrieval versions, by treating the sentence containing the ground-truth span as the target sentence-level answer. We establish baselines using unsupervised term-based information retrieval methods (the BM25 ranking function), as well as two supervised neural models built on pre-trained USE-QA and BERT models. Overall, a classical term-based retrieval approach, BM25, is a strong baseline, and could likely be improved further using additional information retrieval techniques such as normalization and synonym handling. The neural models, however, can be trained end-to-end without much feature engineering, and perform particularly well on tasks with a low degree of question/answer token overlap, or in situations where the context length is limited. The neural model performance can also be improved through the addition of in-domain training data. However, we find that QA tasks are not all alike and having training data in the precise target domain is important.

Acknowledgements

We thank our teammates from Descartes and other Google groups for their feedback and suggestions, particularly DK Choe and Kelvin Guu.

References