Generation-Augmented Retrieval for Open-domain Question Answering
Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, Weizhu Chen
Introduction
Open-domain question answering (OpenQA) aims to answer factoid questions without a pre-specified domain and has numerous real-world applications. In OpenQA, a large collection of documents (e.g., Wikipedia) are often used to seek information pertaining to the questions. One of the most common approaches uses a retriever-reader architecture Chen et al. (2017), which first retrieves a small subset of documents using the question as the query and then reads the retrieved documents to extract (or generate) an answer. The retriever is crucial as it is infeasible to examine every piece of information in the entire document collection (e.g., millions of Wikipedia passages) and the retrieval accuracy bounds the performance of the (extractive) reader.
Early OpenQA systems Chen et al. (2017) use classic retrieval methods such as TF-IDF and BM25 with sparse representations. Sparse methods are lightweight and efficient, but unable to perform semantic matching and fail to retrieve relevant passages without lexical overlap. More recently, methods based on dense representations Guu et al. (2020); Karpukhin et al. (2020) learn to embed queries and passages into a latent vector space, in which text similarity beyond lexical overlap can be measured. Dense retrieval methods can retrieve semantically relevant but lexically different passages and often achieve better performance than sparse methods. However, the dense models are more computationally expensive and suffer from information loss as they condense the entire text sequence into a fixed-size vector that does not guarantee exact matching Luan et al. (2020).
There have been some recent studies on query reformulation with text generation for other retrieval tasks, which, for example, rewrite the queries to context-independent Yu et al. (2020); Lin et al. (2020); Vakulenko et al. (2020) or well-formed Liu et al. (2019) ones. However, these methods require either task-specific data (e.g., conversational contexts, ill-formed queries) or external resources such as paraphrase data Zaiem and Sadat (2019); Wang et al. (2020) that cannot or do not transfer well to OpenQA. Also, some rely on time-consuming training process like reinforcement learning (RL) Nogueira and Cho (2017); Liu et al. (2019); Wang et al. (2020) that is not efficient enough for OpenQA (more discussions in Sec. 2).
In this paper, we propose Generation-Augmented Retrieval (Gar), which augments a query through text generation of a pre-trained language model (PLM). Different from prior studies that reformulate queries, Gar does not require external resources or downstream feedback via RL as supervision, because it does not rewrite the query but expands it with heuristically discovered relevant contexts, which are fetched from PLMs and provide richer background information (Table 2). For example, by prompting a PLM to generate the title of a relevant passage given a query and appending the generated title to the query, it becomes easier to retrieve that relevant passage. Intuitively, the generated contexts explicitly express the search intent not presented in the original query. As a result, Gar with sparse representations achieves comparable or even better performance than state-of-the-art approaches Karpukhin et al. (2020); Guu et al. (2020) with dense representations of the original queries, while being more lightweight and efficient in terms of both training and inference (including the cost of the generation model) (Sec. 6.4).
Specifically, we expand the query (question) by adding relevant contexts as follows. We conduct seq2seq learning with the question as the input and various freely accessible in-domain contexts as the output such as the answer, the sentence where the answer belongs to, and the title of a passage that contains the answer. We then append the generated contexts to the question as the generation-augmented query for retrieval. We demonstrate that using multiple contexts from diverse generation targets is beneficial as fusing the retrieval results of different generation-augmented queries consistently yields better retrieval accuracy.
We conduct extensive experiments on the Natural Questions (NQ) Kwiatkowski et al. (2019) and TriviaQA (Trivia) Joshi et al. (2017) datasets. The results reveal four major advantages of Gar: (1) Gar, combined with BM25, achieves significant gains over the same BM25 model that uses the original queries or existing unsupervised query expansion (QE) methods. (2) Gar with sparse representations (BM25) achieves comparable or even better performance than the current state-of-the-art retrieval methods, such as DPR Karpukhin et al. (2020), that use dense representations. (3) Since Gar uses sparse representations to measure lexical overlapStrictly speaking, Gar with sparse representations handles semantics before retrieval by enriching the queries, while maintaining the advantage of exact matching., it is complementary to dense representations: by fusing the retrieval results of Gar and DPR (denoted as ), we obtain consistently better performance than either method used individually. (4) Gar outperforms DPR in the end-to-end QA performance (EM) when the same extractive reader is used: EM=41.8 (43.8 for ) on NQ and 62.7 on Trivia, creating new state-of-the-art results for extractive OpenQA. Gar also outperforms other retrieval methods under the generative setup when the same generative reader is used: EM=38.1 (45.3 for ) on NQ and 62.2 on Trivia.
Contributions. (1) We propose Generation-Augmented Retrieval (Gar), which augments queries with heuristically discovered relevant contexts through text generation without external supervision or time-consuming downstream feedback. (2) We show that using generation-augmented queries achieves significantly better retrieval and QA results than using the original queries or existing unsupervised QE methods. (3) We show that Gar, combined with a simple BM25 model, achieves new state-of-the-art performance on two benchmark datasets in extractive OpenQA and competitive results in the generative setting.
Related Work
Conventional Query Expansion. Gar shares some merits with query expansion (QE) methods based on pseudo relevance feedback Rocchio (1971); Abdul-Jaleel et al. (2004); Lv and Zhai (2010) in that they both expand the queries with relevant contexts (terms) without the use of external supervision. Gar is superior as it expands the queries with knowledge stored in the PLMs rather than the retrieved passages and its expanded terms are learned through text generation.
Recent Query Reformulation. There are recent or concurrent studies Nogueira and Cho (2017); Zaiem and Sadat (2019); Yu et al. (2020); Vakulenko et al. (2020); Lin et al. (2020) that reformulate queries with generation models for other retrieval tasks. However, these studies are not easily applicable or efficient enough for OpenQA because: (1) They require external resources such as paraphrase data Zaiem and Sadat (2019), search sessions Yu et al. (2020), or conversational contexts Lin et al. (2020); Vakulenko et al. (2020) to form the reformulated queries, which are not available or showed inferior domain-transfer performance in OpenQA Zaiem and Sadat (2019); (2) They involve time-consuming training process such as RL. For example, Nogueira and Cho (2017) reported a training time of 8 to 10 days as it uses retrieval performance in the reward function and conducts retrieval at each iteration. In contrast, Gar uses freely accessible in-domain contexts like passage titles as the generation targets and standard seq2seq learning, which, despite its simplicity, is not only more efficient but effective for OpenQA.
Retrieval for OpenQA. Existing sparse retrieval methods for OpenQA Chen et al. (2017) solely rely on the information of the questions. Gar extends to contexts relevant to the questions by extracting information inside PLMs and helps sparse methods achieve comparable or better performance than dense methods Guu et al. (2020); Karpukhin et al. (2020), while enjoying the simplicity and efficiency of sparse representations. Gar can also be used with dense representations to seek for even better performance, which we leave as future work.
Generative QA. Generative QA generates answers through seq2seq learning instead of extracting answer spans. Recent studies on generative OpenQA Lewis et al. (2020a); Min et al. (2020); Izacard and Grave (2020) are orthogonal to Gar in that they focus on improving the reading stage and directly reuse DPR Karpukhin et al. (2020) as the retriever. Unlike generative QA, the goal of Gar is not to generate perfect answers to the questions but pertinent contexts that are helpful for retrieval. Another line in generative QA learns to generate answers without relevant passages as the evidence but solely the question itself using PLMs Roberts et al. (2020); Brown et al. (2020). Gar further confirms that one can extract factual knowledge from PLMs, which is not limited to the answers as in prior studies but also other relevant contexts.
Generation-Augmented Retrieval
OpenQA aims to answer factoid questions without pre-specified domains. We assume that a large collection of documents (i.e., Wikipedia) are given as the resource to answer the questions and a retriever-reader architecture is used to tackle the task, where the retriever retrieves a small subset of the documents and the reader reads the documents to extract (or generate) an answer. Our goal is to improve the effectiveness and efficiency of the retriever and consequently improve the performance of the reader.
2 Generation of Query Contexts
In Gar, queries are augmented with various heuristically discovered relevant contexts in order to retrieve more relevant passages in terms of both quantity and quality. For the task of OpenQA where the query is a question, we take the following three freely accessible contexts as the generation targets. We show in Sec. 6.2 that having multiple generation targets is helpful in that fusing their results consistently brings better retrieval accuracy.
Context 1: The default target (answer). The default target is the label in the task of interest, which is the answer in OpenQA. The answer to the question is apparently useful for the retrieval of relevant passages that contain the answer itself. As shown in previous work Roberts et al. (2020); Brown et al. (2020), PLMs are able to answer certain questions solely by taking the questions as input (i.e., closed-book QA). Instead of using the generated answers directly as in closed-book QA, Gar treats them as contexts of the question for retrieval. The advantage is that even if the generated answers are partially correct (or even incorrect), they may still benefit retrieval as long as they are relevant to the passages that contain the correct answers (e.g., co-occur with the correct answers).
Context 2: Sentence containing the default target. The sentence in a passage that contains the answer is used as another generation target. Similar to using answers as the generation target, the generated sentences are still beneficial for retrieving relevant passages even if they do not contain the answers, as their semantics is highly related to the questions/answers (examples in Sec. 6.1). One can take the relevant sentences in the ground-truth passages (if any) or those in the positive passages of a retriever as the reference, depending on the trade-off between reference quality and diversity.
Context 3: Title of passage containing the default target. One can also use the titles of relevant passages as the generation target if available. Specifically, we retrieve Wikipedia passages using BM25 with the question as the query, and take the page titles of positive passages that contain the answers as the generation target. We observe that the page titles of positive passages are often entity names of interest, and sometimes (but not always) the answers to the questions. Intuitively, if Gar learns which Wikipedia pages the question is related to, the queries augmented by the generated titles would naturally have a better chance of retrieving those relevant passages.
While it is likely that some of the generated query contexts involve unfaithful or nonfactual information due to hallucination in text generation Mao et al. (2020) and introduce noise during retrieval, they are beneficial rather than harmful overall, as our experiments show that Gar improve both retrieval and QA performance over BM25 significantly. Also, since we generate 3 different (complementary) query contexts and fuse their retrieval results, the distraction of hallucinated content is further alleviated.
3 Retrieval with Generation-Augmented Queries
After generating the contexts of a query, we append them to the query to form a generation-augmented query.One may create a title field during document indexing and conduct multi-field retrieval but here we append the titles to the questions as other query contexts for generalizability. We observe that conducting retrieval with the generated contexts (e.g., answers) alone as queries instead of concatenation is ineffective because (1) some of the generated answers are rather irrelevant, and (2) a query consisting of the correct answer alone (without the question) may retrieve false positive passages with unrelated contexts that happen to contain the answer. Such low-quality passages may lead to potential issues in the following passage reading stage.
If there are multiple query contexts, we conduct retrieval using queries with different generated contexts separately and then fuse their results. The performance of one-time retrieval with all the contexts appended is slightly but not significantly worse. For simplicity, we fuse the retrieval results in a straightforward way: an equal number of passages are taken from the top-retrieved passages of each source. One may also use weighted or more sophisticated fusion strategies such as reciprocal rank fusion Cormack et al. (2009), the results of which are slightly better according to our experiments.We use the fusion tools at https://github.com/joaopalotti/trectools.
Next, one can use any off-the-shelf retriever for passage retrieval. Here, we use a simple BM25 model to demonstrate that Gar with sparse representations can already achieve comparable or better performance than state-of-the-art dense methods while being more lightweight and efficient (including the cost of the generation model), closing the gap between sparse and dense retrieval methods.
OpenQA with Gar
To further verify the effectiveness of Gar, we equip it with both extractive and generative readers for end-to-end QA evaluation. We follow the reader design of the major baselines for a fair comparison, while virtually any existing QA reader can be used with Gar.
For the extractive setup, we largely follow the design of the extractive reader in DPR Karpukhin et al. (2020). Let denote the list of retrieved passages with passage relevance scores . Let denote the top text spans in passage ranked by span relevance scores . Briefly, the DPR reader uses BERT-base Devlin et al. (2019) for representation learning, where it estimates the passage relevance score for each retrieved passage based on the [CLS] tokens of all retrieved passages , and assigns span relevance scores for each candidate span based on the representations of its start and end tokens. Finally, the span with the highest span relevance score from the passage with the highest passage relevance score is chosen as the answer. We refer the readers to Karpukhin et al. (2020) for more details.
Passage-level Span Voting. Many extractive QA methods Chen et al. (2017); Min et al. (2019b); Guu et al. (2020); Karpukhin et al. (2020) measure the probability of span extraction in different retrieved passages independently, despite that their collective signals may provide more evidence in determining the correct answer. We propose a simple yet effective passage-level span voting mechanism, which aggregates the predictions of the spans in the same surface form from different retrieved passages. Intuitively, if a text span is considered as the answer multiple times in different passages, it is more likely to be the correct answer. Specifically, Gar calculates a normalized score for the j-th span in passage during inference as follows: . Gar then aggregates the scores of the spans with the same surface string among all the retrieved passages as the collective passage-level score.We find that the number of spans used for normalization in each passage does not have significant impact on the final performance (we take ) and using the raw or normalized strings for aggregation also perform similarly.
2 Generative Reader
For the generative setup, we use a seq2seq framework where the input is the concatenation of the question and top-retrieved passages and the target output is the desired answer. Such generative readers are adopted in recent methods such as SpanSeqGen Min et al. (2020) and Longformer Beltagy et al. (2020). Specifically, we use BART-large Lewis et al. (2019) as the generative reader, which concatenates the question and top-retrieved passages up to its length limit (1,024 tokens, 7.8 passages on average). Generative Gar is directly comparable with SpanSeqGen Min et al. (2020) that uses the retrieval results of DPR but not comparable with Fusion-in-Decoder (FID) Izacard and Grave (2020) since it encodes 100 passages rather than 1,024 tokens and involves more model parameters.
Experiment Setup
We conduct experiments on the open-domain version of two popular QA benchmarks: Natural Questions (NQ) Kwiatkowski et al. (2019) and TriviaQA (Trivia) Joshi et al. (2017). The statistics of the datasets are listed in Table 1.
2 Evaluation Metrics
Following prior studies Karpukhin et al. (2020), we use top-k retrieval accuracy to evaluate the performance of the retriever and the Exact Match (EM) score to measure the performance of the reader.
Top-k retrieval accuracy is defined as the proportion of questions for which the top-k retrieved passages contain at least one answer span, which is an upper bound of how many questions are “answerable” by an extractive reader.
Exact Match (EM) is the proportion of the predicted answer spans being exactly the same as (one of) the ground-truth answer(s), after string normalization such as article and punctuation removal.
3 Compared Methods
For passage retrieval, we mainly compare with BM25 and DPR, which represent the most used state-of-the-art methods of sparse and dense retrieval for OpenQA, respectively. For query expansion, we re-emphasize that Gar is the first QE approach designed for OpenQA and most of the recent approaches are not applicable or efficient enough for OpenQA since they have task-specific objectives, require external supervision that was shown to transfer poorly to OpenQA, or take many days to train (Sec. 2). We thus compare with a classic unsupervised QE method RM3 Abdul-Jaleel et al. (2004) that does not need external resources for a fair comparison. For passage reading, we compare with both extractive (Min et al., 2019a; Asai et al., 2019; Lee et al., 2019; Min et al., 2019b; Guu et al., 2020; Karpukhin et al., 2020) and generative (Brown et al., 2020; Roberts et al., 2020; Min et al., 2020; Lewis et al., 2020a; Izacard and Grave, 2020) methods when equipping Gar with the corresponding reader.
4 Implementation Details
Retriever. We use Anserini Yang et al. (2017) for text retrieval of BM25 and Gar with its default parameters. We conduct grid search for the QE baseline RM3 Abdul-Jaleel et al. (2004).
Generator. We use BART-large Lewis et al. (2019) to generate query contexts in Gar. When there are multiple desired targets (such as multiple answers or titles), we concatenate them with [SEP] tokens as the reference and remove the [SEP] tokens in the generation-augmented queries. For Trivia, in particular, we use the value field as the generation target of answer and observe better performance. We take the checkpoint with the best ROUGE-1 F1 score on the validation set, while observing that the retrieval accuracy of Gar is relatively stable to the checkpoint selection since we do not directly use the generated contexts but treat them as augmentation of queries for retrieval.
Reader. Extractive Gar uses the reader of DPR with largely the same hyperparameters, which is initialized with BERT-base Devlin et al. (2019) and takes 100 (500) retrieved passages during training (inference). Generative Gar concatenates the question and top-10 retrieved passages, and takes at most 1,024 tokens as input. Greedy decoding is adopted for all generation models, which appears to perform similarly to (more expensive) beam search.
Experiment Results
We evaluate the effectiveness of Gar in three stages: generation of query contexts (Sec. 6.1), retrieval of relevant passages (Sec. 6.2), and passage reading for OpenQA (Sec. 6.3). Ablation studies are mostly shown on the NQ dataset to understand the drawbacks of Gar since it achieves better performance on Trivia.
Automatic Evaluation. To evaluate the quality of the generated query contexts, we first measure their lexical overlap with the ground-truth query contexts. As suggested by the nontrivial ROUGE scores in Table 3, Gar does learn to generate meaningful query contexts that could help the retrieval stage. We next measure the lexical overlap between the query and the ground-truth passage. The ROUGE-1/2/L F1 scores between the original query and ground-truth passage are 6.00/2.36/5.01, and those for the generation-augmented query are 7.05/2.84/5.62 (answer), 13.21/6.99/10.27 (sentence), 7.13/2.85/5.76 (title) on NQ, respectively. Such results further demonstrate that the generated query contexts significantly increase the word overlap between the queries and the positive passages, and thus are likely to improve retrieval results.We use F1 instead of recall to avoid the unfair favor of (longer) generation-augmented query.
Case Studies. In Table 2, we show several examples of the generated query contexts and their ground-truth references. In the first example, the correct album release date appears in both the generated answer and the generated sentence, and the generated title is the same as the Wikipedia page title of the album. In the last two examples, the generated answers are wrong but fortunately, the generated sentences contain the correct answer and (or) other relevant information and the generated titles are highly related to the question as well, which shows that different query contexts are complementary to each other and the noise during query context generation is thus reduced.
2 Generation-Augmented Retrieval
Comparison w. the state-of-the-art. We next evaluate the effectiveness of Gar for retrieval. In Table 4, we show the top-k retrieval accuracy of BM25, BM25 with query expansion (+RM3) Abdul-Jaleel et al. (2004), DPR (Karpukhin et al., 2020), Gar, and (Gar +DPR).
On the NQ dataset, while BM25 clearly underperforms DPR regardless of the number of retrieved passages, the gap between Gar and DPR is significantly smaller and negligible when . When , Gar is slightly better than DPR despite that it simply uses BM25 for retrieval. In contrast, the classic QE method RM3, while showing marginal improvement over the vanilla BM25, does not achieve comparable performance with Gar or DPR. By fusing the results of Gar and DPR in the same way as described in Sec. 3.3, we further obtain consistently higher performance than both methods, with top-100 accuracy 88.9% and top-1000 accuracy 93.2%.
On the Trivia dataset, the results are even more encouraging – Gar achieves consistently better retrieval accuracy than DPR when . On the other hand, the difference between BM25 and BM25 +RM3 is negligible, which suggests that naively considering top-ranked passages as relevant (i.e., pseudo relevance feedback) for QE does not always work for OpenQA. Results on more cutoffs of can be found in App. A.
Effectiveness of diverse query contexts. In Fig. 1, we show the performance of Gar when different query contexts are used to augment the queries. Although the individual performance when using each query context is somewhat similar, fusing their retrieved passages consistently leads to better performance, confirming that different generation-augmented queries are complementary to each other (recall examples in Table 2).
Performance breakdown by question type. In Table 5, we show the top-100 accuracy of the compared retrieval methods per question type on the NQ test set. Again, Gar outperforms BM25 on all types of questions significantly and achieves the best performance across the board, which further verifies the effectiveness of Gar.
3 Passage Reading with Gar
Comparison w. the state-of-the-art. We show the comparison of end-to-end QA performance of extractive and generative methods in Table 6. Extractive Gar achieves state-of-the-art performance among extractive methods on both NQ and Trivia datasets, despite that it is more lightweight and computationally efficient. Generative Gar outperforms most of the generative methods on Trivia but does not perform as well on NQ, which is somewhat expected and consistent with the performance at the retrieval stage, as the generative reader only takes a few passages as input and Gar does not outperform dense retrieval methods on NQ when is very small. However, combining Gar with DPR achieves significantly better performance than both methods or baselines that use DPR as input such as SpanSeqGen (Min et al., 2020) and RAG (Lewis et al., 2020a). Also, Gar outperforms BM25 significantly under both extractive and generative setups, which again shows the effectiveness of the generated query contexts, even if they are heuristically discovered without any external supervision.
The best performing generative method FID Izacard and Grave (2020) is not directly comparable as it takes more (100) passages as input. As an indirect comparison, Gar performs better than FID when FID encodes 10 passages (cf. Fig. 2 in Izacard and Grave (2020)). Moreover, since FID relies on the retrieval results of DPR as well, we believe that it is a low-hanging fruit to replace its input with Gar or and further boost the performance.This claim is later verified by the best systems in the NeurIPS 2020 EfficientQA competition Min et al. (2021). We also observe that, perhaps surprisingly, extractive BM25 performs reasonably well, especially on the Trivia dataset, outperforming many recent state-of-the-art methods.We find that taking 500 passages during reader inference instead of 100 as in Karpukhin et al. (2020) improves the performance of BM25 but not DPR. Generative BM25 also performs competitively in our experiments.
Model Generalizability. Recent studies Lewis et al. (2020b) show that there are significant question and answer overlaps between the training and test sets of popular OpenQA datasets. Specifically, 60% to 70% test-time answers also appear in the training set and roughly 30% test-set questions have a near-duplicate paraphrase in the training set. Such observations suggest that many questions might have been answered by simple question or answer memorization. To further examine model generalizability, we study the per-category performance of different methods using the annotations in Lewis et al. (2020b).
As listed in Table 7, for the No Overlap category, (E) outperforms DPR on the extractive setup and (G) outperforms RAG on the generative setup, which indicates that better end-to-end model generalizability can be achieved by adding Gar for retrieval. also achieves the best EM under the Answer Overlap Only category. In addition, we observe that a closed-book BART model that only takes the question as input performs much worse than additionally taking top-retrieved passages, i.e., (G), especially on the questions that require generalizability. Notably, all methods perform significantly better on the Question Overlap category, which suggests that the high Total EM is mostly contributed by question memorization. That said, appears to be less dependent on question memorization given its lower EM for this category.The same ablation study is also conducted on the retrieval stage and similar results are observed. More detailed discussions can be found in App. A.
4 Efficiency of Gar
Gar is efficient and scalable since it uses sparse representations for retrieval and does not involve time-consuming training process such as RL Nogueira and Cho (2017); Liu et al. (2019). The only overhead of Gar is on the generation of query contexts and the retrieval with generation-augmented (thus longer) queries, whose computational complexity is significantly lower than other methods with comparable retrieval accuracy.
We use Nvidia V100 GPUs and Intel Xeon Platinum 8168 CPUs in our experiments. As listed in Table 8, the training time of Gar is 3 to 6 hours on 1 GPU depending on the generation target. As a comparison, REALM Guu et al. (2020) uses 64 TPUs to train for 200k steps during pre-training alone and DPR Karpukhin et al. (2020) takes about 24 hours to train with 8 GPUs. To build the indices of Wikipedia passages, Gar only takes around 30 min with 35 CPUs, while DPR takes 8.8 hours on 8 GPUs to generate dense representations and another 8.5 hours to build the FAISS index Johnson et al. (2017). For retrieval, Gar takes about 1 min to generate one query context with 1 GPU, 1 min to retrieve 1,000 passages for the NQ test set with answer/title-augmented queries and 2 min with sentence-augmented queries using 35 CPUs. In contrast, DPR takes about 30 min on 1 GPU.
Conclusion
In this work, we propose Generation-Augmented Retrieval and demonstrate that the relevant contexts generated by PLMs without external supervision can significantly enrich query semantics and improve retrieval accuracy. Remarkably, Gar with sparse representations performs similarly or better than state-of-the-art methods based on the dense representations of the original queries. Gar can also be easily combined with dense representations to produce even better results. Furthermore, Gar achieves state-of-the-art end-to-end performance on extractive OpenQA and competitive performance under the generative setup.
Future Extensions
Potential improvements. There is still much space to explore and improve for Gar in future work. For query context generation, one can explore multi-task learning to further reduce computational cost and examine whether different contexts can mutually enhance each other when generated by the same generator. One may also sample multiple contexts instead of greedy decoding to enrich a query. For retrieval, one can adopt more advanced fusion techniques based on both the ranking and score of the passages. As the generator and retriever are largely independent now, it is also interesting to study how to jointly or iteratively optimize generation and retrieval such that the generator is aware of the retriever and generates query contexts more beneficial for the retrieval stage. Last but not least, it is very likely that better results can be obtained by more extensive hyper-parameter tuning.
Applicability to other tasks. Beyond OpenQA, Gar also has great potentials for other tasks that involve text matching such as conversation utterance selection Lowe et al. (2015); Dinan et al. (2020) or information retrieval Nguyen et al. (2016); Craswell et al. (2020). The default generation target is always available for supervised tasks. For example, for conversation utterance selection one can use the reference utterance as the default target and then match the concatenation of the conversation history and the generated utterance with the provided utterance candidates. For article search, the default target could be (part of) the ground-truth article itself. Other generation targets are more task-specific and can be designed as long as they can be fetched from the latent knowledge inside PLMs and are helpful for further text retrieval (matching). Note that by augmenting (expanding) the queries with heuristically discovered relevant contexts extracted from PLMs instead of reformulating them, Gar bypasses the need for external supervision to form the original-reformulated query pairs.
Acknowledgments
We thank Vladimir Karpukhin, Sewon Min, Gautier Izacard, Wenda Qiu, Revanth Reddy, and Hao Cheng for helpful discussions. We thank the anonymous reviewers for valuable comments.
References
Appendix A More Analysis of Retrieval Performance
We show the detailed results of top-k retrieval accuracy of the compared methods in Figs. 2 and 3. Gar performs comparably or better than DPR when on NQ and on Trivia.
We show in Table 9 the retrieval accuracy breakdown using the question-answer overlap categories. The most significant gap between BM25 and other methods is on the Question Overlap category, which coincides with the fact that BM25 is unable to conduct question paraphrasing (semantic matching). Gar helps BM25 to bridge the gap by providing the query contexts and even outperform DPR in this category. Moreover, Gar consistently improves over BM25 on other categories and outperforms DPR as well.